A voice signal processing method and a related device
By obtaining the vibration signal when the user speaks and combining it with the voice signal collected by the sensor, and using adaptive filtering and recurrent neural network models, the problem of insufficient robustness of the existing voice interaction system in separating user voice signals and environmental noise is solved, and better voice recognition effect is achieved.
Patent Information
- Application Number
- CN202080026583.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-05-29
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2040-05-29
AI Technical Summary
Existing voice interaction systems lack robustness in separating user voice signals from environmental noise, making it difficult to effectively suppress environmental noise interference, which affects voice recognition effects.
By obtaining the vibration signal when the user speaks, combining it with the voice signal collected by the sensor, and using the vibration signal as a basis, based on adaptive filtering and recurrent neural network model, the environmental noise interference is suppressed and the target voice information is extracted.
It effectively suppresses environmental noise interference, improves the accuracy and robustness of speech recognition, and achieves better speech recognition effects.
Smart Images

Figure CN114072875B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of audio processing, and in particular to a method for processing a speech signal and related equipment. Background Art
[0002] Human-computer interaction (HCI) primarily studies the information exchange between humans and computers. It mainly includes two parts: human-to-computer and computer-to-human information exchange. It is an interdisciplinary subject closely related to cognitive psychology, ergonomics, multimedia technology, virtual reality technology, etc. In human-computer interaction technology, multimodal interactive devices are interactive devices that use multiple interactive modes in parallel, such as voice interaction, somatosensory interaction, and touch interaction. Human-computer interaction based on multimodal interactive devices: user information is collected through multiple tracking modules (face, gesture, posture, voice, and rhythm) in the interactive device, and a virtual user expression module is formed after understanding, processing, and management to interact with the computer, which can greatly enhance the user's interactive experience.
[0003] With the development of voice technology, many smart devices (such as mobile phones, companion robots, in-vehicle devices, smart speakers, and smart voice assistants) can interact with users through voice. The voice interaction system of smart devices recognizes the user's voice and completes the user's instructions. In these smart devices, microphones are typically used to pick up audio signals from the environment. Audio signals are a mixture of environmental signals. In addition to the voice signal from the user that the smart device hopes to pick up, other signals such as ambient noise and other people's voices are also included.
[0004] In existing implementations, in order to extract the voice signal from a specific user from a mixed signal, a blind separation method can be adopted. This method is essentially a statistical method to separate the sound source. Therefore, it is limited by the actual modeling method and poses great challenges in terms of robustness. Summary of the Invention
[0005] In a first aspect, the present application provides a method for processing a speech signal, the method comprising:
[0006] Obtain user voice signals collected by sensors;
[0007] It should be noted that the user voice signal should not be understood as just the words spoken by the user, but rather as the voice signal including the user's voice. The inclusion of ambient noise in the voice signal can be understood as the presence of a speaking user and other ambient noise (such as other people speaking) in the environment. In this case, the collected voice signal includes the intertwined user voice and ambient noise, and the relationship between the voice signal and ambient noise should not be understood as a simple superposition. In other words, ambient noise should not be understood as an independent signal in the voice signal.
[0008] Obtaining a vibration signal corresponding to the user uttering the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part being a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state;
[0009] It should be noted that the vibration signal corresponding to the user's voice can be obtained based on video extraction.
[0010] Target voice information is obtained according to the vibration signal and the user voice signal collected by the sensor.
[0011] The embodiment of the present application provides a method for processing a speech signal, comprising: obtaining a speech signal of a user collected by a sensor, wherein the speech signal includes environmental noise; obtaining a vibration signal corresponding to when the user utters the speech; wherein the vibration signal is used to represent the vibration characteristics of a body part of the user; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and obtaining target speech information based on the vibration signal and the user speech signal collected by the sensor. In the above manner, the vibration signal is used as the basis for speech recognition. Since the vibration signal does not contain external non-user speech mixed in during complex acoustic transmission, it is less affected by other environmental noises (such as reverberation). Therefore, this part of noise interference can be relatively well suppressed, and a better speech recognition effect can be achieved.
[0012] In an optional implementation, the vibration signal is used to represent a vibration feature corresponding to the vibration generated by the user uttering the voice.
[0013] In an optional implementation, the body part includes at least one of the following: the top of the skull, the face, the throat, or the neck.
[0014] In an optional implementation, obtaining the vibration signal corresponding to the user uttering the voice includes: obtaining a video frame including the user; and extracting the vibration signal corresponding to the user uttering the voice based on the video frame.
[0015] In an optional implementation, the video frames are acquired through a dynamic vision sensor and / or a high-speed camera.
[0016] In an optional implementation, obtaining the target voice information based on the vibration signal and the user voice signal collected by the sensor includes: obtaining a corresponding target audio signal based on the vibration signal; filtering the target audio signal from the user voice signal collected by the sensor based on filtering to obtain a signal to be filtered; filtering the signal to be filtered from the user voice signal collected by the sensor to obtain the target voice information.
[0017] Specifically, the corresponding target audio signal can be restored according to the vibration signal, and based on filtering, the target audio signal is filtered out from the audio signal to obtain a noise signal. After filtering, the filtered signal z'(n) no longer contains the useful signal x'(n), which is basically the external noise except the user's target audio signal s(n); optionally, if multiple cameras (DVS, high-speed camera, etc.) pick up the vibration of a certain person, the target audio signals x1'(n), x2'(n), x3'(n), and x4'(n) recovered from these vibrations are filtered out from the mixed audio signal z(n) in turn according to the above-mentioned adaptive filtering method, that is, the mixed audio signal z'(n) is obtained from which the various x1'(n), x2'(n), x3'(n), and x4'(n) audio components are removed.
[0018] In an optional implementation, the method further includes: obtaining, based on the target voice information, instruction information corresponding to the user's voice signal, the instruction information indicating a semantic intent contained in the user's voice signal. The instruction information can be used to trigger a function corresponding to the semantic intent contained in the user's voice signal, such as opening an application, making a voice call, etc.
[0019] In an optional implementation, obtaining target voice information according to the vibration signal and the user voice signal collected by the sensor includes:
[0020] Based on the vibration signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model; or
[0021] According to the vibration signal, a corresponding target audio signal is obtained; based on the target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
[0022] In an optional implementation, the method further includes:
[0023] Obtaining a brain wave signal of the user corresponding to when the user utters the voice;
[0024] Accordingly, obtaining target voice information according to the vibration signal and the user voice signal collected by the sensor includes:
[0025] Target voice information is obtained according to the vibration signal, the brain wave signal and the user voice signal collected by the sensor.
[0026] In an optional implementation, the method further includes:
[0027] Acquiring, based on the brainwave signal, a motion signal of a vocal tract occlusion portion when the user utters a voice; correspondingly, acquiring target voice information based on the vibration signal, the brainwave signal, and the user voice signal collected by the sensor, includes:
[0028] Target voice information is obtained according to the vibration signal, the motion signal, and the user voice signal collected by the sensor.
[0029] In an optional implementation, obtaining target voice information according to the vibration signal, the brain wave signal, and the user voice signal collected by the sensor includes:
[0030] Based on the vibration signal, the brain wave signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model; or,
[0031] Acquire a corresponding first target audio signal according to the vibration signal;
[0032] According to the brain wave signal, a corresponding second target audio signal is obtained; based on the first target audio signal, the second target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
[0033] In an optional implementation, the target voice information includes a voiceprint feature representing a voice signal of the user.
[0034] In a second aspect, the present application provides a method for processing a speech signal, the method comprising:
[0035] Acquire the user's voice signal collected by the sensor;
[0036] Obtaining a brain wave signal of the user corresponding to when the user utters the voice; and
[0037] Target voice information is obtained according to the brain wave signal and the user voice signal collected by the sensor.
[0038] In an optional implementation, the method further includes:
[0039] Acquiring, based on the brainwave signal, a motion signal of the vocal tract occlusion portion of the user when speaking; correspondingly, acquiring target voice information based on the brainwave signal and the user voice signal collected by the sensor, includes:
[0040] The target voice information is obtained according to the motion signal and the user voice signal collected by the sensor.
[0041] In an optional implementation, obtaining target voice information according to the brain wave signal and the user voice signal collected by the sensor includes:
[0042] Acquiring a corresponding target audio signal according to the brainwave signal;
[0043] Based on filtering, filtering out the target audio signal from the user voice signal collected by the sensor to obtain a signal to be filtered out;
[0044] The signal to be filtered is filtered out from the user voice signal collected by the sensor to obtain the target voice information.
[0045] In an optional implementation, the method further includes:
[0046] Based on the target voice information, instruction information corresponding to the user's voice signal is acquired, where the instruction information indicates a semantic intention contained in the user's voice signal.
[0047] In an optional implementation, obtaining target voice information according to the brain wave signal and the user voice signal collected by the sensor includes:
[0048] Based on the brain wave signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model; or
[0049] According to the brain wave signal, a corresponding target audio signal is obtained; based on the target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
[0050] In an optional implementation, the target voice information includes a voiceprint feature representing a voice signal of the user.
[0051] In a third aspect, the present application provides a method for processing a speech signal, the method comprising:
[0052] Acquire the user's voice signal collected by the sensor;
[0053] Obtaining a vibration signal corresponding to when the user utters the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and
[0054] Voiceprint recognition is performed based on the user voice signal and the vibration signal collected by the sensor.
[0055] In an optional implementation, the vibration signal is used to represent a vibration feature corresponding to the vibration generated by uttering speech.
[0056] In an optional implementation, performing voiceprint recognition based on the user voice signal and the vibration signal collected by the sensor includes:
[0057] performing voiceprint recognition based on the user voice signal collected by the sensor to obtain a first confidence level that the user voice signal collected by the sensor belongs to the user;
[0058] performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the target user;
[0059] A voiceprint recognition result is obtained according to the first confidence level and the second confidence level.
[0060] In an optional implementation, the method further includes:
[0061] Obtaining a brain wave signal of the user corresponding to when the user utters the voice;
[0062] Accordingly, performing voiceprint recognition based on the user voice signal and the vibration signal collected by the sensor includes:
[0063] Voiceprint recognition is performed based on the user voice signal, the vibration signal, and the brainwave signal collected by the sensor.
[0064] In an optional implementation, performing voiceprint recognition based on the user voice signal, the vibration signal, and the brainwave signal collected by the sensor includes:
[0065] performing voiceprint recognition based on the user voice signal collected by the sensor to obtain a first confidence level that the user voice signal collected by the sensor belongs to the user;
[0066] performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the user;
[0067] performing voiceprint recognition based on the brainwave signal to obtain a third confidence level that the user voice signal collected by the sensor belongs to the user;
[0068] A voiceprint recognition result is obtained according to the first confidence level, the second confidence level, and the third confidence level.
[0069] In a fourth aspect, the present application provides a speech signal processing device, the device comprising:
[0070] Environmental voice acquisition module, used to obtain user voice signals collected by sensors;
[0071] a vibration signal acquisition module, configured to acquire a vibration signal corresponding to the user uttering the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part being a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and
[0072] The voice information acquisition module is used to obtain target voice information according to the vibration signal and the user voice signal collected by the sensor.
[0073] In an optional implementation, the vibration signal is used to represent a vibration feature corresponding to the vibration generated by the user uttering the voice.
[0074] In an optional implementation, the body part includes at least one of the following: the top of the skull, the face, the throat, or the neck.
[0075] In an optional implementation, the vibration signal acquisition module is configured to acquire a video frame including the user; and extract, based on the video frame, a vibration signal corresponding to when the user utters a voice.
[0076] In an optional implementation, the video frames are acquired through a dynamic vision sensor and / or a high-speed camera.
[0077] In an optional implementation, the voice information acquisition module is used to obtain a corresponding target audio signal based on the vibration signal; based on filtering, filter out the target audio signal from the user voice signal collected by the sensor to obtain a noise-removed signal; and filter out the noise signal to be filtered from the user voice signal collected by the sensor to obtain the target voice information.
[0078] In an optional implementation, the apparatus further includes:
[0079] The instruction information acquisition module is used to acquire instruction information corresponding to the user's voice signal based on the target voice information, where the instruction information indicates the semantic intention contained in the user's voice signal.
[0080] In an optional implementation, the voice information acquisition module is used to obtain the target voice information through a recurrent neural network model based on the vibration signal and the user voice signal collected by the sensor; or, based on the vibration signal, obtain the corresponding target audio signal; based on the target audio signal and the user voice signal collected by the sensor, obtain the target voice information through a recurrent neural network model.
[0081] In an optional implementation, the apparatus further includes:
[0082] The brain wave signal acquisition module is used to obtain the brain wave signal of the user corresponding to when the user utters the voice; correspondingly, the voice information acquisition module is used to obtain target voice information based on the vibration signal, the brain wave signal and the user voice signal collected by the sensor.
[0083] In an optional implementation, the apparatus further includes:
[0084] The motion signal acquisition module is used to obtain the motion signal of the vocal tract occlusion part when the user speaks according to the brain wave signal; correspondingly, the voice information acquisition module is used to obtain target voice information according to the vibration signal, the motion signal and the user voice signal collected by the sensor.
[0085] In an optional implementation, the voice information acquisition module is configured to obtain the target voice information through a recurrent neural network model based on the vibration signal, the brain wave signal, and the user voice signal collected by the sensor; or
[0086] Acquire a corresponding first target audio signal according to the vibration signal;
[0087] According to the brain wave signal, a corresponding second target audio signal is obtained; based on the first target audio signal, the second target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
[0088] In an optional implementation, the target voice information includes a voiceprint feature representing a voice signal of the user.
[0089] In a fifth aspect, the present application provides a speech signal processing device, the device comprising:
[0090] An environmental voice acquisition module is used to obtain the user's voice signal collected by the sensor;
[0091] an electroencephalogram signal acquisition module, configured to acquire an electroencephalogram signal of the user corresponding to when the user utters the speech; and
[0092] The voice information acquisition module is used to obtain target voice information based on the brain wave signal and the user voice signal collected by the sensor.
[0093] In an optional implementation, the apparatus further includes:
[0094] The motion signal acquisition module is used to obtain the motion signal of the vocal tract occlusion part of the user when speaking based on the brain wave signal; correspondingly, the voice information acquisition module is used to obtain the target voice information based on the motion signal and the user voice signal collected by the sensor.
[0095] In an optional implementation, the voice information acquisition module is used to acquire a corresponding target audio signal based on the brain wave signal;
[0096] Based on filtering, filtering out the target audio signal from the user voice signal collected by the sensor to obtain a signal to be filtered out;
[0097] The signal to be filtered is filtered out from the user voice signal collected by the sensor to obtain the target voice information.
[0098] In an optional implementation, the apparatus further includes:
[0099] The instruction information acquisition module is used to acquire instruction information corresponding to the user's voice signal based on the target voice information, where the instruction information indicates the semantic intention contained in the user's voice signal.
[0100] In an optional implementation, the voice information acquisition module is configured to obtain the target voice information through a recurrent neural network model based on the brain wave signal and the user voice signal collected by the sensor; or
[0101] According to the brain wave signal, a corresponding target audio signal is obtained; based on the target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
[0102] In an optional implementation, the target voice information includes a voiceprint feature representing a voice signal of the user.
[0103] In a sixth aspect, the present application provides a speech signal processing device, the device comprising:
[0104] Environmental voice acquisition module, used to obtain user voice signals collected by sensors;
[0105] a vibration signal acquisition module, configured to acquire a vibration signal corresponding to the user uttering the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part being a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and
[0106] The voiceprint recognition module is used to perform voiceprint recognition based on the user voice signal and the vibration signal collected by the sensor.
[0107] In an optional implementation, the vibration signal is used to represent a vibration feature corresponding to the vibration generated by uttering speech.
[0108] In an optional implementation, the voiceprint recognition module is configured to perform voiceprint recognition based on the user voice signal collected by the sensor to obtain a first confidence level that the user voice signal collected by the sensor belongs to the user;
[0109] performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the target user;
[0110] A voiceprint recognition result is obtained according to the first confidence level and the second confidence level.
[0111] In an optional implementation, the apparatus further includes:
[0112] an electroencephalogram signal acquisition module, configured to acquire an electroencephalogram signal of the user corresponding to the user uttering the speech;
[0113] Correspondingly, the voiceprint recognition module is used to perform voiceprint recognition based on the user voice signal, the vibration signal and the brainwave signal collected by the sensor.
[0114] In an optional implementation, the voiceprint recognition module is configured to perform voiceprint recognition based on the user voice signal collected by the sensor to obtain a first confidence level that the user voice signal collected by the sensor belongs to the user;
[0115] performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the user;
[0116] performing voiceprint recognition based on the brainwave signal to obtain a third confidence level that the user voice signal collected by the sensor belongs to the user;
[0117] A voiceprint recognition result is obtained according to the first confidence level, the second confidence level, and the third confidence level.
[0118] In a seventh aspect, the present application provides an autonomous driving vehicle, which may include a processor coupled to a memory, the memory storing program instructions, and which, when executed by the processor, implements the method described in the first aspect. For details regarding the steps performed by the autonomous driving vehicle in each possible implementation of the first aspect performed by the processor, please refer to the first aspect and will not be repeated here.
[0119] In an eighth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the method described in the first aspect above.
[0120] In a ninth aspect, the present application provides a circuit system, which includes a processing circuit, and the processing circuit is configured to execute the method described in the first aspect above.
[0121] In a tenth aspect, the present application provides a computer program which, when executed on a computer, enables the computer to execute the method described in the first aspect above.
[0122] In an eleventh aspect, the present application provides a chip system, which includes a processor for supporting a server or a threshold value acquisition device to implement the functions involved in the above aspects, for example, sending or processing the data and / or information involved in the above methods. In one possible design, the chip system also includes a memory, which is used to store program instructions and data necessary for the server or communication device. The chip system can be composed of a chip or can include a chip and other discrete devices.
[0123] The embodiment of the present application provides a method for processing a speech signal, comprising: obtaining a speech signal of a user collected by a sensor, wherein the speech signal includes environmental noise; obtaining a vibration signal corresponding to when the user utters the speech; wherein the vibration signal is used to represent the vibration characteristics of a body part of the user; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and obtaining target speech information based on the vibration signal and the user speech signal collected by the sensor. In the above manner, the vibration signal is used as the basis for speech recognition. Since the vibration signal does not contain external non-user speech mixed in during complex acoustic transmission, it is less affected by other environmental noises (such as reverberation). Therefore, this part of noise interference can be relatively well suppressed, and a better speech recognition effect can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0124] Figure 1a It is a schematic diagram of a smart device;
[0125] Figure 1bA graphical user interface diagram of a mobile phone provided in an embodiment of the present application;
[0126] Figure 2 This is an illustration of an application scenario of an embodiment of the present invention;
[0127] Figure 3 and Figure 4 Another application scenario provided by the embodiment of the present application is illustrated;
[0128] Figure 5 A schematic diagram of the structure of an electronic device;
[0129] Figure 6 This is a schematic diagram of the software structure of the electronic device according to an embodiment of the present application;
[0130] Figure 7 This is a flow chart of a method for processing speech signals provided in an embodiment of the present application;
[0131] Figure 8 A schematic diagram of a system architecture;
[0132] Figure 9 A schematic diagram of the structure of an RNN;
[0133] Figure 10 A schematic diagram of the structure of an RNN;
[0134] Figure 11 A schematic diagram of the structure of an RNN;
[0135] Figure 12 A schematic diagram of the structure of an RNN;
[0136] Figure 13 A schematic diagram of the structure of an RNN;
[0137] Figure 14 A flow chart of a method for processing a speech signal provided in an embodiment of the present application;
[0138] Figure 15 A flow chart of a method for processing a speech signal provided in an embodiment of the present application;
[0139] Figure 16 This application provides a structural diagram of a speech signal processing device;
[0140] Figure 17 This application provides a structural diagram of a speech signal processing device;
[0141] Figure 18 This application provides a structural diagram of a speech signal processing device;
[0142] Figure 19A schematic diagram of the structure of an execution device provided in an embodiment of the present application;
[0143] Figure 20 This is a structural diagram of a training device provided in an embodiment of the present application;
[0144] Figure 21 A schematic diagram of the structure of the chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0145] The embodiments of the present invention will be described below with reference to the accompanying drawings.
[0146] The terms "first," "second," "third," and "fourth," etc., in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, rather than to describe a specific order. In addition, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0147] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0148] As used in this specification, the terms "component," "module," "system," and the like are used to represent computer-related entities, hardware, firmware, a combination of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. By way of illustration, both an application running on a computing device and a computing device can be a component. One or more components can reside in a process and / or an execution thread, and a component can be located on a computer and / or distributed between two or more computers. In addition, these components can be executed from various computer-readable media having various data structures stored thereon. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component on a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0149] The technical solution in this application will be described below with reference to the accompanying drawings.
[0150] The speech signal processing method provided in the embodiments of the present application can be applied to scenarios such as human-computer interaction related to speech recognition and voiceprint recognition. Specifically, the speech signal processing method provided in the embodiments of the present application can be applied to speech recognition and voiceprint recognition. The following briefly introduces the speech recognition scenario and the voiceprint recognition scenario respectively.
[0151] Scenario 1: Human-computer interaction based on speech recognition:
[0152] Speech recognition (ASR), also known as automatic speech recognition, aims to convert the vocabulary content in human speech into computer-readable input, such as keystrokes, binary codes, or character sequences.
[0153] In one scenario, the present application can be applied to a device with a voice interaction function. In this embodiment, "having a voice interaction function" can be a function that can be implemented on the device, which can recognize the user's voice and trigger corresponding functions based on the voice, thereby realizing voice interaction with the user. The device with a voice interaction function can be, for example, a smart device such as a speaker, an alarm clock, a watch, a robot, or an in-vehicle device, or a portable device such as a mobile phone, a tablet, an AR augmented reality device, or a VR virtual reality device.
[0154] In one embodiment of the application, a device with a voice interaction function may include an audio sensor and a video sensor, wherein the audio sensor can collect audio signals in the environment, and the video sensor can collect videos in a certain area; the audio signal may include audio signals emitted by one or more users when speaking and other noise signals in the environment, and the video may include the one or more users speaking. Furthermore, the audio signal emitted by one or more users when speaking can be extracted based on the audio signal and the video. Furthermore, a voice interaction function with the user can be implemented based on the extracted audio signal. How to extract the audio signal emitted by one or more users when speaking based on the audio signal and the video will be described in detail in subsequent embodiments and will not be repeated here.
[0155] In one embodiment of the present application, the audio sensor and video sensor may not be components of the device with voice interaction functionality itself, but may be independent components or components integrated into other devices. In such a case, the device with voice interaction functionality may only obtain the audio signals in the environment captured by the audio sensor or only obtain the video within a certain area captured by the video sensor. Then, based on the audio signals and video, the device may extract the audio signals emitted by one or more users when speaking. Furthermore, based on the extracted audio signals, the device may implement voice interaction with the user.
[0156] Furthermore, the audio sensor is a component of the device with voice interaction function, while the video sensor is not a component of the device with voice interaction function; or, the audio sensor is not a component of the device with voice interaction function, and the video sensor is not a component of the device with voice interaction function; or, the audio sensor is not a component of the device with voice interaction function, while the video sensor is a component of the device with voice interaction function.
[0157] For example, the device with voice interaction function may be Figure 1a Smart devices such as Figure 1a As shown, when the smart device recognizes the voice "piupiupiu," it does not perform any action. For example, if the smart device recognizes the voice "turn on the air conditioner," it performs the action associated with the voice "turn on the air conditioner": turning on the air conditioner. For example, if the smart device recognizes the sound of a user whistling, i.e., a whistle, it performs the action associated with the whistle: turning on the light. For example, if the smart device recognizes the voice "turn on the light," it does not perform any action. For example, if the smart device recognizes the voice "sleep" in whisper mode, it performs the action associated with the voice "sleep" in whisper mode: switching to sleep mode. The voice "piupiupiu," the whistle, and the voice "sleep" in whisper mode are special voices. The voices "turn on the air conditioner," "turn on the light," and so on are normal voices. Normal voices are those that can be semantically recognized and that vibrate the vocal cords when pronounced. Special voices are those that differ from normal voices. For example, special voices are those that do not vibrate the vocal cords when pronounced, i.e., voiceless voices. Another example is voices that lack semantic meaning.
[0158] For example, the device with voice interaction function can be a device with display function, such as a mobile phone, see Figure 1b , Figure 1b A graphical user interface (GUI) of a mobile phone provided in an embodiment of the present application is shown as follows: Figure 1bAs shown in the figure, the GUI is the display interface when the mobile phone interacts with the user. When the mobile phone detects the user's voice wake-up word "Xiao Yi Xiao Yi", the mobile phone can display the text display window 101 of the voice assistant on the desktop, and the mobile phone can remind the user "Hi, I'm listening" through window 101. It should be understood that the mobile phone can also broadcast "Hi, I'm listening" to the user in voice while displaying the text reminder through window 101 or 102.
[0159] In some scenarios, the device with voice interaction function may be a system composed of multiple devices.
[0160] Figure 2 The figure shows an application scenario of an embodiment of the present application. Figure 2 The application scenario in can also be called a smart home scenario. Figure 2 The application scenario in may include at least one electronic device (eg, electronic device 210, electronic device 220, electronic device 230, electronic device 240, electronic device 250), electronic device 260, and electronic device 270. Figure 2 In the example, electronic device 210 may be a television. Electronic device 220 may be a speaker. Electronic device 230 may be a monitoring device. Electronic device 240 may be a watch. Electronic device 250 may be a smart microphone. Electronic device 260 may be a mobile phone or tablet. Electronic devices may also be wireless communication devices, such as routers and gateways. Figure 2 The electronic devices 210, 220, 230, 240, 250, and 260 can perform uplink and downlink transmissions with the electronic devices via wireless communication protocols. For example, the electronic devices can send information to the electronic devices 210, 220, 230, 240, 250, and 260, and can also receive information sent by the electronic devices 210, 220, 230, 240, 250, and 260.
[0161] It should be noted that the embodiments of the present application can be applied to application scenarios including one or more wireless communication devices and multiple electronic devices, and the present application does not limit this.
[0162] In the embodiment of the present application, the device with voice interaction function can be any electronic device in the smart home system, such as a TV, a speaker, a watch, a smart microphone, a mobile phone or a tablet computer, etc. Any electronic device in the smart home system can include an audio sensor or a video sensor. After acquiring audio information or video in the environment, the audio information or video can be transmitted to the device with voice interaction function or to a server on the cloud side based on a wireless communication device. Figure 2(not shown), the device with voice interaction function can extract the audio signals emitted by one or more users when speaking based on the audio information and video. Then, the voice interaction function with the user can be realized based on the extracted audio signals; or the server on the cloud side can extract the audio signals emitted by one or more users when speaking based on the audio information and video, and transmit the extracted audio signals to the device with voice interaction function, and then the device with voice interaction function can realize the voice interaction function with the user based on the extracted audio signals.
[0163] In one example, an application scenario includes electronic device 210, electronic device 260, and electronic device 210. Electronic device 260 is a television, and electronic device 210 is a mobile phone. The router is used to enable wireless communication between the television and the mobile phone. The device with voice interaction functionality can be a mobile phone. The television can be equipped with a video sensor, and the mobile phone can be equipped with an audio sensor. After capturing video, the television can transmit the video to the mobile phone. The mobile phone can extract audio signals emitted by one or more users based on the audio information and video. Furthermore, based on the extracted audio signals, voice interaction with the user can be implemented.
[0164] In one example, an application scenario includes electronic device 220, electronic device 260, and electronic device 220. Electronic device 260 is a speaker, electronic device 220 is a mobile phone, and electronic device 260 is a router. The router is used to enable wireless communication between the speaker and the mobile phone. The device with voice interaction functionality can be a mobile phone, which can be equipped with a video sensor, and the speaker can be equipped with an audio sensor. After the speaker acquires audio information, it can transmit the audio information to the mobile phone. The mobile phone can extract audio signals emitted by one or more users based on the audio information and video. Furthermore, based on the extracted audio signals, voice interaction with the user can be implemented.
[0165] In one example, the application scenario includes electronic device 230, electronic device 260, and electronic device 230. Electronic device 230 is a monitoring device, electronic device 260 is a mobile phone, and electronic device 260 is a router. The router is used to enable wireless communication between the monitoring device and the mobile phone. The device with voice interaction functionality can be a mobile phone. The monitoring device can be equipped with a video sensor, and the mobile phone can be equipped with an audio sensor. After the monitoring device captures the video, it can transmit the video to the mobile phone. The mobile phone can extract audio signals emitted by one or more users when speaking based on the audio information and video. Furthermore, based on the extracted audio signals, voice interaction functionality with the user can be implemented.
[0166] In one example, the application scenario includes electronic device 250, electronic device 260, and electronic device 250. Electronic device 250 is a microphone, electronic device 260 is a mobile phone, and electronic device 260 is a router. The router is used to implement wireless communication between the microphone and the mobile phone. The device with voice interaction function can be a microphone, the mobile phone can be provided with a video sensor, and the microphone can be provided with an audio sensor. After the mobile phone captures the video, it can transmit the video to the microphone. The microphone can extract the audio signal emitted by one or more users when speaking based on the audio information and video. Furthermore, the voice interaction function with the user can be implemented based on the extracted audio signal.
[0167] It should be noted that the above description of the product form is only for illustration. In actual applications, the deployment form of the video sensor and audio sensor can be flexibly set.
[0168] Figure 3 、 Figure 4 This is another application scenario provided by the embodiments of the present application. Figure 3 、 Figure 4 The application scenario in can also be called an intelligent driving scenario. Figure 3 、 Figure 4 The application scenario in may include electronic devices, which include device 310, device 320, device 330, device 340, and device 350. The electronic device may be a driving system (also referred to as an in-vehicle system). Device 310 may be a display screen. Device 320 may be a microphone. Device 330 may be a speaker. Device 340 may be a camera. Device 350 may be a seat adjustment device. Electronic device 360 may be a mobile phone or a tablet computer. The electronic device may receive data sent by device 310, device 320, device 330, device 340, and device 350. Moreover, the electronic device and the electronic device 360 may communicate through a wireless communication protocol. For example, the electronic device may send a signal to the electronic device 360, and may also receive a signal sent by the electronic device 360.
[0169] It should be noted that the embodiments of the present application can be applied to application scenarios including driving systems and multiple electronic devices, and this application does not limit this.
[0170] In one example, the application scenario includes device 320, device 330, electronic device 360 and electronic device (driving system). Device 320 is a microphone, device 340 is a camera, electronic device 360 is a tablet computer, and the electronic device is a driving system. Among them, the driving system is used to wirelessly communicate with the mobile phone, and is also used to drive the microphone to collect audio signals, and drive the camera to collect video. The driving system can drive the microphone to collect audio signals, and send the audio signals collected by the microphone to the tablet computer. The driving system can drive the camera to collect video, and send the video collected by the camera to the tablet computer; the tablet computer can extract the audio signals emitted when one or more users speak based on the audio information and video. Furthermore, the voice interaction function with the user can be implemented based on the extracted audio signal.
[0171] It should be noted that in intelligent driving scenarios, video sensors can be deployed independently, for example, set at a preset position in the car, so that they can capture videos within a preset area. For example, video sensors can be set on the windshield or on the seat, and can then capture videos of users in a certain seat.
[0172] In one embodiment of the present application, a device with a speech recognition function may be a head-mounted portable device, such as an AR / VR device. The head-mounted portable device may be provided with an audio sensor and a brainwave acquisition device. The audio sensor may acquire audio signals, and the brainwave acquisition device may acquire brainwave signals. The head-mounted portable device may then extract audio signals emitted by one or more users when speaking based on the audio signals and brainwave signals. Furthermore, a speech interaction function with the user may be implemented based on the extracted audio signals.
[0173] It should be noted that the aforementioned audio sensor and brainwave acquisition device may not be components of the device with voice interaction functionality itself, but rather may be independent components or components integrated into other devices. In such cases, the device with voice interaction functionality may only acquire the audio signals in the environment captured by the audio sensor or only acquire the brainwave signals within a certain area captured by the brainwave acquisition device. Furthermore, based on the audio and brainwave signals, the device may extract the audio signals emitted by one or more users when speaking. Furthermore, based on the extracted audio signals, voice interaction with the user may be implemented.
[0174] Furthermore, the audio sensor is a component of the device with voice interaction function, while the brainwave acquisition device is not a component of the device with voice interaction function; or, the audio sensor is not a component of the device with voice interaction function, and the brainwave acquisition device is not a component of the device with voice interaction function; or, the audio sensor is not a component of the device with voice interaction function, while the brainwave acquisition device is a component of the device with voice interaction function.
[0175] It is understandable that Figures 1a to 4 The application scenarios described above are only several exemplary implementations in the embodiments of the present invention. The application scenarios in the embodiments of the present invention include but are not limited to the above application scenarios.
[0176] 2. Voiceprint recognition scenarios:
[0177] A voiceprint is a sound wave spectrum that carries speech information, as detected by electroacoustic instruments. It is a biometric characteristic composed of over a hundred characteristic dimensions, including wavelength, frequency, and intensity. Voiceprint recognition analyzes the characteristics of one or more speech signals to identify unknown voices. Simply put, it's a technology that determines whether a particular sentence was spoken by a specific person. Voiceprints can be used to determine the speaker's identity, allowing for targeted responses.
[0178] In addition, the present application can also be applied to the scenario of speech denoising. The speech signal processing method in the present application can be used in audio input devices that require speech denoising, such as headphones, microphones (independent microphones or microphones on terminal devices, etc.). Users can speak into the audio input. Through the speech signal processing method in the present application, the audio input device can extract the speech signal emitted by the user from the audio input including ambient noise.
[0179] It should be understood that the examples given here are only for the purpose of facilitating the understanding of the application scenarios of the embodiments of the present application, and are not intended to be exhaustive. The embodiments of the present application are described below in conjunction with the accompanying drawings. It is understood by those skilled in the art that with the development of technology and the emergence of new scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0180] The following describes an electronic device provided by an embodiment of the present application, a user interface for such an electronic device, and an embodiment for using such an electronic device. In some embodiments, the electronic device may be a portable electronic device that also includes other functions such as a personal digital assistant and / or a music player function, such as a mobile phone, a tablet computer, a wearable electronic device with a wireless communication function (such as a smart watch), etc. Exemplary embodiments of portable electronic devices include but are not limited to portable electronic devices equipped with or other operating systems. The above-mentioned portable electronic device may also be other portable electronic devices, such as a laptop computer (Laptop), etc. It should also be understood that in some other embodiments, the above-mentioned electronic device may not be a portable electronic device, but may be a desktop computer, a television, a speaker, a monitoring device, a camera, a display, a microphone, a seat adjustment device, a fingerprint recognition device, an in-vehicle driving system, etc.
[0181] For example, Figure 5 1 shows a schematic structural diagram of an electronic device 100. The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a microphone 170C, a sensor module 180, a button 190, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface.
[0182] It should be understood that the structures illustrated in the embodiments of the present application do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0183] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). The different processing units may be independent components or integrated into one or more processors. In some embodiments, the electronic device 100 may also include one or more processors 110. The controller may generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution. In other embodiments, the processor 110 may also include a memory for storing instructions and data. Exemplarily, the memory in the processor 110 may be a cache memory. This memory may store instructions or data that have just been used or are being recycled by the processor 110. If the processor 110 needs to use the instruction or data again, it can directly call it from the memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the electronic device 100 in processing data or executing instructions.
[0184] In some embodiments, the processor 110 may include one or more interfaces. The interface may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. Among them, the USB interface is an interface that complies with the USB standard specification, and specifically can be a MiniUSB interface, a MicroUSB interface, a USB Type C interface, etc. The USB interface can be used to connect a charger to charge the electronic device 100, and can also be used to transmit data between the electronic device 100 and peripheral devices. The USB interface can also be used to connect headphones to play audio through the headphones.
[0185] It is understood that the interface connection relationship between the modules illustrated in the embodiments of the present application is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.
[0186] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. A GPU is a microprocessor for image processing that connects display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs that execute program instructions to generate or modify display information.
[0187] Display screen 194 is used to display images, videos, and the like. Display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED). In some embodiments, electronic device 100 may include one or more display screens 194.
[0188] The display screen 194 of the electronic device 100 can be a flexible screen. Currently, flexible screens have attracted much attention due to their unique characteristics and huge potential. Compared with traditional screens, flexible screens have the characteristics of strong flexibility and bendability, which can provide users with new interaction methods based on bendability and meet users' more demands for electronic devices. For electronic devices equipped with foldable displays, the foldable displays on the electronic devices can be switched between a small screen in a folded form and a large screen in an unfolded form at any time. Therefore, users are using the split-screen function on electronic devices equipped with foldable displays more and more frequently.
[0189] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.
[0190] The ISP processes data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. It can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 193.
[0191] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device 100 may include one or more cameras 193.
[0192] The camera 193 in the embodiment of the present application may be a high-speed camera or a dynamic vision sensor (DVS).
[0193] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.
[0194] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.
[0195] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU can enable intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.
[0196] The external memory interface 120 can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 via the external memory interface 120 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.
[0197] The internal memory 121 can be used to store one or more computer programs, which include instructions. The processor 110 can execute the above instructions stored in the internal memory 121, so that the electronic device 100 performs the method for displaying the screen off provided in some embodiments of the present application, as well as various applications and data processing. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store an operating system; the program storage area may also store one or more applications (such as a gallery, contacts, etc.). The data storage area may store data created during the use of the electronic device 100 (such as photos, contacts, etc.). In addition, the internal memory 121 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more disk storage components, a flash memory component, a universal flash storage (UFS), etc. In some embodiments, the processor 110 can execute the instructions stored in the internal memory 121, and / or the instructions stored in the memory provided in the processor 110, so that the electronic device 100 performs the method for displaying the screen off provided in the embodiments of the present application, as well as other applications and data processing. The electronic device 100 can implement audio functions such as music playback and recording through an audio module, a speaker, a receiver, a microphone, an earphone interface, and an application processor.
[0198] The sensor module 180 may include an acceleration sensor 180E, a fingerprint sensor 180H, an ambient light sensor 180L, and the like.
[0199] Accelerometer 180E can detect the magnitude of acceleration of electronic device 100 in all directions (generally three axes). It can also detect the magnitude and direction of gravity when electronic device 100 is stationary. It can also be used to identify the electronic device's posture, enabling applications such as switching between landscape and portrait modes and pedometers.
[0200] Ambient light sensor 180L is used to sense ambient light brightness. Electronic device 100 can adaptively adjust the brightness of display screen 194 based on the perceived ambient light. Ambient light sensor 180L can also be used to automatically adjust white balance when taking photos. Ambient light sensor 180L can also work with proximity light sensor 180G to detect whether electronic device 100 is in a pocket to prevent accidental touches.
[0201] The fingerprint sensor 180H is used to collect fingerprints. The electronic device 100 can use the collected fingerprint characteristics to implement fingerprint unlocking, access application locks, fingerprint photography, fingerprint call answering, etc.
[0202] The brain wave sensor 195 can collect brain wave signals.
[0203] The buttons 190 include a power button, a volume button, and the like. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.
[0204] Figure 6 This is a block diagram of the software structure of the electronic device 100 according to an embodiment of the present application. The layered architecture divides the software into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers: the application layer, the application framework layer, the Android runtime and system libraries, and the kernel layer. The application layer may include a series of application packages.
[0205] like Figure 6 As shown, the application package may include applications such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc.
[0206] The application framework layer provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions.
[0207] like Figure 6 As shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, and the like.
[0208] The window manager is used to manage window programs. The window manager can obtain the display size, determine whether there is a status bar, lock the screen, take screenshots, etc.
[0209] Content providers are used to store and retrieve data and make it accessible to applications. The data may include videos, images, audio, calls made and received, browsing history and bookmarks, phone books, etc.
[0210] The view system includes visual controls, such as those for displaying text and images. The view system is used to build applications. A display interface can consist of one or more views. For example, a display interface containing a text notification icon might include a view for displaying text and a view for displaying images.
[0211] The phone manager is used to provide communication functions of the electronic device 100, such as management of call status (including answering, hanging up, etc.).
[0212] The resource manager provides various resources for applications, such as localized strings, icons, images, layout files, video files, and so on.
[0213] The Notification Manager allows applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically without user interaction. For example, the Notification Manager is used to notify users of completed downloads and message reminders. The Notification Manager can also display notifications in the top status bar of the system as icons or scrolling text, such as notifications from background applications, or as dialog windows on the screen. Examples include text messages in the status bar, beeps, vibrations on electronic devices, and flashing indicator lights.
[0214] The system library can include multiple functional modules, such as a surface manager, media libraries, a 3D graphics processing library (such as OpenGL ES), and a 2D graphics engine (such as SGL).
[0215] The surface manager is used to manage the display subsystem and provide fusion of 2D and 3D layers for multiple applications.
[0216] The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library can support a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0217] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.
[0218] A 2D graphics engine is a drawing engine for 2D drawings.
[0219] The kernel layer is the layer between hardware and software. The kernel layer includes at least display driver, camera driver, audio driver, and sensor driver.
[0220] For ease of understanding, the following examples of this application will be described with Figure 5 and Figure 6 Taking the device of the structure shown as an example, the element pressing method provided in the embodiment of the present application is specifically described in combination with the accompanying drawings.
[0221] Reference Figure 7 , Figure 7 A flow chart of a speech signal processing method provided in an embodiment of the present application is shown as follows: Figure 7 As shown in the figure, the speech signal processing method provided in the embodiment of the present application includes:
[0222] 701. Obtain a user voice signal collected by a sensor.
[0223] In an embodiment of the present application, a user's voice signal collected by a sensor from the environment can be obtained, and the voice signal includes environmental noise; the voice signal hereinafter can also be expressed as a voice signal.
[0224] It should be noted that the user voice signal should not be understood as merely the words spoken by the user, but should be understood as the voice signal including the voice uttered by the user.
[0225] It should be noted that when a voice signal includes ambient noise, it can be understood that the presence of the user speaking and other ambient noise (such as other people speaking) in the environment. In this case, the collected voice signal includes the intertwined user voice and ambient noise. The relationship between the voice signal and ambient noise should not be understood as a simple superposition. In other words, it should not be understood that ambient noise exists as an independent signal in the voice signal.
[0226] In an embodiment of the present application, an audio sensor (such as a microphone or a microphone array) can collect user voice signals from the environment. The user voice signal is a mixed signal z(n) in the environment. In addition to the voice signal s1(n) emitted by the user that is desired to be picked up, there are other signals, such as environmental noise n(n), other people's voices s2(n), etc., that is, z(n) = s1(n) + s2(n) + n(n). In scenarios where voice interaction and voice denoising are required, we hope to extract the voice signal emitted by the user from the voice signal in the environment collected by the audio sensor, that is, to separate the voice signal s1(n) emitted by the user from the mixed signal z(n).
[0227] It should be noted that the execution subject of step 701 can be a device with voice interaction function or a voice input device; taking the execution subject as an example of a device with voice interaction function, in one implementation, the device with voice interaction function can be integrated with an audio sensor, and then the audio sensor can obtain an audio signal including the user's voice signal; in one implementation, the audio sensor may not be integrated into the device with voice interaction function, for example, the audio sensor can be integrated into other devices, or as an independent device (such as an independent microphone), the audio sensor can transmit the collected audio signal to the device with voice interaction function, then the device with voice interaction function can obtain the audio signal.
[0228] Optionally, the audio sensor can specifically pick up audio signals coming from a certain direction, such as directional pickup in the direction of the user, thereby eliminating some external noise as much as possible (but there is still noise). Directional acquisition requires a microphone array or a vector microphone. Here, taking the microphone array as an example, a beamforming method can be used. A beamformer can be used to implement this, which can include delay-sum beamforming and filter-sum beamforming. Specifically, assume that the input signal of the microphone array is z i (n), the filter transfer coefficient is w i (n), then the output of the filter-sum beamformer system is:
[0229]
[0230] Where M is the number of microphones. When the filter coefficient is just a single weighting constant, filter-sum beamforming is simplified to delay-sum beamforming, that is:
[0231]
[0232] Among them, τ i Denotes the delay compensation obtained by estimation. By controlling τ i The value of can point the array beam in any direction to pick up the audio signal in that direction. If you do not want to pick up the audio signal in a certain direction, you can control the beam pointing to not include that direction. The audio signal collected after the pickup direction control is z(n).
[0233] In addition, the product form of the voice input device can refer to the product form of the above-mentioned device with voice interaction function, which will not be repeated here.
[0234] 702. Obtain a vibration signal corresponding to when the user utters the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state.
[0235] In an embodiment of the present application, a vibration signal corresponding to when a user utters a voice may be obtained, wherein the vibration signal is used to represent the vibration characteristics of a body part of the user when uttering the voice signal.
[0236] It should be noted that there is no strict timing limit between step 701 and step 702. Step 701 can be executed before or after step 702, or at the same time, and this application does not limit it.
[0237] In the embodiment of the present application, the vibration signal corresponding to the user's voice can be obtained based on video extraction;
[0238] The action of extracting the vibration signal from the video frame may be performed by a device having a voice interaction function or a voice input device; taking the case where the action of extracting the vibration signal from the video frame may be performed by a device having a voice interaction function as an example:
[0239] In one implementation, a device with a voice interaction function may be integrated with a video sensor that can capture video frames including the user. Accordingly, the device with a voice interaction function may extract a vibration signal corresponding to the user based on the video frames.
[0240] In one implementation, a video sensor can be set up independently from a device with a voice interaction function. The video sensor can collect video frames including the user and send the video frames to the device with a voice interaction function. Accordingly, the device with a voice interaction function can extract the vibration signal corresponding to the user based on the video frame; In one implementation, a video sensor can be set up independently from a device with a voice interaction function. The video sensor can collect video frames including the user and send the video frames to the device with a voice interaction function. Accordingly, the device with a voice interaction function can extract the vibration signal corresponding to the user based on the video frame.
[0241] The action of extracting the vibration signal from the video frame may be performed by a server on the cloud side or other devices on the terminal side;
[0242] In one implementation, a video sensor can be integrated into a device with a voice interaction function. The video sensor can collect video frames including the user and send the video frames to a cloud-side server or other end-side devices. Correspondingly, the cloud-side server or other end-side devices can extract the vibration signal corresponding to the user based on the video frame and send the vibration signal to the device with a voice interaction function.
[0243] In one implementation, a video sensor can be set independently from a device with a voice interaction function. The video sensor can collect video frames including the user and send the video frames to a cloud-side server or other end-side devices. Correspondingly, the cloud-side server or other end-side devices can extract the vibration signal corresponding to the user based on the video frame and send the vibration signal to the device with a voice interaction function.
[0244] It should be noted that the above descriptions of the execution entities of the actions of extracting vibration signals from video frames are only some exemplary examples, and this application does not limit them.
[0245] In one implementation, the video frames are captured by a dynamic visual sensor and / or a high-speed camera. Taking the case of video frames captured by a dynamic visual sensor as an example, in this embodiment of the present application, the dynamic visual sensor can capture video frames including the top of the skull, face, throat, or neck of the user while they are speaking.
[0246] In one implementation, the number of dynamic vision sensors that capture video frames may be one or more;
[0247] Among them, when the number of dynamic vision sensors that capture video frames is one, the dynamic vision sensor can capture video frames including the entire body of the user or local body parts. Among them, in the implementation where the dynamic vision sensor captures video frames including local body parts, the dynamic vision sensor can only select the parts that vibrate accordingly based on the vocal behavior when the user is in a vocal state to capture video frames. The body parts can be, for example, the top of the skull, face, throat or neck.
[0248] In one implementation, the video acquisition direction of the dynamic vision sensor can be pre-set. For example, in the application scenario of the intelligent driving system, the dynamic vision sensor can be set at a preset position in the car, and the video acquisition direction of the dynamic vision sensor can be set to face the preset body part of the user. Taking the preset body part as the face as an example, the video acquisition direction of the dynamic vision sensor can be towards the preset area of the driving seat, which is usually the area where the face is located when there is a person sitting in the driving seat.
[0249] In one implementation, the dynamic vision sensor can capture video frames that include the entire user's body. The video capture direction of the dynamic vision sensor can also be pre-set. For example, in an intelligent driving system application scenario, the dynamic vision sensor can be set at a pre-set position in the vehicle and set to face the driver's seat.
[0250] In one implementation, there are multiple dynamic vision sensors, and each dynamic vision sensor can pre-set its video acquisition direction so that each dynamic vision sensor can capture a video frame including a body part, where the body part is the part that vibrates accordingly based on the vocalization behavior when the user is in a vocal state. For example, in the application scenario of an intelligent driving system, dynamic vision sensors can be deployed in front and behind the headrest (to capture video frames of people facing forward and backward in the car), on the car frame (to capture video frames of people facing left and right), and under the windshield (to capture video frames of people in the front row).
[0251] In the embodiments of the present application, the video frames of different body parts can be collected using the same sensor, such as using high-speed cameras, or using dynamic vision sensors, or using a mixture of these two sensors, which is not limited in the present application.
[0252] In the application scenarios of smart homes, dynamic vision sensors can be deployed on TVs, smart screens, smart speakers, etc. In the application scenarios of smartphones, dynamic vision sensors can be deployed on mobile phones, such as based on the front or rear camera of the mobile phone.
[0253] In the embodiment of the present application, the vibration signal represents the original characteristics of the human voice; optionally, there may be multiple vibration signals, such as: a vibration signal x1(n) of the head, a vibration signal x2(n) of the throat, a vibration signal x3(n) of the face, a vibration signal x4(n) of the neck, and so on. In the embodiment of the present application, the corresponding target audio signal can be recovered based on the vibration signals.
[0254] In an embodiment of the present application, the vibration signal is used to represent the vibration characteristics of the body part of the user when emitting the voice signal. The vibration characteristics can be vibration characteristics directly obtained from the video or vibration characteristics that are only related to the vocal vibration after other action interference is filtered out.
[0255] For video frames captured by a high-speed camera, filters in different directions can be used to decompose the image into a pyramid of images at different scales and orientations. Specifically, the image is first filtered with a low-pass filter to obtain a low-pass residual image. This low-pass residual image is then continuously downsampled to images of different scales. Each image scale is then filtered with bandpass filters in different directions to obtain response maps for each direction. The amplitude and phase of the response maps are then calculated, and the local motion information of the current frame t is calculated. The first frame is used as the reference frame. Based on the pyramid results, the phase difference between the decomposition results of the current frame and the reference frame at different pixel positions at different scales and orientations is calculated to quantify the magnitude of local motion at each pixel. Based on the local motion of each pixel, the global motion information of the current frame is calculated. The global motion information can be obtained by weighted averaging the local motion information. The weight is the amplitude corresponding to the scale, direction, and pixel position. The weighted sum of all pixels in this direction and scale is taken to obtain global motion information of different scales and directions. By summing the above global information, the global motion information of this image frame can be obtained. Based on the above steps, a motion magnitude value can be calculated for each image frame. Based on the continuous frame rate, the amplitude corresponding to each frame is used as the audio sampling value to obtain a preliminary restored audio signal. After high-pass filtering, the restored audio signal x'(n) is obtained. Optionally, if there are multiple vibration signals, the corresponding target audio signals x1'(n), x2'(n), x3'(n), and x4'(n) can be restored separately based on the above method.
[0256] Regarding video frames captured by dynamic vision sensors: Since dynamic vision sensors operate on the principle that each pixel independently responds to changes in light intensity, they compare the current light intensity with the intensity at the time of the previous event. When the difference between the two (i.e., the difference) exceeds a threshold, a new event is generated. Each event includes the pixel coordinates, the emission time, and the light intensity polarity. The light intensity polarity represents the trend of light intensity change, typically using +1 or "On" to indicate an increase in light intensity, and -1 or "Off" to indicate a decrease in light intensity. Since dynamic vision sensors have no concept of exposure, pixels continuously monitor and respond to light intensity, achieving microsecond temporal resolution. Furthermore, dynamic vision sensors are sensitive to motion and almost insensitive to static areas. They can be used to capture the vibration of an object, thereby enabling sound recovery. This results in an audio signal recovered at a specific pixel location. High-pass filtering removes low-frequency, non-audio vibration interference, resulting in the signal x'(n), which can represent the audio signal. The audio signals recovered from multiple pixels, such as all pixels, can be weighted and summed to obtain the weighted average audio signal x'(n) recovered by the dynamic vision sensor. If there are multiple sensors or multiple target locations, they can be recovered separately to obtain the independently recovered target audio signals x1'(n), x2'(n), x3'(n), and x4'(n).
[0257] 703. Obtain target voice information according to the vibration signal and the user voice signal collected by the sensor.
[0258] In an embodiment of the present application, a corresponding target audio signal can be restored based on the vibration signal; based on filtering, the target audio signal is filtered out from the audio signal to obtain a signal to be filtered out; and the signal to be filtered out is filtered out from the voice signal to obtain target voice information, wherein the target voice information is the voice signal obtained after removing the environmental noise.
[0259] Specifically, the corresponding target audio signal can be restored according to the vibration signal, and based on filtering, the target audio signal is filtered out from the audio signal to obtain a signal to be filtered. After filtering, the filtered signal z'(n) basically does not contain the useful signal x'(n), which is basically the external noise except the user's target audio signal s(n); optionally, if multiple cameras (DVS, high-speed camera, etc.) pick up the vibration of a certain person, the target audio signals x1'(n), x2'(n), x3'(n), and x4'(n) recovered from these vibrations are filtered out from the mixed audio signal z(n) in turn according to the above-mentioned adaptive filtering method, that is, a mixed audio signal z'(n) is obtained from which various x1'(n), x2'(n), x3'(n), and x4'(n) audio components are removed.
[0260] In an embodiment of the present application, the signal to be filtered can be filtered out from the audio signal to obtain the user's voice signal; in one implementation, a noise spectrum can be obtained (that is, z'(n) is considered to be background noise other than the target voice signal s(n)): and z'(n) is transformed into the frequency domain, such as fast Fourier transform (FFT) to obtain the noise spectrum; the target audio signal z(n) is transformed into the frequency domain, such as FFT to obtain the frequency spectrum, and then the frequency spectrum is subtracted from the noise spectrum to obtain the signal spectrum of the enhanced voice; finally, the signal spectrum is subjected to inverse fast Fourier transform (IFFT) to obtain the user's voice signal, that is, the voice enhanced signal.
[0261] In one implementation, the signal to be filtered is filtered out from the audio signal by using an adaptive filtering method.
[0262] It should be noted that the above method of filtering the signal to be filtered from the audio signal to obtain the speech signal after the environmental noise is removed is only some examples and is not limited in this application.
[0263] In one implementation, command information corresponding to the user's voice signal may also be obtained based on the target voice information, where the command information indicates the semantic intent contained in the user's voice signal. The command information may be used to trigger a function corresponding to the semantic intent contained in the user's voice signal, such as opening an application, making a voice call, etc.
[0264] In one implementation, the target voice information may be obtained through a neural network model based on the vibration signal and the voice signal.
[0265] In one implementation, a corresponding target audio signal is obtained based on the vibration signal; and the target voice information is obtained through a neural network model based on the target audio signal and the voice signal. That is, the input of the neural network model can also be the target audio signal recovered from the vibration signal.
[0266] The following describes a system architecture provided by an embodiment of the present application.
[0267] See attached Figure 8, an embodiment of the present invention provides a system architecture 200. The data acquisition device 260 is used to collect audio data and store it in the database 230; wherein, the audio data may include noise-free audio, vibration signals (or target audio signals recovered from vibration signals) and noisy audio signals; wherein, a person can speak / play audio in a quiet environment, and the audio signal at this time can be recorded with an ordinary microphone as "noise-free audio", recorded as s(n). At the same time, a vibration sensor (can be multiple) is used to face the person's head, face, throat, neck, etc., to collect video frames during this period and obtain corresponding vibration signals, recorded as x(n). If there are multiple sensors, the signals can be recorded as x1(n), x2(n), x3(n), x4(n), etc. Or the target audio signal recovered from the vibration signal. Various types of noise can be added to the "noise-free audio" to obtain a "noisy audio signal", recorded as sn(n).
[0268] Training device 220 generates target model / rule 201 based on the audio data maintained in database 230. The following describes in more detail how training device 220 obtains target model / rule 201 based on the audio data. Target model / rule 201 can obtain the target voice information or the user's voice signal based on the vibration signal and the audio signal.
[0269] Here, the training process description is adaptively added based on the actual embodiment. If the inventive point is not in the training process, the training process description in the following example will be used. If the training process is improved, please use the improved training process instead of the training process description below.
[0270] The training device can use a deep neural network to train the data to generate a target model / rule 201. The work of each layer in the deep neural network can be expressed by a mathematical expression To describe: From a physical perspective, the work of each layer in a deep neural network can be understood as completing the transformation from input space to output space (i.e., from the row space to the column space of a matrix) through five operations on the input space (a set of input vectors). These five operations include: 1. Dimensionality increase / decrease; 2. Zoom in / out; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are represented by Completed, operation 4 is completed by +b, and operation 5 is implemented by a(). The word "space" is used here because the object being classified is not a single thing, but a class of things, and space refers to the collection of all individuals of this class of things. Among them, W is a weight vector, and each value in this vector represents the weight value of a neuron in this layer of the neural network. This vector W determines the spatial transformation from the input space to the output space mentioned above, that is, the weight of each layer controls how to transform the space. The purpose of training a deep neural network is to eventually obtain the weight matrix of all layers of the trained neural network (the weight matrix formed by many layers of vectors W). Therefore, the training process of a neural network is essentially learning how to control spatial transformation, and more specifically learning the weight matrix.
[0271] Because we want the output of a deep neural network to be as close as possible to the value we really want to predict, we can compare the current network's predicted value with the target value we really want, and then update the weight vector of each layer of the neural network based on the difference between the two (of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the deep neural network). For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value, and this adjustment is continued until the neural network can predict the target value we really want. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value." This is the loss function or objective function, which are important equations used to measure the difference between the predicted value and the target value. Taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then, the training of the deep neural network becomes a process of minimizing this loss as much as possible.
[0272] The target model / rule obtained by the training device 220 can be applied to different systems or devices. Figure 8 In the embodiment, the execution device 210 is configured with an I / O interface 212 for data interaction with external devices, and a “user” can input data into the I / O interface 212 through a client device 240 .
[0273] The execution device 210 can call data, code, etc. in the data storage system 250 , and can also store data, instructions, etc. in the data storage system 250 .
[0274] The calculation module 211 processes the input data using the target model / rules 201 .
[0275] Finally, the I / O interface 212 returns the processing result (the user's instruction information or the user's voice signal) to the client device 240 and provides it to the user.
[0276] More deeply, the training device 220 can generate corresponding target models / rules 201 based on different data for different goals to provide users with better results.
[0277] In the attached Figure 2 In the case shown in , the user can manually specify the data to be input into the execution device 210, for example, by operating in the interface provided by the I / O interface 212. In another case, the client device 240 can automatically input data into the I / O interface 212 and obtain the results. If the automatic data input of the client device 240 requires user authorization, the user can set the corresponding permissions in the client device 240. The user can view the results output by the execution device 210 on the client device 240, and the specific presentation form can be specific methods such as display, sound, and action. The client device 240 can also serve as a data acquisition terminal to store the collected audio data in the database 230.
[0278] It is worth noting that Figure 2 This is only a schematic diagram of a system architecture provided by an embodiment of the present invention. The positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in the attached Figure 2 In the embodiment, the data storage system 250 is an external memory relative to the execution device 210. In other cases, the data storage system 250 can also be placed in the execution device 210.
[0279] Next, the neural network model provided by the embodiment of the present application is described from the training side:
[0280] During the training data preparation phase, a person can speak or play audio in a quiet environment. A standard microphone is used to record the audio signal, which is referred to as "noise-free audio," denoted as s(n). Simultaneously, a vibration sensor (or multiple sensors) is pointed at the person's head, face, throat, neck, and other areas to capture visual signals, denoted as x(n). If multiple sensors are used, the signals can be denoted as x1(n), x2(n), x3(n), x4(n), and so on. The aforementioned algorithm is used to recover the audio signal from x(n), yielding the "visual audio signal" x'(n) captured and recovered by the visual microphone. If multiple sensors are used, the recovered audio signals are x1'(n), x2'(n), x3'(n), and x4'(n). Various types of noise are added to the "noise-free audio" to obtain the "noisy audio signal," denoted as sn(n).
[0281] Based on the collected data, a deep model is trained to learn the mapping relationship between this "noisy audio signal" (z(n) collected by the microphone) and "visual vibration audio signal" (x(n) collected by the visual vibration sensor, etc.) to the "noise-free audio signal" (enhanced speech signal s'(n)).
[0282] A deep neural network that considers temporal relationships can be used, such as a recurrent neural network (RNN) or a long short-term memory network (LSTM).
[0283] First, some definitions of technical terms involved in the embodiments of this application are given:
[0284] (1) Neural Network
[0285] A neural network can be composed of neural units. A neural unit can refer to an operation unit with xs and intercept 1 as input. The output of the operation unit can be:
[0286]
[0287] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting many of the above-mentioned single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0288] (2) Deep Neural Networks
[0289] A deep neural network (DNN) can be understood as a neural network with many hidden layers. There is no special metric for "many" here. The multi-layer neural networks and deep neural networks we often talk about are essentially the same thing. According to the position of different layers of DNN, the neural network inside the DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the coefficient W and the offset vector So how are the specific parameters defined in DNN? First, let's look at the definition of coefficient W. Take a three-layer DNN as an example, for example: the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscripts correspond to the output of the third layer index 2 and the input of the second layer index 4. In summary, the coefficients from the kth neuron in the L-1th layer to the jth neuron in the Lth layer are defined as Note that the input layer has no W parameter. In deep neural networks, more hidden layers allow the network to better capture complex real-world situations. Theoretically, a model with more parameters is more complex and has greater "capacity," meaning it can handle more complex learning tasks.
[0290] (3) Convolutional Neural Network (CNN) is a deep neural network with a convolutional structure. A convolutional neural network contains a feature extractor consisting of a convolution layer and a subsampling layer. The feature extractor can be regarded as a filter, and the convolution process can be regarded as using a trainable filter to convolve an input image or convolution feature plane (feature map). The convolution layer refers to the neuron layer in the convolutional neural network that performs convolution processing on the input signal. In the convolution layer of the convolutional neural network, a neuron can only be connected to some neurons in the adjacent layer. A convolution layer usually contains several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights here are the convolution kernels. Shared weights can be understood as the way to extract image information is independent of position. The implicit principle is that the statistical information of a part of the image is the same as that of other parts. This means that the image information learned in a part can also be used in another part. Therefore, for all positions on the image, we can use the same learned image information. In the same convolutional layer, multiple convolution kernels can be used to extract different image information. Generally speaking, the more convolution kernels there are, the richer the image information reflected by the convolution operation.
[0291] Convolution kernels can be initialized as matrices of random size, and during the training process of the convolutional neural network, the convolution kernels can be learned to obtain reasonable weights. In addition, the direct benefit of shared weights is that they reduce the number of connections between the layers of the convolutional neural network, while also reducing the risk of overfitting.
[0292] (4) Recurrent Neural Networks (RNN)
[0293] RNNs are designed to process sequential data. In traditional neural network models, layers are fully connected, from the input layer to the hidden layer to the output layer, and the nodes within each layer are disconnected. However, these common neural networks are inadequate for many problems. For example, if you want to predict the next word in a sentence, you generally need to use the previous word, because the previous and next words in a sentence are not independent. RNNs are called recurrent neural networks because the current output of a sequence is also related to the previous output. Specifically, the network remembers previous information and applies it to the calculation of the current output. In other words, the nodes between hidden layers are no longer disconnected but connected, and the input of the hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous moment. In theory, RNNs can process sequence data of any length.
[0294] Training an RNN is similar to training a traditional ANN (artificial neural network). The same back propagation error algorithm is used, but there is a slight difference. If the RNN is expanded, the parameters W, U, and V are shared, while traditional neural networks are not. Furthermore, when using the gradient descent algorithm, the output of each step depends not only on the network state at the current step, but also on the state of the network at the previous steps. For example, at t = 4, three steps need to be passed backward, and various gradients need to be added to the three steps. This learning algorithm is called back propagation through time (BPTT).
[0295] Given the existence of artificial neural networks and convolutional neural networks, why do we still need recurrent neural networks? The reason is simple. Both convolutional and artificial neural networks assume that elements are independent of each other, and that inputs and outputs are also independent, like cats and dogs. However, in the real world, many elements are interconnected, such as the changes in stock prices over time. For example, someone said, "I love traveling, and my favorite place is Yunnan. I must visit __ someday." Everyone knows to fill in the blank with "Yunnan." This is because we infer this information based on the context, but achieving this is quite difficult. Therefore, recurrent neural networks were developed. Their essence is that they possess memory, just like humans. Therefore, their output depends on the current input and memory.
[0296] Figure 9It is a schematic diagram of the structure of an RNN. Each circle can be regarded as a unit, and each unit does the same thing. Therefore, it can be folded into the form of the left half of the figure. To explain the RNN in one sentence, it is that a unit structure is reused.
[0297] The RNN is a sequence-to-sequence model. Suppose \(x_{t - 1}\), \(x_t\), \(x_{t + 1}\) is an input: "I am Chinese". Then \(o_{t - 1}\), \(o_t\) should correspond to "am", "Chinese" respectively. What is the most likely next word? That is, the probability that \(o_{t + 1}\) should be "person" is relatively high.
[0298] Therefore, we can make the following definitions:
[0299] \(X_t\): represents the input at time \(t\), \(o_t\): represents the output at time \(t\), \(S_t\): represents the memory at time \(t\). Because the output at the current time is determined by the memory and the input at the current time. Just like when you are in your senior year, your knowledge is the combination of the knowledge learned in your senior year (current input) and the things learned in your junior year and before (memory). The RNN is similar in this regard. What neural networks are best at is integrating a lot of content together through a series of parameters and then learning these parameters. Therefore, the basis of the RNN is defined as: \(S_t = f(U * X_t+W * S_{t - 1})\);
[0300] The \(f()\) function is an activation function in the neural network. But why add it? For example, if you have learned a very good problem-solving method in college, do you still need to use the problem-solving method from junior high school? Obviously not. The idea of the RNN is the same way. Since it can remember, of course it only remembers important information and forgets other unimportant information. But what is most suitable for filtering information in a neural network? It must be the activation function. Therefore, an activation function is applied here to make a non-linear mapping to filter information. This activation function may be tanh or others.
[0301] Suppose you are about to graduate from your senior year and want to take the postgraduate entrance examination. Do you first remember the content you have learned and then take the exam, or just take a few books to take the exam directly? Obviously, the idea of the RNN is to carry the memory \(S_t\) at the current time when predicting. If you want to predict the probability of the next word after “I am Chinese”, it is already obvious here. It is most appropriate to use softmax to predict the probability of each word appearing. But the prediction cannot be directly made with a matrix. Therefore, a weight matrix \(V\) is also carried during prediction. It is expressed by the formula:
[0302] \(o_t=\text{softmax}(VS_t)\) where \(o_t\) represents the output at time \(t\).
[0303] (5) Backpropagation algorithm
[0304] Convolutional neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial super-resolution model during training, reducing the reconstruction error loss of the super-resolution model. Specifically, the forward propagation of the input signal to the output generates an error loss. This error loss information is then backpropagated to update the parameters of the initial super-resolution model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by the error loss, aiming to obtain the optimal super-resolution model parameters, such as the weight matrix.
[0305] Taking the neural network model in the embodiment of the present application as RNN as an example, the neural network structure in the embodiment of the present application can be as follows:
[0306] In one implementation, referring to Figure 10 , RNN whose input is the target audio signal obtained according to the vibration signal, and the user voice signal collected by the sensor. Among them, the target audio signal includes multiple moments and the signal sampling value corresponding to each moment, and the user voice signal collected by the sensor includes multiple moments and the signal sampling value corresponding to each moment. In one implementation, the signal sampling value of the target audio signal and the signal sampling value of the user voice signal collected by the sensor can be combined at each moment to obtain a new audio signal. The new audio signal includes multiple moments and the signal sampling value corresponding to each moment, wherein the signal sampling value corresponding to each moment is obtained by combining the signal sampling value of the target audio signal and the signal sampling value of the user voice signal (this application does not limit the specific combination method). The new audio signal obtained after the combination can be used as the input of the recurrent neural network.
[0307] In one implementation, the target audio signal may be {x0, x1, x2, ..., x t}; where x t is the signal sampling value of the target audio signal at time t, and the user voice signal can be {y0, y1, y2, ..., y t}; where y t is the signal sampling value of the user's voice signal at time t. When combining, the signal sampling values at corresponding times can be combined, for example, {x t ,y t} is the result of combining the signal sampling values at time t, then the new audio signal obtained by combining the target audio signals can be {{x0, y0}, {x1, y1}, {x2, y2}, ..., {x t ,y t}}.
[0308] It should be noted that the above-mentioned combination of input audio signals is only an illustration. In actual applications, the combined audio signal can express the timing characteristics of the signal sampling value. This application does not limit the specific combination method.
[0309] It should be noted that the input of the model can be a combined audio signal. In another implementation, the input of the model can be a user voice signal and a target audio signal collected by a sensor. In this case, the combination of the audio signals can be implemented by the model itself. In another implementation, it can be a user voice signal collected by a sensor and a vibration signal corresponding to when the user utters the voice. The process of converting the vibration signal into the target audio signal can be implemented by the model itself, that is, the model can first convert the vibration signal into the target audio signal and then combine the audio signals.
[0310] After the combined audio signal obtained above is input into the RNN, the target speech information can be output. The target speech information can include multiple moments and the signal sampling value corresponding to each moment. For example, the target speech information can be {k0, k1, k2, ..., k l It should be noted that the number of signal sampling values (number of time instants) included in the target voice information may be the same as or different from the number of signal sampling values (number of time instants) included in the input audio signal. For example, if the target voice information is only the voice information related to the human voice in the user voice signal collected by the sensor, the number of signal sampling values included in the target voice information may be smaller than the number of signal sampling values included in the input audio signal.
[0311] s t is the state of the hidden layer at step t, which is the memory unit of the network. t According to the output x of the current input layer t and the state s of the hidden layer in the previous step t-1 Perform calculations. t =f(Ux t +Ws t-1 ), where f is generally a nonlinear activation function, such as tanh or ReLU function; when calculating s0, that is, the hidden layer state of the first moment feature, s is needed -1 , but it does not exist and is generally set to 0 vector in implementation; o t is the output of step t, o t =g(Vs t ). g is a linear or nonlinear function.
[0312] During the model training process, the data in the training sample database can be input into the initialized neural network model for training, wherein the training sample database includes a pair of "speech signal with environmental noise", "target audio signal" and corresponding "noise-free audio signal", and the initialized neural network model includes weights and biases; during the K-th training process, the denoised audio signal s'(n) is extracted from the audio features of the noisy audio signal and the visual audio signal of the sample through the neural network model that has been adjusted K-1 times, where K is an integer greater than 0; after the K-th training, the error value between the denoised audio signal s'(n) extracted from the sample and the noise-free audio signal s(n) is obtained; based on the error value between the denoised audio signal and the noise-free audio signal extracted from the sample of the sample video frame, the weights and biases used in the K+1-th training process are adjusted.
[0313] The audio signal obtained from the above combination contains two dimensions: the user voice signal collected by the sensor and the target audio signal. Since both are audio signals (the audio signal x'(n) must first be decoded based on the vibration signal x(n) to recover the audio signal), the MFCC coefficients of the audio feature vectors can be extracted for each. MFCC features are the most widely used basic features in speech recognition and speaker recognition. MFCC features are based on the characteristics of the human ear: the human ear's perception of sound frequencies above approximately 1000 Hz does not follow a linear relationship, but rather an approximately linear relationship on a logarithmic frequency scale. MFCCs are cepstral parameters extracted in the Mel-scale frequency domain. The Mel-scale describes this nonlinear characteristic of the human ear's frequency response.
[0314] The extraction of MFCC features can include the following steps: Preprocessing: It consists of pre-emphasis and frame windowing. The purpose of pre-emphasis is to eliminate the influence of nasal radiation during pronunciation, and to enhance the high-frequency part of the speech through a high-pass filter. Since the speech signal is short-term and stable, the speech signal is divided into short time periods through frame windowing, and each short time period is called a frame. At the same time, in order to avoid the loss of dynamic information of the speech signal, there must be an overlapping area between adjacent frames. The FFT transform changes the time domain signal after frame windowing to the frequency domain to obtain the spectrum feature X(k). After filtering the speech frame spectrum feature X(k) through the above-mentioned Mel filter bank, the energy of each subband is obtained, and then the logarithm is taken to obtain the Mel-frequency logarithmic energy spectrum S(m); S(m) is subjected to discrete cosine transform (DCT) to obtain the MFCC coefficient C(n). When constructing the feature vector, if there are multiple visual vibration audio signals, then: after averaging the multiple signals, extract the MFCC coefficients as the "visual audio signal"; extract the MFCC coefficients for each visual / vibration audio signal separately, and then combine and string these coefficients together to form a larger feature vector. Extract the MFCC coefficients for each visual / vibration audio signal separately, and then take the average of these MFCC coefficients, and use the averaged MFCC coefficients as the feature vector of the "visual audio signal".
[0315] In addition to constructing the feature vector using all the audio information as described in the above embodiment, it can also be constructed directly based on the vibration signal obtained from the video and the audio information. At this time, the constructed audio feature vector still contains two dimensions, namely "noisy audio signal", but "visual vibration audio signal" is replaced by "visual vibration signal", where the signal is acquired as follows: If it is a high-speed camera, for each frame of the image, four scales r (such as 1, 1 / 4, 1 / 16, 1 / 64) and four directions θ (such as up, down, left and right) are used. For each scale and direction, the value is as follows: In this way, 16 eigenvalues of vibration information are obtained, which can form the eigenvector For a DVS sensor (brain-inspired camera), within a specific audio frame interval, such as T, a random pixel-by-pixel vibration offset S(t) can be selected within a subinterval of T / N. This selection is repeated within each subinterval to obtain 16 vibration offset values. These offset values are then combined into a feature vector, which is then used as the audio signal feature for training based on the RNN neural network described above. Furthermore, features extracted from the noisy audio signal, the vibration recovery signal, and the vibration signal can be combined into an audio signal feature vector for training.
[0316] In one implementation, instead of using feature vectors, features are extracted directly from the original multimodal data and applied in the network. Specifically, a deep network can be trained to learn the mapping relationship between "noisy audio signals" and "vibration signals" to noise-free speech signals. RNNs contain input units, and the corresponding input set is labeled {x0, x1, x2, ..., x t , x t+1 , ...}, and the output set of the output unit is labeled {o0, o1, o2, ..., o t , o t+1 ,...}. RNNs also contain hidden units whose output sets are labeled {s0, s1, s2, ..., s t , s t+1 ,...}. x t Represents the input of step t=1, 2, 3..., corresponding to the fused feature signal at time t. Here is the "noisy audio signal" and "visual vibration signals" The connected multi-mode mixed signal x t =[sn (0) ,...,sn (t), sv (0) ,...,sv (t) ].
[0317] s t is the state of the hidden layer at step t, which is the memory unit of the network. t According to the output x of the current input layer t and the state s of the hidden layer in the previous step t-1 Perform calculations. t =f(Ux t +Ws t-1 ), where f is generally a nonlinear activation function, such as tanh or ReLU function; when calculating s0, that is, the hidden layer state of the first moment feature, s is needed -1 , but it does not exist and is generally set to 0 vector in implementation; o t is the output of step t, o t =g(Vs t ), function g is a linear or nonlinear function.
[0318] In a specific training process, data in a training sample database can be input into an initialized neural network model for training, wherein the training sample database includes a pair of "noisy audio signals", "visual vibration signals" and corresponding "noise-free audio signals", and the initialized neural network model includes weights and biases; during the K-th training process, the denoised audio signal s'(n) is extracted from the noisy audio signal and visual vibration signal of the sample through learning by the neural network model that has been adjusted K-1 times, where K is an integer greater than 0; after the K-th training, the error value between the denoised audio signal s'(n) extracted from the sample and the noise-free audio signal s(n) is obtained; based on the error value between the denoised audio signal and the noise-free audio signal extracted from the sample of the sample video frame, the weights and biases used in the K+1-th training process are adjusted. It should be noted that when training the model, the vibration audio signal x'(n) (and or x1'(n), x2'(n), x3'(n), x4'(n)) can be used instead of the visual vibration signal x(n) (and or x1(n), x2(n), x3(n), x4(n)) for training, that is, the audio signal restored from the vibration signal is fused with the audio signal collected by the microphone for training.
[0319] In an embodiment of the present application, the output of the neural network model can be a voice signal obtained after removing environmental noise, or a user's command information, wherein the command information is determined based on the user's voice signal, and the command information is used to indicate the user's intention carried in the user's voice signal. The device with voice interaction function can trigger the corresponding function based on the command information, such as opening a certain application, etc.
[0320] The embodiment of the present application provides a method for processing a speech signal, comprising: obtaining a speech signal of a user collected by a sensor, wherein the speech signal includes environmental noise; obtaining a vibration signal corresponding to when the user utters the speech; wherein the vibration signal is used to represent the vibration characteristics of a body part of the user; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and obtaining target speech information based on the vibration signal and the user speech signal collected by the sensor. In the above manner, the vibration signal is used as the basis for speech recognition. Since the vibration signal does not contain external non-user speech mixed in during complex acoustic transmission, it is less affected by other environmental noises (such as reverberation). Therefore, this part of noise interference can be relatively well suppressed, and a better speech recognition effect can be achieved.
[0321] In one implementation, the user's brainwave signal corresponding to the user uttering the voice can also be obtained; accordingly, the target voice information can be obtained based on the vibration signal, the brainwave signal, and the voice signal. The user's brainwave signal can be obtained based on a brainwave pickup device, which can be headphones, glasses, or other ear-worn devices.
[0322] In this embodiment, a mapping relationship table between brain wave signals and vocal tract occlusion movements can be established, and brain wave signals and vocal tract occlusion movement signals of a person when reading various morphemes and sentences can be collected. The brain wave signals are collected by an EEG acquisition device (for example, including electrodes, front-end analog amplifiers, analog-to-digital conversion, EEG signal processing and other modules) in multiple brain regions according to different frequency bands. The movement signals of the vocal tract occlusion can be collected by an electromyography acquisition device or an optical imaging device. After that, a mapping relationship between the brain wave signals and the movement signals of the vocal tract occlusion of a person when using different corpus materials can be established. At this time, after obtaining the brain wave signal of the user corresponding to when the user utters the voice, the movement signal of the vocal tract occlusion corresponding to the brain wave signal can be obtained based on the mapping relationship between the brain wave signal and the movement signal of the vocal tract occlusion.
[0323] In this embodiment, the brain wave signal can be converted into the joint movement (motion signal) of the vocal tract occlusion, and then these decoded movements are converted into speech signals. That is, the brain wave signal is first converted into the movement of the vocal tract occlusion part, which involves the anatomical structure of speech production (such as the movement signals of the lips, tongue, larynx and mandible). In order to realize the conversion and mapping of the brain wave signal to the movement of the vocal tract occlusion part, it is necessary to associate a large number of vocal tract movements with their neural activities when a person speaks. This association can be established based on the established recurrent neural network and a large number of previously collected vocal tract movement and speech recording data sets, and the movement signal of the vocal tract occlusion part can be converted into a speech signal.
[0324] In one implementation, the target voice information may be obtained through a neural network model based on the vibration signal, the motion signal, and the voice signal.
[0325] In one implementation, referring to Figure 11 , a corresponding first target audio signal can be obtained based on the vibration signal; a corresponding second target audio signal can be obtained based on the motion signal; and the target voice information can be obtained using a neural network model based on the first target audio signal, the second target audio signal, and the voice signal. Specific implementation details can be found in the description of the neural network model in the above embodiment and will not be repeated here.
[0326] In another implementation, a speech signal can be obtained based on direct mapping of the brain wave signal, and further, the target speech information can be obtained through a neural network model based on the vibration signal, the brain wave signal and the speech signal.
[0327] In one implementation, referring to Figure 11 , a corresponding first target audio signal can be obtained based on the vibration signal; a corresponding second target audio signal can be obtained based on the brainwave signal; and the target voice information can be obtained using a neural network model based on the first target audio signal, the second target audio signal, and the voice signal. Specific implementation details can be found in the description of the neural network model in the above embodiment and will not be repeated here.
[0328] In one implementation, the target voice information includes a voiceprint feature representing a voice signal of the user.
[0329] In one implementation, the target voice information can be obtained through a neural network model based on the vibration signal and the voice signal. The target voice information can be used to represent the voiceprint features of the user's voice signal. Furthermore, the target voice information can be processed based on a fully connected layer to obtain a voiceprint recognition result.
[0330] In one implementation, referring to Figure 12 , a corresponding target audio signal can be obtained according to the vibration signal; based on the target audio signal and the voice signal, the target voice information is obtained through a neural network model.
[0331] In one implementation, the target voice information can be obtained through a neural network model based on the vibration signal, the brain wave signal and the voice signal.
[0332] In one implementation, referring to Figure 13 , a corresponding first target audio signal can be obtained according to the vibration signal; a corresponding second target audio signal can be obtained according to the brain wave signal; based on the first target audio signal, the second target audio signal and the voice signal, the target voice information is obtained through a neural network model.
[0333] The construction method of the model can refer to the above Figure 10 The description in the corresponding embodiments will not be repeated here.
[0334] In an embodiment of the present application, the vibration signal when the user speaks is used as the basis for voiceprint recognition. Since the vibration signal is less affected by other noise interference (such as reverberation interference, etc.), it can express the original audio characteristics of the user's speech. Therefore, the present application uses the vibration signal as the basis for voiceprint recognition, which has better recognition effect and stronger reliability.
[0335] Reference Figure 14 , Figure 14 A flow chart of a method for processing a speech signal provided in an embodiment of the present application is shown as follows: Figure 14 As shown in , the method includes:
[0336] 1401. Acquire a user's voice signal collected by a sensor.
[0337] The specific description of step 1401 can refer to the description of step 701 and will not be repeated here.
[0338] 1402. Obtain a brain wave signal of the user corresponding to when the user utters the voice.
[0339] For a detailed description of step 1402 , reference may be made to the detailed description related to the brain wave signal in the above embodiment, which will not be repeated here.
[0340] 1403. Obtain target voice information according to the brain wave signal and the user voice signal collected by the sensor.
[0341] In this embodiment of the present application, after obtaining the user's voice signal and the user's brainwave signal, target voice information can be obtained based on the brainwave signal and the voice signal. Unlike step 703 described above, this embodiment is based on the brainwave signal and the voice signal. The description of step 703 in the above embodiment for obtaining the target voice information based on the brainwave signal and the voice signal can be referenced and will not be repeated here.
[0342] In an embodiment of the present application, the movement signal of the vocal tract occlusion part of the user when speaking can also be obtained based on the brain wave signal; further, the target voice information can be obtained based on the movement signal and the voice signal.
[0343] Optionally, in one implementation, the target voice information is a voice signal obtained after environmental noise removal processing, and the corresponding target audio signal can be obtained based on the brain wave signal; based on filtering, the target audio signal is filtered out from the voice signal to obtain the signal to be filtered; and the signal to be filtered out is filtered out from the voice signal to obtain the target voice information.
[0344] Optionally, in one implementation, instruction information corresponding to the user's voice signal may be acquired based on the target voice information, where the instruction information indicates the semantic intention contained in the user's voice signal.
[0345] Optionally, in one implementation, the target voice information can be obtained through a neural network model based on the brain wave signal and the voice signal; or, the corresponding target audio signal can be obtained according to the brain wave signal; the target voice information can be obtained through a neural network model based on the target audio signal and the voice signal; wherein, the target voice information is a voice signal obtained after removing environmental noise processing or instruction information corresponding to the user's voice signal.
[0346] In one implementation, the target voice information includes a voiceprint feature representing a voice signal of the user.
[0347] An embodiment of the present application provides a method for processing speech signals, the method comprising: obtaining a speech signal of a user collected by a sensor; obtaining a brain wave signal of the user corresponding to when the user utters the speech; and obtaining target speech information based on the brain wave signal and the user speech signal collected by the sensor. In the above manner, the vibration signal is used as the basis for speech recognition. Since the vibration signal does not contain external non-user speech mixed in during complex acoustic transmission, it is less affected by other environmental noises (such as reverberation). Therefore, this part of the noise interference can be relatively well suppressed, and a better speech recognition effect can be achieved.
[0348] Reference Figure 15 , Figure 15 A flow chart of a method for processing a speech signal provided in an embodiment of the present application is shown as follows: Figure 15 As shown in , the method includes:
[0349] 1501. Acquire a user's voice signal collected by a sensor.
[0350] The specific description of step 1501 can refer to the description of step 701 and will not be repeated here.
[0351] 1502. Obtain a vibration signal corresponding to the user uttering the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state;
[0352] The specific description of step 1502 can refer to the description of step 702 and will not be repeated here.
[0353] 1503. Perform voiceprint recognition based on the user voice signal and the vibration signal collected by the sensor.
[0354] In one implementation, the vibration signal is used to represent a vibration feature corresponding to the sound-generating vibration.
[0355] In one implementation, voiceprint recognition is performed based on the user voice signal collected by the sensor to obtain a first confidence level that the user voice signal collected by the sensor belongs to the user; voiceprint recognition is performed based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the target user; and a voiceprint recognition result is obtained based on the first confidence level and the second confidence level. For example, the first confidence level and the second confidence level may be weighted to obtain the voiceprint recognition result.
[0356] In one implementation, the brain wave signal of the user corresponding to when the user utters the voice can be obtained; based on the brain wave signal, the movement signal of the vocal tract occlusion part when the user utters the voice can be obtained; and then, voiceprint recognition can be performed based on the user voice signal, the vibration signal and the movement signal collected by the sensor.
[0357] In one implementation, voiceprint recognition can be performed based on the user's voice signal collected by the sensor to obtain a first confidence level that the user's voice signal collected by the sensor belongs to the user; voiceprint recognition can be performed based on the vibration signal to obtain a second confidence level that the user's voice signal collected by the sensor belongs to the user; voiceprint recognition can be performed based on the brainwave signal to obtain a third confidence level that the user's voice signal collected by the sensor belongs to the user; and a voiceprint recognition result can be obtained based on the first confidence level, the second confidence level, and the third confidence level. For example, the first confidence level, the second confidence level, and the third confidence level can be weighted to obtain the voiceprint recognition result.
[0358] In one implementation, a voiceprint recognition result may be obtained through a neural network model based on the audio signal, the vibration signal, and the brainwave signal.
[0359] In this embodiment, if multiple audio signals are restored (including multiple vibration information or multiple target audio signals restored by brain wave signals), the audio x'(n), y'(n), x1'(n), x2'(n), x3'(n), and x4'(n) can be restored separately and voiceprint recognition can be performed separately. The final result is then provided by weighted summing the respective voiceprint recognition results:
[0360] VP=h1*x1+h2*x2+h3*x3+h4*x4+h5*x+h6*y+h7*s; where x1, x2, x3, x4, x, y, and s represent the recognition results of the vibration signal, brainwave signal, and audio signal, respectively. h1, h2, h3, h4, h5, h6, and h7 represent the weighting of the corresponding recognition results, which can be flexibly selected. If the final recognition result VP exceeds the preset threshold VP_TH, it means that the audio voiceprint result obtained during vibration pickup has passed.
[0361] In an embodiment of the present application, the vibration signal when the user speaks is used as the basis for voiceprint recognition. Since the vibration signal is less affected by other noise interference (such as reverberation interference, etc.), it can express the original audio characteristics of the user's speech. Therefore, the present application uses the vibration signal as the basis for voiceprint recognition, which has better recognition effect and stronger reliability.
[0362] Reference Figure 16 , Figure 16 This application provides a structural diagram of a speech signal processing device, such as Figure 16 As shown in FIG. 1 , the apparatus 1600 includes:
[0363] The environmental voice acquisition module 1601 is used to obtain the user voice signal collected by the sensor;
[0364] a vibration signal acquisition module 1602 for acquiring a vibration signal corresponding to the user uttering the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part being a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and
[0365] The voice information acquisition module 1603 is configured to obtain target voice information according to the vibration signal and the user voice signal collected by the sensor.
[0366] In an optional implementation, the vibration signal is used to represent a vibration feature corresponding to the vibration generated by uttering speech.
[0367] In an optional implementation, the body part includes at least one of the following: the top of the skull, the face, the throat, or the neck.
[0368] In an optional implementation, the vibration signal acquisition module 1602 is configured to acquire a video frame including the user; and extract, based on the video frame, a vibration signal corresponding to when the user utters a voice.
[0369] In an optional implementation, the video frames are acquired through a dynamic vision sensor and / or a high-speed camera.
[0370] In an optional implementation, the target voice information is a voice signal obtained after environmental noise removal processing, and the voice information acquisition module 1603 is used to obtain the corresponding target audio signal based on the vibration signal; based on filtering, the target audio signal is filtered out from the user voice signal collected by the sensor to obtain the noise signal to be filtered out; and the noise signal to be filtered out is filtered out from the user voice signal collected by the sensor to obtain the target voice information.
[0371] In an optional implementation, the apparatus further includes:
[0372] The instruction information acquisition module is used to acquire instruction information corresponding to the user's voice signal based on the target voice information, where the instruction information indicates the semantic intention contained in the user's voice signal.
[0373] In an optional implementation, the voice information acquisition module 1603 is used to obtain the target voice information through a recurrent neural network model based on the vibration signal and the user voice signal collected by the sensor; or, based on the vibration signal, obtain the corresponding target audio signal; based on the target audio signal and the user voice signal collected by the sensor, obtain the target voice information through a recurrent neural network model.
[0374] In an optional implementation, the apparatus further includes:
[0375] The brain wave signal acquisition module is used to obtain the brain wave signal of the user corresponding to when the user utters the voice; correspondingly, the voice information acquisition module is used to obtain target voice information based on the vibration signal, the brain wave signal and the user voice signal collected by the sensor.
[0376] In an optional implementation, the apparatus further includes:
[0377] The motion signal acquisition module is used to obtain the motion signal of the vocal tract occlusion part when the user speaks according to the brain wave signal; correspondingly, the voice information acquisition module is used to obtain target voice information according to the vibration signal, the motion signal and the user voice signal collected by the sensor.
[0378] In an optional implementation, the voice information acquisition module 1603 is configured to obtain the target voice information through a recurrent neural network model based on the vibration signal, the brain wave signal, and the user voice signal collected by the sensor; or
[0379] Acquire a corresponding first target audio signal according to the vibration signal;
[0380] According to the brain wave signal, a corresponding second target audio signal is obtained; based on the first target audio signal, the second target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
[0381] In an optional implementation, the target voice information includes a voiceprint feature representing a voice signal of the user.
[0382] The embodiment of the present application provides a speech signal processing device, which includes: an environmental speech acquisition module for acquiring a user speech signal collected by a sensor; a vibration signal acquisition module for acquiring a vibration signal corresponding to when the user utters the speech; wherein the vibration signal is used to represent the vibration characteristics of the user's body part; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and a speech information acquisition module for obtaining target speech information based on the vibration signal and the user speech signal collected by the sensor. In the above manner, the vibration signal is used as the basis for speech recognition. Since the vibration signal does not contain the external non-user speech mixed in during complex acoustic transmission, it is less affected by other environmental noises (such as reverberation). Therefore, this part of the noise interference can be relatively well suppressed, and a better speech recognition effect can be achieved.
[0383] Reference Figure 17 , Figure 17 This application provides a structural diagram of a speech signal processing device, such as Figure 17 As shown in FIG. 1 , the apparatus 1700 includes:
[0384] The environmental voice acquisition module 1701 is used to obtain the user's voice signal collected by the sensor;
[0385] The brain wave signal acquisition module 1702 is used to acquire the brain wave signal of the user corresponding to when the user utters the voice; and
[0386] The voice information acquisition module 1703 is configured to obtain target voice information based on the brain wave signal and the user voice signal collected by the sensor.
[0387] In an optional implementation, the apparatus further includes:
[0388] The motion signal acquisition module is used to obtain the motion signal of the vocal tract occlusion part of the user when speaking based on the brain wave signal; correspondingly, the voice information acquisition module is used to obtain the target voice information based on the motion signal and the user voice signal collected by the sensor.
[0389] In an optional implementation, the voice information acquisition module is used to acquire a corresponding target audio signal based on the brain wave signal;
[0390] Based on filtering, filtering out the target audio signal from the user voice signal collected by the sensor to obtain a signal to be filtered;
[0391] The signal to be filtered out is filtered out from the user voice signal collected by the sensor to obtain the target voice information.
[0392] In an optional implementation, the apparatus further includes:
[0393] The instruction information acquisition module is used to acquire instruction information corresponding to the user's voice signal based on the target voice information, where the instruction information indicates the semantic intention contained in the user's voice signal.
[0394] In an optional implementation, the voice information acquisition module is configured to obtain the target voice information through a recurrent neural network model based on the brain wave signal and the user voice signal collected by the sensor; or
[0395] According to the brain wave signal, a corresponding target audio signal is obtained; based on the target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
[0396] In an optional implementation, the target voice information includes a voiceprint feature representing a voice signal of the user.
[0397] The embodiment of the present application provides a speech signal processing device, which includes: an environmental speech acquisition module for acquiring a user's speech signal collected by a sensor; a brain wave signal acquisition module for acquiring the user's brain wave signal corresponding to when the user utters the speech; and a speech information acquisition module for obtaining target speech information based on the brain wave signal and the user's speech signal collected by the sensor. In the above manner, the brain wave signal is used as the basis for speech recognition. Since the brain wave signal does not contain the external non-user speech mixed in during complex acoustic transmission, it is less affected by other environmental noises (such as reverberation). Therefore, this part of the noise interference can be relatively well suppressed, and a better speech recognition effect can be achieved.
[0398] Reference Figure 18 , Figure 18 This application provides a structural diagram of a speech signal processing device, such as Figure 18 As shown in FIG. 1 , the apparatus 1800 includes:
[0399] The environmental voice acquisition module 1801 is used to obtain the user voice signal collected by the sensor;
[0400] a vibration signal acquisition module 1802 for acquiring a vibration signal corresponding to the user uttering the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part being a part that vibrates correspondingly based on the vocalization behavior when the user is in a vocalization state; and
[0401] The voiceprint recognition module 1803 is configured to perform voiceprint recognition based on the user voice signal and the vibration signal collected by the sensor.
[0402] In an optional implementation, the vibration signal is used to represent a vibration feature corresponding to the vibration generated by uttering speech.
[0403] In an optional implementation, the voiceprint recognition module is configured to perform voiceprint recognition based on the user voice signal collected by the sensor to obtain a first confidence level that the user voice signal collected by the sensor belongs to the user;
[0404] performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the target user;
[0405] A voiceprint recognition result is obtained according to the first confidence level and the second confidence level.
[0406] In an optional implementation, the apparatus further includes:
[0407] an electroencephalogram signal acquisition module, configured to acquire an electroencephalogram signal of the user corresponding to the user uttering the speech;
[0408] Correspondingly, the voiceprint recognition module is used to perform voiceprint recognition based on the user voice signal, the vibration signal and the brainwave signal collected by the sensor.
[0409] In an optional implementation, the voiceprint recognition module is configured to perform voiceprint recognition based on the user voice signal collected by the sensor to obtain a first confidence level that the user voice signal collected by the sensor belongs to the user;
[0410] performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the user;
[0411] performing voiceprint recognition based on the brainwave signal to obtain a third confidence level that the user voice signal collected by the sensor belongs to the user;
[0412] A voiceprint recognition result is obtained according to the first confidence level, the second confidence level, and the third confidence level.
[0413] An embodiment of the present application provides a speech signal processing device, the device comprising: an environmental speech acquisition module for acquiring a user speech signal acquired by a sensor; a vibration signal acquisition module for acquiring a vibration signal corresponding to when the user utters the speech; wherein the vibration signal is used to represent the vibration characteristics of a body part of the user; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and a voiceprint recognition module for performing voiceprint recognition based on the user speech signal and the vibration signal acquired by the sensor. In an embodiment of the present application, the vibration signal when the user speaks is used as the basis for voiceprint recognition. Since the vibration signal is less affected by other noises (such as reverberation interference, etc.), it can express the original audio characteristics of the user's speech. Therefore, the present application uses the vibration signal as the basis for voiceprint recognition, which has better recognition effect and higher reliability.
[0414] Next, an execution device provided by an embodiment of the present application is introduced, wherein the execution device may be a device with a voice interaction function or a voice input device in the above embodiment, see Figure 19 , Figure 19 This is a structural diagram of an execution device provided in an embodiment of the present application. The execution device 1900 can be specifically manifested as a mobile phone, a tablet, a laptop, a smart wearable device, a server, etc., which is not limited here. Among them, the execution device 1900 can be deployed with Figure 10 The task scheduling device described in the corresponding embodiment is used to implement Figure 10 The function of task scheduling in the corresponding embodiment. Specifically, the execution device 1900 includes: a receiver 1901, a transmitter 1902, a processor 1903 and a memory 1904 (wherein the number of the processor 1903 in the execution device 1900 can be one or more, Figure 19 (taking one processor as an example), the processor 1903 may include an application processor 19031 and a communication processor 19032. In some embodiments of the present application, the receiver 1901, the transmitter 1902, the processor 1903 and the memory 1904 may be connected via a bus or other means.
[0415] Memory 1904 may include read-only memory and random access memory, and provides instructions and data to processor 1903. A portion of memory 1904 may also include non-volatile random access memory (NVRAM). Memory 1904 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0416] Processor 1903 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0417] The methods disclosed in the above embodiments of the present application can be applied to or implemented by processor 1903. Processor 1903 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits or software instructions in processor 1903. The above processor 1903 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 1903 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in conjunction with the embodiments of the present application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 1904, and processor 1903 reads information from memory 1904 and, in conjunction with its hardware, completes the steps of the above method.
[0418] Receiver 1901 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 1902 can be used to output digital or character information through the first interface. Transmitter 1902 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 1902 can also include a display device such as a display screen.
[0419] In one embodiment of the present application, the processor 1903 is configured to execute Figure 7 、 Figure 14 as well as Figure 15 The speech signal processing method executed by the execution device in the corresponding embodiment.
[0420] The present application also provides a training device. Figure 20 , Figure 20 This is a structural diagram of a training device provided in an embodiment of the present application. Specifically, the training device 2000 is implemented by one or more servers. The training device 2000 may have relatively large differences due to different configurations or performances. It may include one or more central processing units (CPUs) 2020 (for example, one or more processors) and memory 2032, and one or more storage media 2030 (for example, one or more mass storage devices) storing application programs 2042 or data 2044. Among them, the memory 2032 and the storage medium 2030 can be temporary storage or permanent storage. The program stored in the storage medium 2030 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the training device. Furthermore, the central processing unit 2020 can be configured to communicate with the storage medium 2030 to execute a series of instruction operations in the storage medium 2030 on the training device 2000.
[0421] The training device 2000 may also include one or more power supplies 2026, one or more wired or wireless network interfaces 2050, one or more input and output interfaces 2058; or, one or more operating systems 2041, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0422] In an embodiment of the present application, the central processing unit 2020 is used to execute the steps related to the neural network model training method in the above embodiment.
[0423] An embodiment of the present application also provides a computer program product, which, when running on a computer, enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0424] A computer-readable storage medium is also provided in an embodiment of the present application, which stores a program for signal processing. When the computer-readable storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0425] The execution device, training device or terminal device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the data processing method described in the above embodiment, or so that the chip in the training device executes the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0426] For details, please refer to Figure 21 , Figure 21 This is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 2100. NPU 2100 is mounted on the host CPU as a coprocessor and is assigned tasks by the host CPU. The core of the NPU is the arithmetic circuit 2103, which is controlled by the controller 2104 to extract matrix data from the memory and perform multiplication operations.
[0427] In some implementations, the arithmetic circuit 2103 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 2103 is a two-dimensional systolic array. The arithmetic circuit 2103 may also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 2103 is a general-purpose matrix processor.
[0428] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 2102 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 2101 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 2108.
[0429] Unified memory 2106 is used to store input and output data. Weight data is directly transferred to weight memory 2102 through the Direct Memory Access Controller (DMAC) 2105. Input data is also transferred to unified memory 2106 through the DMAC.
[0430] BIU stands for Bus Interface Unit, i.e., bus interface unit 2110 , which is used for interaction between the AXI bus and DMAC and instruction fetch buffer (IFB) 2109 .
[0431] The bus interface unit 2110 (BIU) is used for the instruction fetch memory 2109 to obtain instructions from the external memory, and is also used for the storage unit access controller 2105 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0432] DMAC is mainly used to transfer input data in the external memory DDR to the unified memory 2106 or to transfer weight data to the weight memory 2102 or to transfer input data to the input memory 2101.
[0433] The vector calculation unit 2107 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 2103, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of feature planes.
[0434] In some implementations, the vector calculation unit 2107 can store the processed output vector in the unified memory 2106. For example, the vector calculation unit 2107 can apply a linear function or a nonlinear function to the output of the operation circuit 2103, such as linear interpolation of the feature plane extracted by the convolution layer, or accumulate a vector of values to generate an activation value. In some implementations, the vector calculation unit 2107 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 2103, for example, for use in subsequent layers in a neural network.
[0435] An instruction fetch buffer 2109 connected to the controller 2104 is used to store instructions used by the controller 2104;
[0436] Unified memory 2106, input memory 2101, weight memory 2102, and instruction fetch memory 2109 are all on-chip memories. External memories are private to the NPU hardware architecture.
[0437] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.
[0438] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0439] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0440] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0441] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A speech signal processing method, characterized in that: The method comprises: Obtain user voice signals collected by sensors; Obtaining a vibration signal corresponding to when the user utters the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and Obtaining target voice information according to the vibration signal and the user voice signal collected by the sensor; The method further comprises: Obtaining a brain wave signal of the user corresponding to when the user utters the voice; Accordingly, obtaining target voice information according to the vibration signal and the user voice signal collected by the sensor includes: Target voice information is obtained according to the vibration signal, the brain wave signal and the user voice signal collected by the sensor.
2. The method according to claim 1, characterized in that The vibration signal is used to represent a vibration feature corresponding to the vibration generated by the user uttering the voice.
3. The method according to claim 1, characterized in that The body part includes at least one of the following: the top of the skull, the face, the throat, or the neck.
4. The method according to claim 1, wherein The obtaining of a vibration signal corresponding to the user uttering the voice comprises: Acquire a video frame including the user; Extracting, according to the video frame, a vibration signal corresponding to when the user utters the voice.
5. The method according to claim 4, characterized in that The video frames are acquired through a dynamic vision sensor and / or a high-speed camera.
6. The method according to any one of claims 1 to 5, characterized in that: The obtaining target voice information according to the vibration signal and the user voice signal collected by the sensor includes: Acquiring a corresponding target audio signal according to the vibration signal; Based on filtering, filtering out the target audio signal from the user voice signal collected by the sensor to obtain a signal to be filtered; The signal to be filtered out is filtered out from the user voice signal collected by the sensor to obtain the target voice information.
7. The method according to any one of claims 1 to 5, characterized in that: The obtaining target voice information according to the vibration signal and the user voice signal collected by the sensor includes: Based on the vibration signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model; or According to the vibration signal, a corresponding target audio signal is obtained; based on the target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
8. The method according to claim 1, characterized in that The method further comprises: Acquiring, based on the brainwave signal, a motion signal of a vocal tract occlusion portion when the user utters a voice; correspondingly, acquiring target voice information based on the vibration signal, the brainwave signal, and the user voice signal collected by the sensor, includes: Target voice information is obtained according to the vibration signal, the motion signal, and the user voice signal collected by the sensor.
9. The method according to claim 1, characterized in that The obtaining target voice information according to the vibration signal, the brain wave signal, and the user voice signal collected by the sensor includes: Based on the vibration signal, the brain wave signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model; or, Acquire a corresponding first target audio signal according to the vibration signal; According to the brain wave signal, a corresponding second target audio signal is obtained; based on the first target audio signal, the second target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
10. The method according to any one of claims 1 to 5, characterized in that: The target voice information includes a voiceprint feature representing a voice signal of the user.
11. A speech signal processing method, characterized in that: The method comprises: Obtain user voice signals collected by sensors; Obtaining a brain wave signal of the user corresponding to when the user utters the voice; and Target voice information is obtained according to the brain wave signal and the user voice signal collected by the sensor.
12. The method according to claim 11, characterized in that The method further comprises: Acquiring, based on the brainwave signal, a motion signal of the vocal tract occlusion portion of the user when speaking; correspondingly, acquiring target voice information based on the brainwave signal and the user voice signal collected by the sensor, includes: The target voice information is obtained according to the motion signal and the user voice signal collected by the sensor.
13. The method according to claim 11 or 12, characterized in that The step of obtaining target voice information based on the brain wave signal and the user voice signal collected by the sensor includes: Acquiring a corresponding target audio signal according to the brainwave signal; Based on filtering, filtering out the target audio signal from the user voice signal collected by the sensor to obtain a signal to be filtered; The signal to be filtered out is filtered out from the user voice signal collected by the sensor to obtain the target voice information.
14. The method according to claim 12, characterized in that The step of obtaining target voice information based on the brain wave signal and the user voice signal collected by the sensor includes: Based on the brain wave signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model; or According to the brain wave signal, a corresponding target audio signal is obtained; based on the target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
15. The method according to claim 11, 12 or 14, characterized in that The target voice information includes a voiceprint feature representing a voice signal of the user.
16. A speech signal processing method, characterized in that: The method comprises: Obtain user voice signals collected by sensors; Obtaining a vibration signal corresponding to when the user utters the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part is a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and performing voiceprint recognition based on the user voice signal and the vibration signal collected by the sensor; The method further comprises: Obtaining a brain wave signal of the user corresponding to when the user utters the voice; Accordingly, performing voiceprint recognition based on the user voice signal and the vibration signal collected by the sensor includes: Voiceprint recognition is performed based on the user voice signal, the brainwave signal, and the vibration signal collected by the sensor.
17. The method according to claim 16, characterized in that The performing voiceprint recognition based on the user voice signal and the vibration signal collected by the sensor includes: performing voiceprint recognition based on the user voice signal collected by the sensor to obtain a first confidence level that the user voice signal collected by the sensor belongs to the user; performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the target user; A voiceprint recognition result is obtained according to the first confidence level and the second confidence level.
18. The method according to claim 16, characterized in that The performing voiceprint recognition based on the user voice signal, the vibration signal, and the brainwave signal collected by the sensor includes: performing voiceprint recognition based on the user voice signal collected by the sensor to obtain a first confidence level that the user voice signal collected by the sensor belongs to the user; performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the user; performing voiceprint recognition based on the brainwave signal to obtain a third confidence level that the user voice signal collected by the sensor belongs to the user; A voiceprint recognition result is obtained according to the first confidence level, the second confidence level, and the third confidence level.
19. A speech signal processing device, characterized in that: The device comprises: Environmental voice acquisition module, used to obtain user voice signals collected by sensors; a vibration signal acquisition module, configured to acquire a vibration signal corresponding to the user uttering the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part being a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and a voice information acquisition module, configured to obtain target voice information based on the vibration signal and the user voice signal collected by the sensor; The device further comprises: The brain wave signal acquisition module is used to obtain the brain wave signal of the user corresponding to when the user utters the voice; correspondingly, the voice information acquisition module is used to obtain target voice information based on the vibration signal, the brain wave signal and the user voice signal collected by the sensor.
20. The device according to claim 19, characterized in that The vibration signal is used to represent a vibration feature corresponding to the vibration generated by the user uttering the voice.
21. The device according to claim 19 or 20, characterized in that The voice information acquisition module is configured to acquire a corresponding target audio signal based on the vibration signal; and filter the target audio signal from the user voice signal collected by the sensor based on filtering to obtain a signal to be filtered out of noise; The signal to be filtered out is filtered out from the user voice signal collected by the sensor to obtain the target voice information.
22. The device according to claim 19 or 20, characterized in that The voice information acquisition module is used to obtain the target voice information through a recurrent neural network model based on the vibration signal and the user voice signal collected by the sensor; or to obtain the corresponding target audio signal according to the vibration signal; and to obtain the target voice information through a recurrent neural network model based on the target audio signal and the user voice signal collected by the sensor.
23. The device according to claim 19 or 20, characterized in that The target voice information includes a voiceprint feature representing a voice signal of the user.
24. A speech signal processing device, characterized in that: The device comprises: An environmental voice acquisition module is used to obtain the user's voice signal collected by the sensor; an electroencephalogram signal acquisition module, configured to acquire an electroencephalogram signal of the user corresponding to when the user utters the speech; and The voice information acquisition module is used to obtain target voice information based on the brain wave signal and the user voice signal collected by the sensor.
25. The device according to claim 24, characterized in that The device further comprises: The motion signal acquisition module is used to obtain the motion signal of the vocal tract occlusion part of the user when speaking based on the brain wave signal; correspondingly, the voice information acquisition module is used to obtain the target voice information based on the motion signal and the user voice signal collected by the sensor.
26. The device according to claim 24 or 25, characterized in that The voice information acquisition module is used to obtain the corresponding target audio signal based on the brain wave signal; based on filtering, filter the target audio signal from the user voice signal collected by the sensor to obtain the signal to be filtered; filter the signal to be filtered from the user voice signal collected by the sensor to obtain the target voice information.
27. The device according to claim 24 or 25, characterized in that The voice information acquisition module is configured to obtain the target voice information through a recurrent neural network model based on the brain wave signal and the user voice signal collected by the sensor; or According to the brain wave signal, a corresponding target audio signal is obtained; based on the target audio signal and the user voice signal collected by the sensor, the target voice information is obtained through a recurrent neural network model.
28. The device according to claim 24 or 25, characterized in that The target voice information includes a voiceprint feature representing a voice signal of the user.
29. A speech signal processing device, characterized in that: The device comprises: Environmental voice acquisition module, used to obtain user voice signals collected by sensors; a vibration signal acquisition module, configured to acquire a vibration signal corresponding to the user uttering the voice; wherein the vibration signal is used to represent a vibration characteristic of a body part of the user; the body part being a part that vibrates accordingly based on the vocalization behavior when the user is in a vocalization state; and a voiceprint recognition module, configured to perform voiceprint recognition based on the user voice signal and the vibration signal collected by the sensor; The device further comprises: an electroencephalogram signal acquisition module, configured to acquire an electroencephalogram signal of the user corresponding to the user uttering the speech; Correspondingly, the voiceprint recognition module is used to perform voiceprint recognition based on the user voice signal, the vibration signal and the brainwave signal collected by the sensor.
30. The device according to claim 29, characterized in that The voiceprint recognition module is configured to perform voiceprint recognition based on the user voice signal collected by the sensor, and obtain a first confidence level that the user voice signal collected by the sensor belongs to the user; performing voiceprint recognition based on the vibration signal to obtain a second confidence level that the user voice signal collected by the sensor belongs to the target user; A voiceprint recognition result is obtained according to the first confidence level and the second confidence level.
31. A system, characterized in that Including processor and memory; The memory stores program instructions, and when the program instructions stored in the memory are executed by the processor, the method according to any one of claims 1 to 18 is implemented.
32. A computer-readable storage medium, characterized in that The invention comprises a program which, when run on a computer, causes the computer to execute the method according to any one of claims 1 to 18.
33. A computer program, characterized in that When the method is run on a computer, the computer is caused to execute the method according to any one of claims 1 to 18.
Citation Information
Patent Citations
System, device, and method of voice-based user authentication utilizing a challenge
US20180232511A1
Sound processing apparatus
US20200005770A1
Hearing device configured to utilize non-audio information to process audio signals
US20200077206A1