Voice interaction method, apparatus, device, storage medium, and program product
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-09
- Publication Date
- 2026-08-11
AI Technical Summary
[0007]基于此,本发明提供一种语音交互方法的多个实施例,其中至少一个实施例用于解决穿戴设备之间无交互性,且通过穿戴设备进行语音通信时噪声较大的技术问题
[0085]This application provides a voice interaction method applied to a first wearable device. The method includes: generating a link switching request in response to a first interaction command, the link switching request indicating that the uplink audio link be switched to a first audio input source; acquiring a first audio signal and a second audio signal, the first audio signal originating from the first audio input source and the second audio signal originating from a second audio input source; performing noise reduction processing on the first audio signal and the second audio signal to generate a processed audio interaction signal; and transmitting the audio interaction signal to a main control terminal for forwarding to a remote communication end. In implementation, this method actively initiates a link switching by the first wearable device in response to the interaction command, directionally switching the uplink audio link to the first audio input source, while simultaneously integrating auxiliary signals from the second audio input source for collaborative noise reduction processing, which helps improve the clarity of uplink speech and the ability to suppress environmental noise. On the other hand, this method centralizes the collaborative processing logic of voice interaction within the first wearable device, utilizing the first wearable device to complete the link switching in response to the interaction command, and integrating multiple voice information collected from the first and second audio input sources for local noise reduction processing, thereby eliminating the need to occupy the processing resources of the main control terminal and helping to reduce the overall system power consumption and communication load. Meanwhile, the architecture based on multi-device collaborative acquisition and centralized processing helps to reduce the number of sensors such as microphones on each wearable device, thereby effectively reducing device weight, space occupation, and manufacturing costs, and significantly improving the battery life of wearable devices, ultimately achieving a balance between lightweight design and high-quality voice interaction.
Smart Images

Figure CN122551789A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent communication electronic equipment technology, and in particular to a voice interaction method, apparatus, device, storage medium, and program product. Background Technology
[0002] With the rapid development of mobile computing, sensor technology, and low-power wireless communication technology, traditional consumer electronics devices are gradually evolving from a handheld-centric model to one that naturally integrates with the human body. Against this backdrop, smart wearable devices have emerged. The fundamental driving force behind these products is to meet users' urgent needs for more convenient, immediate, and personalized information interaction and health management experiences. They aim to overcome the limitations of smartphones and other devices, which require active user operation, are handheld, and experience interruptions in information access. By being worn directly on the body, they achieve seamless, continuous, and contextualized delivery of information and services, thus becoming a key carrier connecting the digital and physical worlds and enhancing human perception and capabilities.
[0003] Currently, smart wearable devices have diversified into various product forms, serving different application scenarios. Among them, the two most representative categories are augmented reality (AR) glasses and smart rings. AR smart glasses, as head-mounted devices integrating sensing, computing, and display functions, are gradually moving from concept to practical application. Interactive scenarios based on AR glasses are also increasing, enriching people's lives and changing their lifestyles. Smart rings based on micro-sensors have also developed rapidly in recent years, and their use is gradually enriching people's daily lives.
[0004] Among related technologies, existing smart wearable device technologies, whether AR glasses that overlay environmental information or smart rings that focus on vital sign perception, have established basic technical frameworks and application models. In applications, voice interaction, as one of the fundamental human-computer interaction methods, has become an important input and control means for smart wearable devices.
[0005] However, current voice interaction methods for wearable devices have the following technical problems:
[0006] Currently, there is no interaction between wearable devices, and voice communication through wearable devices is noisy, which needs to be optimized. Summary of the Invention
[0007] Based on this, the present invention provides several embodiments of a voice interaction method, wherein at least one embodiment is used to solve the technical problem of no interaction between wearable devices and high noise when voice communication is performed through wearable devices.
[0008] In a first aspect, this application provides a voice interaction method applied to a first wearable device, the method comprising:
[0009] In response to the first interactive command, a link switching request is generated, which is used to indicate that the uplink audio link be switched to the first audio input source;
[0010] Acquire a first audio signal and a second audio signal, wherein the first audio signal originates from the first audio input source and the second audio signal originates from the second audio input source;
[0011] The first audio signal and the second audio signal are subjected to noise reduction processing to generate a processed audio interaction signal;
[0012] The audio interaction signal is transmitted to the main control terminal for forwarding to the remote communication terminal.
[0013] In one embodiment, the second audio input source is applied to a second wearable device, and the first wearable device and the second wearable device are communicatively connected. The step of performing noise reduction processing on the first audio signal and the second audio signal to generate a processed audio interaction signal includes:
[0014] The information difference between the first audio signal and the second audio signal is obtained, and the information difference represents the signal difference generated by the first audio input source and the second audio input source receiving the same sound source.
[0015] Based on the signal difference, the first audio signal and the second audio signal are subjected to noise reduction processing to generate the noise-reduced audio interaction signal.
[0016] In one embodiment, generating a link switching request in response to a first interactive instruction includes:
[0017] If not in a call state, responding to the first interactive command triggers a call request to a preset remote communication terminal;
[0018] After the call is established, the uplink audio link is switched to the first audio input source associated with the first interaction command.
[0019] In one embodiment, generating a link switching request in response to a first interactive instruction includes:
[0020] If in a call state, the uplink audio link is switched to the first audio input source in response to the second interaction command;
[0021] In response to the third interactive command, the uplink audio link is switched to the second audio input source.
[0022] In one embodiment, before performing noise reduction processing on the first audio signal and the second audio signal to generate the processed audio interaction signal, the method further includes:
[0023] The first audio signal and the second audio signal are processed based on a preset correlation algorithm to determine the time delay information between the first audio signal and the second audio signal;
[0024] Alignment compensation is performed on the first audio signal and the second audio signal based on the time delay information.
[0025] In one embodiment, the target speech is a homologous signal in both the first and second audio signals, while the noise is a non-homologous signal. The alignment compensation of the first and second audio signals based on the time delay information includes:
[0026] The cross-correlation sequence between the first audio signal and the second audio signal is calculated based on the cross-correlation function, which is:
[0027]
[0028] in, The first audio signal. This is the second audio signal;
[0029] Determine the value k that maximizes the cross-correlation sequence, and use the number of sampling points corresponding to the value k as the time delay information.
[0030] In one embodiment, calculating the cross-correlation sequence between the first audio signal and the second audio signal based on the cross-correlation function includes:
[0031] Based on a preset target audio segment, the first audio signal and the second audio signal are filtered.
[0032] The time delay information is calculated based on a preset weighted cross-correlation function, wherein the weighted cross-correlation function is:
[0033]
[0034] in, The first audio signal The Fourier transform representation, The second audio signal The conjugate representation of the Fourier transform, where IFFT represents the inverse Fourier transform.
[0035] In one embodiment, the noise reduction processing of the first audio signal and the second audio signal to generate the processed audio interaction signal includes:
[0036] Perform speech activity detection on the first audio signal to identify speech segments and non-speech segments in the first audio signal;
[0037] In the speech segment, the target audio component in the target direction is enhanced by a preset beamforming algorithm;
[0038] In the noise components of the non-speech segment or the speech segment, adaptive noise cancellation is performed based on a reference signal to obtain the denoised audio interaction signal, wherein the reference signal is determined based on the second audio signal.
[0039] In one embodiment, the method further includes:
[0040] During the transmission of the audio interaction signal, the connection status of the communication link used to receive the second audio signal is monitored;
[0041] In response to the interruption of the communication link, the uplink audio link is switched to a single-source mode that transmits only the first audio signal.
[0042] Secondly, this application also provides a voice interaction device, comprising:
[0043] An interactive switching module is used to generate a link switching request in response to a first interactive command. The link switching request is used to indicate that the uplink audio link be switched to the first audio input source.
[0044] The signal acquisition module is used to acquire a first audio signal and a second audio signal, wherein the first audio signal originates from the first audio input source and the second audio signal originates from the second audio input source.
[0045] The audio processing module is used to perform noise reduction processing on the first audio signal and the second audio signal to generate a processed audio interaction signal;
[0046] The signal transmission module is used to transmit the audio interaction signal to the main control terminal for forwarding to the remote communication terminal.
[0047] In one embodiment, the second audio input source is applied to a second wearable device, the first wearable device and the second wearable device are communicatively connected, and the audio processing module is further configured to:
[0048] The information difference between the first audio signal and the second audio signal is obtained, and the information difference represents the signal difference generated by the first audio input source and the second audio input source receiving the same sound source.
[0049] Based on the signal difference, the first audio signal and the second audio signal are subjected to noise reduction processing to generate the noise-reduced audio interaction signal.
[0050] In one embodiment, the interaction switching module is further configured to:
[0051] If not in a call state, responding to the first interactive command triggers a call request to a preset remote communication terminal;
[0052] After the call is established, the uplink audio link is switched to the first audio input source associated with the first interaction command.
[0053] In one embodiment, the interaction switching module is further configured to:
[0054] If in a call state, the uplink audio link is switched to the first audio input source in response to the second interaction command;
[0055] In response to the third interactive command, the uplink audio link is switched to the second audio input source.
[0056] In one embodiment, prior to the audio processing module, the following is also included:
[0057] The correlation algorithm module is used to process the first audio signal and the second audio signal based on a preset correlation algorithm to determine the time delay information between the first audio signal and the second audio signal;
[0058] The delay compensation module is used to perform alignment compensation on the first audio signal and the second audio signal based on the delay information.
[0059] In one embodiment, the target speech is a homologous signal in both the first and second audio signals, while the noise is a non-homologous signal. The time delay compensation module is further configured to:
[0060] The cross-correlation sequence between the first audio signal and the second audio signal is calculated based on the cross-correlation function, which is:
[0061]
[0062] in, The first audio signal. This is the second audio signal;
[0063] Determine the value k that maximizes the cross-correlation sequence, and use the number of sampling points corresponding to the value k as the time delay information.
[0064] In one embodiment, the delay compensation module is further configured to:
[0065] Based on a preset target audio segment, the first audio signal and the second audio signal are filtered.
[0066] The time delay information is calculated based on a preset weighted cross-correlation function, wherein the weighted cross-correlation function is:
[0067]
[0068] in, The first audio signal The Fourier transform representation, The second audio signal The conjugate representation of the Fourier transform, where IFFT represents the inverse Fourier transform.
[0069] In one embodiment, the audio processing module is further configured to:
[0070] Perform speech activity detection on the first audio signal to identify speech segments and non-speech segments in the first audio signal;
[0071] In the speech segment, the target audio component in the target direction is enhanced by a preset beamforming algorithm;
[0072] In the noise components of the non-speech segment or the speech segment, adaptive noise cancellation is performed based on a reference signal to obtain the denoised audio interaction signal, wherein the reference signal is determined based on the second audio signal.
[0073] In one embodiment, the device further includes:
[0074] A connection detection module is used to monitor the connection status of the communication link used to receive the second audio signal during the transmission of the audio interaction signal.
[0075] A mode switching module is used to switch the uplink audio link to a single-source mode that transmits only the first audio signal in response to the interruption of the communication link.
[0076] Thirdly, this application also provides a wearable device, including:
[0077] The audio acquisition module is used to acquire the first audio signal of the target object;
[0078] An interactive input module is used to receive interactive instructions input by the target object;
[0079] The communication module is connected to both the audio acquisition module and the interactive input module. The communication module includes a first communication unit and a second communication unit. The first communication unit is used to establish a first communication connection with the main control terminal, and the second communication unit is used to establish a second communication connection with another wearable device.
[0080] The control module is connected to the audio acquisition module, the interactive input module, and the communication module. The control module is configured to implement the steps of the method described in any one of the first aspects when executed, in order to switch the communication link and perform voice interaction.
[0081] Fourthly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the voice interaction method as described in any embodiment of the first aspect.
[0082] Fifthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the voice interaction method as described in any embodiment of the first aspect.
[0083] Sixthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the voice interaction method as described in any embodiment of the first aspect.
[0084] The aforementioned voice interaction method, apparatus, device, storage medium, and program product, derived from the technical features in the embodiments, can achieve the following beneficial effects to address the technical problems raised in the background art:
[0085] This application provides a voice interaction method applied to a first wearable device. The method includes: generating a link switching request in response to a first interaction command, the link switching request indicating that the uplink audio link be switched to a first audio input source; acquiring a first audio signal and a second audio signal, the first audio signal originating from the first audio input source and the second audio signal originating from a second audio input source; performing noise reduction processing on the first audio signal and the second audio signal to generate a processed audio interaction signal; and transmitting the audio interaction signal to a main control terminal for forwarding to a remote communication end. In implementation, this method actively initiates a link switching by the first wearable device in response to the interaction command, directionally switching the uplink audio link to the first audio input source, while simultaneously integrating auxiliary signals from the second audio input source for collaborative noise reduction processing, which helps improve the clarity of uplink speech and the ability to suppress environmental noise. On the other hand, this method centralizes the collaborative processing logic of voice interaction within the first wearable device, utilizing the first wearable device to complete the link switching in response to the interaction command, and integrating multiple voice information collected from the first and second audio input sources for local noise reduction processing, thereby eliminating the need to occupy the processing resources of the main control terminal and helping to reduce the overall system power consumption and communication load. Meanwhile, the architecture based on multi-device collaborative acquisition and centralized processing helps to reduce the number of sensors such as microphones on each wearable device, thereby effectively reducing device weight, space occupation, and manufacturing costs, and significantly improving the battery life of wearable devices, ultimately achieving a balance between lightweight design and high-quality voice interaction. Attached Figure Description
[0086] To more clearly illustrate the technical solutions in the embodiments of this application or the conventional technology, the drawings used in the description of the embodiments or the conventional technology will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0087] Figure 1 This is a schematic diagram of the application environment of a voice interaction method in one embodiment of this application;
[0088] Figure 2 This is a schematic diagram showing the connection of an application environment for a voice interaction method in one embodiment of this application;
[0089] Figure 3 This is a schematic diagram of the structure of a wearable device in one embodiment of this application;
[0090] Figure 4 This is a schematic diagram of the first process of a voice interaction method in one embodiment of this application;
[0091] Figure 5 This is a second flowchart illustrating a voice interaction method in another embodiment of this application;
[0092] Figure 6 This is a schematic diagram of the third process of a voice interaction method in another embodiment of this application;
[0093] Figure 7 This is a schematic diagram of the fourth process of a voice interaction method in another embodiment of this application;
[0094] Figure 8 This is a fifth flowchart illustrating a voice interaction method in another embodiment of this application;
[0095] Figure 9 This is a schematic diagram of the sixth process of a voice interaction method in another embodiment of this application;
[0096] Figure 10 This is a schematic diagram of the seventh process of a voice interaction method in another embodiment of this application;
[0097] Figure 11 This is a schematic diagram of the eighth process of a voice interaction method in another embodiment of this application;
[0098] Figure 12 This is a ninth flowchart illustrating a voice interaction method in another embodiment of this application;
[0099] Figure 13 This is a schematic diagram of the audio signal processing flow in a specific embodiment of this application;
[0100] Figure 14 This is a structural block diagram of a voice interaction device in one embodiment;
[0101] Figure 15 This is an internal structural diagram of a computer device in one embodiment.
[0102] Explanation of reference numerals in the attached figures: 100, first wearable device; 110, audio acquisition module; 120, interactive input module; 130, communication module; 140, control module; 150, audio playback module; 200, second wearable device; 300, main control terminal. Detailed Implementation
[0103] To facilitate understanding of this application, a more complete description will be provided below with reference to the accompanying drawings, which illustrate embodiments of the present application. However, the present application can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that the disclosure of this application will be thorough and complete.
[0104] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0105] It is understood that the terms "first," "second," etc., used herein may be used to describe various elements, but these elements are not limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of this application, a first resistor may be referred to as a second resistor, and similarly, a second resistor may be referred to as a first resistor. Both the first resistor and the second resistor are resistors, but they are not the same resistor.
[0106] It is understood that the term "connection" in the following embodiments should be understood as "electrical connection," "communication connection," etc., if the connected circuits, modules, units, etc., have electrical signal or data transmission with each other.
[0107] It is understandable that "at least one" refers to one or more, and "multiple" refers to two or more. "At least a part of an element" refers to part or all of an element.
[0108] When used herein, the singular forms of “a,” “an,” and “the” may also include the plural forms unless the context clearly indicates otherwise. It should also be understood that the terms “comprising / including” or “having,” etc., specify the presence of the stated features, wholes, steps, operations, components, parts, or combinations thereof, but do not preclude the possibility of the presence or addition of one or more other features, wholes, steps, operations, components, parts, or combinations thereof. Meanwhile, the term “and / or” as used in this specification includes any and all combinations of the associated listed items.
[0109] At least one embodiment of this application was developed by the inventor based on his understanding and research into the following issues:
[0110] Among related technologies, existing smart wearable device technologies, whether AR glasses that overlay environmental information or smart rings that focus on vital sign perception, have established basic technical frameworks and application models. In applications, voice interaction, as one of the fundamental human-computer interaction methods, has become an important input and control means for smart wearable devices.
[0111] However, current voice interaction methods for wearable devices have the following technical problems:
[0112] Currently, there is no interaction between wearable devices, and voice communication through wearable devices is noisy, which needs to be optimized.
[0113] Based on the above problems, this application provides several embodiments of a voice interaction method, wherein at least one embodiment is used to solve the technical problem of no interaction between wearable devices and high noise when voice communication is performed through wearable devices.
[0114] The voice interaction method provided in this application embodiment can be applied to, for example... Figure 1 and Figure 2 The application environment shown. Figure 1 and Figure 2 The image shows a voice interaction collaboration system based on multiple wearable devices, including a first wearable device, a second wearable device, and a main control terminal.
[0115] The first wearable device is connected to the second wearable device.
[0116] The main control terminal is connected to both the first wearable device and the second wearable device.
[0117] For example, the main control terminal is provided with a bidirectional communication link between itself and both the first wearable device and the second wearable device. The bidirectional communication link may include a control link and an audio transmission link.
[0118] For example, it can be as follows Figure 3 As shown, the first wearable device or the second wearable device may include an audio acquisition module, an interactive input module, and a communication module.
[0119] The audio acquisition module is used to acquire the first audio signal of the target object.
[0120] For example, the audio acquisition module can refer to the hardware component in a wearable device used to capture sound wave signals and convert them into electrical signals, and may include a microphone and related peripheral circuitry.
[0121] The interactive input module is used to receive interactive instructions input by the target object.
[0122] For example, the interactive input module can refer to a hardware unit in a wearable device that receives user operation commands, converts physical or touch operations into electrical signals, and transmits them to the control module. Interactive commands can refer to control commands triggered by the user through the interactive input module, including single click, double click, long press, swipe, etc., used to start, switch, or terminate the voice interaction process.
[0123] The communication module is connected to the audio acquisition module and the interactive input module. The communication module includes a first communication unit and a second communication unit. The first communication unit is used to establish a first communication connection with the main control terminal, and the second communication unit is used to establish a second communication connection with another wearable device.
[0124] For example, the communication module can refer to a hardware unit in the wearable device responsible for wireless data transmission and reception. The communication module can support data transmission and control command interaction between the wearable device and the main control terminal and other wearable devices. The first communication unit can refer to a subsystem for establishing a communication connection with the main control terminal (such as a smartphone), and can be implemented based on Bluetooth / BLE or Wi-Fi protocols. The second communication unit can refer to a subsystem for establishing point-to-point communication with other wearable devices, and can be implemented based on Wi-Fi P2P, UWB, or proprietary wireless protocols.
[0125] The control module is connected to the audio acquisition module, the interactive input module, and the communication module. The control module is configured to respond to interactive commands, switch communication links, and perform voice interaction.
[0126] Specifically, wearable devices can refer to intelligent electronic devices that can be worn on a certain part of the human body (such as the hand, head, ear, wrist, etc.). Wearable devices have the attribute of being wearable and can be worn on a certain part of the human body, such as the hand, head, ear, wrist, etc., and their form can include smart rings, AR glasses, etc. In this embodiment, we can take a smart ring as the first wearable device, AR glasses as the second wearable device, and a mobile phone as the main control terminal as an example for explanation. Other cases are similar and will not be described in detail here.
[0127] In one specific exemplary embodiment, the smart ring may integrate a miniature microphone (MIC) for collecting ambient sound. It may also integrate a miniature button for initiating or discontinuing the voice interaction process. The smart ring may further integrate an MCU control module, Bluetooth (Bluetooth Classic) / BLE (Bluetooth Low Energy) modules, and WiFi modules for related interactive control and BT / BLE communication / audio stream transmission. The AR glasses may integrate a speaker (SPK) and a microphone (MIC) for picking up ambient sound or playing related audio. In addition, the AR glasses may integrate an MCU control module, BT / BLE modules, and WiFi modules for related control and BT / BLE communication / audio stream transmission. The smart ring and AR glasses can connect to a mobile phone via BT / BLE. BLE is used to transmit control information, and BT is used to transmit audio data. When the user presses the button on the smart ring to issue a related interactive command, the mobile app will switch the audio / video link based on a pre-set program to achieve the corresponding voice interaction function. In this way, in scenarios where private calls are made using a smart ring, the microphone on the smart ring can collect relatively clear and clean original speech, while the microphone on the AR glasses can pick up human voices and environmental noise, and transmit the picked-up sound to the smart ring via P2P. The smart ring then integrates and processes the sound collected by the AR glasses and its own microphone, thereby achieving call noise reduction.
[0128] In one embodiment, such as Figure 4 As shown, a voice interaction collaboration method based on multiple wearable devices is provided, which can be applied to... Figure 1 Taking the first wearable device in the example, the explanation includes the following steps:
[0129] Step 402: In response to the first interactive instruction, generate a link switching request, which is used to indicate that the uplink audio link be switched to the first audio input source.
[0130] Interactive commands can refer to control commands triggered by the user through the interactive input module, including single click, double click, long press, swipe, etc., used to start, switch, or terminate the voice interaction process. A first interactive command can refer to an interactive action corresponding to a specific control command; in this embodiment, the first interactive command corresponds to a link switching request.
[0131] The uplink audio link can refer to the signal path through which audio signals are transmitted from the first wearable device to the remote end of the call via the main control terminal during a voice call.
[0132] The first audio input source can refer to an audio acquisition unit associated with the first wearable device, such as a microphone. In one application scenario, the first audio input source is located close to the user's mouth, enabling the acquisition of high signal-to-noise ratio speech signals.
[0133] For example, in a call scenario, when a user presses a button on the first wearable device, the first wearable device can notify the master terminal of the button event via BLE (Bluetooth Low Energy). After receiving the button event from the first wearable device, the master terminal can initiate the corresponding processing flow. In a non-call scenario, if the user clicks a button on the first wearable device, the first wearable device synchronizes the button event to the master terminal via BLE. After receiving the button event, the master terminal can actively initiate a call. In this way, communication link switching can be performed based on the triggered processing flow.
[0134] In one application scenario, let's take a smart ring as the first wearable device and AR glasses as the second wearable device as an example. The user wears both the smart ring and the AR glasses, initially making a hands-free call through the AR glasses. Due to a noisy environment, the user wants to switch to a private call mode to avoid being overheard. At this point, the user can trigger the first interaction command by clicking a touch button on the surface of the smart ring. After detecting the click event, the smart ring generates a link switching request, switching the uplink audio link from the AR glasses' microphone to the smart ring's microphone.
[0135] Step 404: Acquire a first audio signal and a second audio signal, wherein the first audio signal originates from the first audio input source and the second audio signal originates from the second audio input source.
[0136] The first audio signal may refer to the raw audio data acquired by the audio acquisition unit built into the first wearable device. The second audio signal may refer to the raw audio data acquired by the audio acquisition unit built into the second wearable device.
[0137] For example, the first wearable device can acquire a first audio signal and a second audio signal (the second audio signal is transmitted to the first wearable device from a second audio input source). The first audio signal can provide a relatively clean speech reference, reducing speech distortion after noise reduction, while the second audio signal can provide a noisy environmental reference, thereby assisting in the extraction of environmental noise features. The first and second audio signals are complementary, thus providing the basic input for subsequent noise reduction algorithms based on information difference.
[0138] Step 406: Perform noise reduction processing on the first audio signal and the second audio signal to generate a processed audio interaction signal.
[0139] Among them, noise reduction processing can refer to the process of processing the acquired audio signal to suppress or eliminate environmental noise components and retain or enhance the target speech components.
[0140] For example, the first wearable device can acquire a first audio signal and a second audio signal, and achieve noise suppression based on the information difference between the two audio signals to obtain a processed audio interaction signal.
[0141] Step 408: Transmit the audio interaction signal to the main control terminal for forwarding to the remote communication terminal.
[0142] In this context, the remote end of communication can refer to the user or device at the other end during a call, that is, the receiving device at the other end that is conducting voice communication with the current user.
[0143] By implementing the above-described voice interaction collaboration method based on multiple wearable devices, the following beneficial effects can be achieved:
[0144] This application provides a voice interaction method applied to a first wearable device. The method includes: generating a link switching request in response to a first interaction command, the link switching request indicating that the uplink audio link be switched to a first audio input source; acquiring a first audio signal and a second audio signal, the first audio signal originating from the first audio input source and the second audio signal originating from a second audio input source; performing noise reduction processing on the first audio signal and the second audio signal to generate a processed audio interaction signal; and transmitting the audio interaction signal to a main control terminal for forwarding to a remote communication end. In implementation, this method actively initiates a link switching by the first wearable device in response to the interaction command, directionally switching the uplink audio link to the first audio input source, while simultaneously integrating auxiliary signals from the second audio input source for collaborative noise reduction processing, which helps improve the clarity of uplink speech and the ability to suppress environmental noise. On the other hand, this method centralizes the collaborative processing logic of voice interaction within the first wearable device, utilizing the first wearable device to complete the link switching in response to the interaction command, and integrating multiple voice information collected from the first and second audio input sources for local noise reduction processing, thereby eliminating the need to occupy the processing resources of the main control terminal and helping to reduce the overall system power consumption and communication load. Meanwhile, the architecture based on multi-device collaborative acquisition and centralized processing helps to reduce the number of sensors such as microphones on each wearable device, thereby effectively reducing device weight, space occupation, and manufacturing costs, and significantly improving the battery life of wearable devices, ultimately achieving a balance between lightweight design and high-quality voice interaction.
[0145] In one embodiment, it can be as follows Figure 5As shown, the second audio input source is applied to the second wearable device, and the first wearable device and the second wearable device are communicatively connected. Step 406 includes:
[0146] Step 502: Obtain the information difference between the first audio signal and the second audio signal shown.
[0147] The information difference characterizes the signal difference generated when the first audio input source and the second audio input source receive the same sound source. The information difference can refer to the physical differences between the first and second audio signals due to different acquisition locations, including the time difference of sound waves arriving at different microphones, the amplitude difference of energy attenuation caused by different sound wave propagation distances, and the phase difference. The information difference can reflect the different distribution characteristics of target speech and noise in the two signals.
[0148] Step 504: Based on the signal difference, perform noise reduction processing on the first audio signal and the second audio signal to generate the noise-reduced audio interaction signal.
[0149] In this embodiment, information difference is extracted by analyzing the time, amplitude or phase difference between the two audio signals. Based on the information difference, adaptive noise reduction is performed with the first audio signal with a high signal-to-noise ratio as the main signal and the second audio signal as the noise reference. This fully utilizes the physical information difference generated by the spatial separation of the two devices to achieve accurate noise suppression, thereby improving the noise reduction effect while reducing the dependence on the number of microphones on a single device.
[0150] In one embodiment, such as Figure 6 As shown, step 402 includes:
[0151] Step 602: If not in a call state, respond to the first interactive instruction and trigger a call request to a preset communication remote end.
[0152] Step 604: After the call is established, switch the uplink audio link to the first audio input source associated with the first interaction command.
[0153] In this embodiment, in the absence of a call, the system automatically initiates a call to a preset remote device in response to the first interactive command. After the call is established, the uplink is seamlessly switched to the first audio input source, thereby realizing one-click automatic connection from the absence of a call to the private call mode. Users do not need to manually dial or switch audio channels, which significantly simplifies the operation process and improves the call initiation efficiency in multi-device collaborative scenarios.
[0154] In one embodiment, such as Figure 7 As shown, step 402 includes:
[0155] Step 702: If in a call state, in response to the second interaction command, switch the uplink audio link to the first audio input source.
[0156] Step 704: In response to the third interactive command, switch the uplink audio link to the second audio input source.
[0157] In this embodiment, during a call, the uplink is switched to the first audio input source in response to the second interactive command, enabling a private call mode; in response to the third interactive command, the uplink is switched back to the second audio input source, restoring hands-free mode. This helps users dynamically switch audio input sources as needed during calls, flexibly adapting to private and hands-free scenarios, and improving application flexibility.
[0158] In one embodiment, such as Figure 8 As shown, before step 406, the procedure further includes:
[0159] Step 802: Process the first audio signal and the second audio signal based on a preset correlation algorithm to determine the time delay information between the first audio signal and the second audio signal.
[0160] Correlation algorithms, in this context, refer to mathematical methods used to measure the change in similarity between two signals over a relative time offset. Correlation algorithms calculate the degree of correlation between one signal and another at different time delays, finding the time offset that makes the two signals most similar. Time delay information refers to the time difference between the arrival of the same sound source at two different acquisition locations; this information can be expressed in units of the number of sampling points or milliseconds.
[0161] For example, the first wearable device can process the first audio signal and the second audio signal based on a preset correlation algorithm to determine the time delay information between the first audio signal and the second audio signal.
[0162] Step 804: Perform alignment compensation on the first audio signal and the second audio signal based on the time delay information.
[0163] Alignment compensation refers to the process of time-shifting one audio signal based on pre-obtained time delay information, so that the target speech from the same source in the two signals approximates overlap on the time axis. Alignment compensation can be achieved through integer sample point shifting or sub-sampling precision interpolation.
[0164] For example, the first wearable device may perform alignment compensation on the first audio signal and the second audio signal based on the time delay information.
[0165] In this embodiment, the time delay information of the two audio signals is calculated based on the correlation algorithm, and alignment compensation is performed based on the time delay. This helps to eliminate the time difference of the target speech caused by different acquisition positions, so that the same source speech acquired by the two devices is accurately synchronized on the time axis. This provides an alignment basis for subsequent adaptive filtering with the second audio signal as a noise reference, avoids speech cancellation, and ensures the noise reduction effect.
[0166] In one embodiment, such as Figure 9 As shown, the target speech is a homologous signal in both the first and second audio signals, while the noise is a non-homologous signal. Step 804 includes:
[0167] Step 902: Calculate the cross-correlation sequence between the first audio signal and the second audio signal based on the cross-correlation function, wherein the cross-correlation function is:
[0168]
[0169] in, The first audio signal. This is the second audio signal.
[0170] The cross-correlation function refers to a mathematical function used to measure the similarity between two discrete signals at different relative time offsets. The cross-correlation sequence refers to the result of calculating the cross-correlation function in the discrete time domain; it is a numerical sequence with time offset k as the independent variable and correlation as the dependent variable. The cross-correlation sequence can be obtained by progressively sliding one signal relative to another and calculating the sum of the products at each sliding position.
[0171] For example, the first wearable device may calculate the cross-correlation sequence between the first audio signal and the second audio signal based on a cross-correlation function.
[0172] Step 904: Determine the k value that makes the cross-correlation sequence reach its maximum value, and use the number of sampling points corresponding to the k value as the time delay information.
[0173] Here, the value of k refers to the sliding offset of the second audio signal relative to the first audio signal in the cross-correlation sequence, and the unit is the number of sampling points. A positive k indicates that the second signal lags behind the first signal, and a negative k indicates that the second signal leads the first signal. Each k value can correspond to a correlation value, and the sequence length is determined by the signal length and the sliding range.
[0174] For example, the first wearable device can determine a value k that maximizes the cross-correlation sequence. This value k can represent the relative offset of the two signals with the highest correlation, i.e., the time difference of the target speech. In this case, the time delay information can be expressed as follows:
[0175]
[0176] In this embodiment, the cross-correlation sequence of the two signals is calculated based on the cross-correlation function, and the sampling point-level time delay information is obtained by locating the k value corresponding to the peak. This embodiment utilizes the high correlation between the target speech and the two devices to accurately estimate the sound wave propagation time difference based on a clear mathematical principle. The calculation is simple and the physical meaning is clear, which helps to provide a quantifiable time delay benchmark for subsequent signal alignment and ensures the effectiveness of dual-device collaborative noise reduction.
[0177] In one embodiment, such as Figure 10 As shown, step 902 includes:
[0178] Step 1002: Based on the preset target audio segment, filter the first audio signal and the second audio signal.
[0179] The preset target speech judgment can refer to the frequency band corresponding to the speech information to be collected. In this embodiment, it can refer to the preset frequency range of the main energy distribution of human speech, such as 300Hz to 3400Hz. The specific frequency range can be determined by the technicians according to the actual use needs, which will not be elaborated here.
[0180] Filtering refers to the process of selecting the frequency of an audio signal using digital filters, such as bandpass filters and bandlimited filters. Filtering allows signal components within a preset frequency band to pass through while attenuating or blocking interference components outside the band. Typical digital filters can be implemented based on IIR filters, FIR filters, or FFT frequency domain filtering.
[0181] Step 1004: Calculate the time delay information based on a preset weighted cross-correlation function, wherein the weighted cross-correlation function is:
[0182]
[0183] in, The first audio signal The Fourier transform representation, The second audio signal The conjugate representation of the Fourier transform, where IFFT represents the inverse Fourier transform.
[0184] The weighted cross-correlation function refers to an improved algorithm that optimizes time delay estimation performance by weighting the frequency domain based on the basic cross-correlation function. The weighted cross-correlation function can enhance the contribution of high signal-to-noise ratio (SNR) frequency components in the signal and suppress interference from low SNR components, thereby improving the robustness and peak sharpness of time delay estimation.
[0185] For example, the use of the GCC-PHAT (Generalized Cross-Correlation-Phase Transform) function can be used as an example. GCC-PHAT can eliminate the influence of amplitude information by normalizing the cross-power spectrum, retaining only the phase information. GCC-PHAT helps to achieve excellent time delay estimation performance in low signal-to-noise ratio and reverberation environments.
[0186] In this embodiment, filtering the audio segments preserves the effective components and removes high- and low-frequency interference, which helps improve signal quality. By using a weighted cross-correlation function to calculate time delay information, peak values are sharpened and noise interference is suppressed in low signal-to-noise ratio environments. This embodiment helps enhance the noise robustness of time delay estimation, ensuring accurate alignment of dual-channel speech in noisy environments and guaranteeing the effectiveness of noise reduction processing.
[0187] In one embodiment, such as Figure 11 As shown, step 406 includes:
[0188] Step 1102: Perform speech activity detection on the first audio signal to identify speech segments and non-speech segments in the first audio signal.
[0189] Speech activity detection refers to algorithms used to determine whether human speech exists in an audio signal. In one embodiment, speech activity detection can divide the signal into speech segments and non-speech segments by analyzing parameters such as short-time energy, zero-crossing rate, and spectral characteristics.
[0190] The speech segment refers to the time interval determined by speech activity detection to contain user speech activity. Within this interval, the signal mainly consists of target speech and environmental noise. The non-speech segment refers to the time interval determined by speech activity detection to not contain user speech activity. Within this interval, the signal mainly consists of environmental noise, which can be used for noise characteristic modeling.
[0191] For example, the first wearable module can perform voice activity detection on the first audio signal and identify voice segments and non-voice segments in the first audio signal.
[0192] Step 1104: In the speech segment, the target audio component in the target direction is enhanced by a preset beamforming algorithm.
[0193] Beamforming can refer to a spatial filtering method that uses the phase difference between signals collected by multiple microphones to form a main lobe of a beam in the target direction to improve the signal gain in that direction, while simultaneously forming nulls in the interference direction to suppress interference.
[0194] Step 1106: In the noise component of the non-speech segment or the speech segment, adaptive noise cancellation is performed based on a reference signal to obtain the denoised audio interaction signal, wherein the reference signal is determined based on the second audio signal.
[0195] For example, the reference signal may refer to an auxiliary signal used for noise cancellation, which is related to the noise component in the main signal but not to the target speech component.
[0196] In this embodiment, a segmented noise suppression strategy is used to enhance the target human voice in the speech segment and adaptively cancel it using a reference signal in the non-speech or noise segment. This helps to improve the targeting and effectiveness of noise reduction and also helps to enhance noise suppression capabilities while avoiding speech distortion, so that the system can still provide a clear and natural voice interaction experience in complex acoustic environments.
[0197] In one embodiment, such as Figure 12 As shown, the method further includes:
[0198] Step 1202: During the transmission of the audio interaction signal, monitor the connection status of the communication link used to receive the second audio signal.
[0199] Step 1204: In response to the interruption of the communication link, switch the uplink audio link to a single-source mode that transmits only the first audio signal.
[0200] In this embodiment, the connection status of the second audio signal transmission link is monitored in real time, and the system automatically switches to a single-source transmission mode using only the first audio signal when the link is interrupted. This embodiment provides a degradation fallback mechanism when collaborative noise reduction fails, ensuring uninterrupted calls while reducing device power consumption and enhancing the system robustness and reliability in multi-device collaborative scenarios.
[0201] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0202] Based on the same inventive concept, this application also provides a voice interaction device for implementing the voice interaction method described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations of one or more embodiments of the voice interaction device provided below can be found in the limitations of the voice interaction method described above, and will not be repeated here.
[0203] In one embodiment, such as Figure 14 As shown, a voice interaction device is provided, including: an interaction switching module, a signal acquisition module, an audio processing module, and a signal transmission module, wherein:
[0204] An interactive switching module is used to generate a link switching request in response to a first interactive command. The link switching request is used to indicate that the uplink audio link be switched to the first audio input source.
[0205] The signal acquisition module is used to acquire a first audio signal and a second audio signal, wherein the first audio signal originates from the first audio input source and the second audio signal originates from the second audio input source.
[0206] The audio processing module is used to perform noise reduction processing on the first audio signal and the second audio signal to generate a processed audio interaction signal;
[0207] The signal transmission module is used to transmit the audio interaction signal to the main control terminal for forwarding to the remote communication terminal.
[0208] In one embodiment, the second audio input source is applied to a second wearable device, the first wearable device and the second wearable device are communicatively connected, and the audio processing module is further configured to:
[0209] The information difference between the first audio signal and the second audio signal is obtained, and the information difference represents the signal difference generated by the first audio input source and the second audio input source receiving the same sound source.
[0210] Based on the signal difference, the first audio signal and the second audio signal are subjected to noise reduction processing to generate the noise-reduced audio interaction signal.
[0211] In one embodiment, the interaction switching module is further configured to:
[0212] If not in a call state, responding to the first interactive command triggers a call request to a preset remote communication terminal;
[0213] After the call is established, the uplink audio link is switched to the first audio input source associated with the first interaction command.
[0214] In one embodiment, the interaction switching module is further configured to:
[0215] If in a call state, the uplink audio link is switched to the first audio input source in response to the second interaction command;
[0216] In response to the third interactive command, the uplink audio link is switched to the second audio input source.
[0217] In one embodiment, prior to the audio processing module, the following is also included:
[0218] The correlation algorithm module is used to process the first audio signal and the second audio signal based on a preset correlation algorithm to determine the time delay information between the first audio signal and the second audio signal;
[0219] The delay compensation module is used to perform alignment compensation on the first audio signal and the second audio signal based on the delay information.
[0220] In one embodiment, the target speech is a homologous signal in both the first and second audio signals, while the noise is a non-homologous signal. The time delay compensation module is further configured to:
[0221] The cross-correlation sequence between the first audio signal and the second audio signal is calculated based on the cross-correlation function, which is:
[0222]
[0223] in, The first audio signal. This is the second audio signal;
[0224] Determine the value k that maximizes the cross-correlation sequence, and use the number of sampling points corresponding to the value k as the time delay information.
[0225] In one embodiment, the delay compensation module is further configured to:
[0226] Based on a preset target audio segment, the first audio signal and the second audio signal are filtered.
[0227] The time delay information is calculated based on a preset weighted cross-correlation function, wherein the weighted cross-correlation function is:
[0228]
[0229] in, The first audio signal The Fourier transform representation, The second audio signal The conjugate representation of the Fourier transform, where IFFT represents the inverse Fourier transform.
[0230] In one embodiment, the audio processing module is further configured to:
[0231] Perform speech activity detection on the first audio signal to identify speech segments and non-speech segments in the first audio signal;
[0232] In the speech segment, the target audio component in the target direction is enhanced by a preset beamforming algorithm;
[0233] In the noise components of the non-speech segment or the speech segment, adaptive noise cancellation is performed based on a reference signal to obtain the denoised audio interaction signal, wherein the reference signal is determined based on the second audio signal.
[0234] In one embodiment, the device further includes:
[0235] A connection detection module is used to monitor the connection status of the communication link used to receive the second audio signal during the transmission of the audio interaction signal.
[0236] A mode switching module is used to switch the uplink audio link to a single-source mode that transmits only the first audio signal in response to the interruption of the communication link.
[0237] The various modules in the aforementioned voice interaction device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0238] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 15As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a voice interaction method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.
[0239] Those skilled in the art will understand that Figure 15 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0240] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0241] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0242] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0243] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0244] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0245] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0246] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A voice interaction method, characterized in that, Applied to a first wearable device, the method includes: In response to the first interactive command, a link switching request is generated, which is used to indicate that the uplink audio link be switched to the first audio input source; Acquire a first audio signal and a second audio signal, wherein the first audio signal originates from the first audio input source and the second audio signal originates from the second audio input source; The first audio signal and the second audio signal are subjected to noise reduction processing to generate a processed audio interaction signal; The audio interaction signal is transmitted to the main control terminal for forwarding to the remote communication terminal.
2. The voice interaction method of claim 1, wherein, The second audio input source is applied to the second wearable device, and the first wearable device and the second wearable device are communicatively connected. The step of performing noise reduction processing on the first audio signal and the second audio signal to generate the processed audio interaction signal includes: The information difference between the first audio signal and the second audio signal is obtained, and the information difference represents the signal difference generated by the first audio input source and the second audio input source receiving the same sound source. Based on the signal difference, the first audio signal and the second audio signal are subjected to noise reduction processing to generate the noise-reduced audio interaction signal.
3. The voice interaction method of claim 1, wherein, The step of generating a link switching request in response to the first interactive command includes: If not in a call state, responding to the first interactive command triggers a call request to a preset remote communication terminal; After the call is established, the uplink audio link is switched to the first audio input source associated with the first interaction command.
4. The voice interaction method of claim 1, wherein, The step of generating a link switching request in response to the first interactive command includes: If in a call state, the uplink audio link is switched to the first audio input source in response to the second interaction command; In response to the third interactive command, the uplink audio link is switched to the second audio input source.
5. The voice interaction method of claim 1, wherein, Before performing noise reduction processing on the first audio signal and the second audio signal to generate the processed audio interaction signal, the method further includes: The first audio signal and the second audio signal are processed based on a preset correlation algorithm to determine the time delay information between the first audio signal and the second audio signal; Alignment compensation is performed on the first audio signal and the second audio signal based on the time delay information.
6. The voice interaction method of claim 5, wherein, The target speech is a homologous signal in both the first and second audio signals, while the noise is a non-homologous signal. The alignment compensation of the first and second audio signals based on the time delay information includes: The cross-correlation sequence between the first audio signal and the second audio signal is calculated based on the cross-correlation function, which is: wherein, is the first audio signal, is the second audio signal; Determine the value k that maximizes the cross-correlation sequence, and use the number of sampling points corresponding to the value k as the time delay information.
7. The voice interaction method of claim 6, wherein, The step of calculating the cross-correlation sequence between the first audio signal and the second audio signal based on the cross-correlation function includes: Based on a preset target audio segment, the first audio signal and the second audio signal are filtered. The time delay information is calculated based on a preset weighted cross-correlation function, wherein the weighted cross-correlation function is: in, The first audio signal The Fourier transform representation, The second audio signal The conjugate representation of the Fourier transform, where IFFT represents the inverse Fourier transform.
8. The voice interaction method according to claim 1, characterized in that, The noise reduction process performed on the first audio signal and the second audio signal to generate the processed audio interaction signal includes: Perform speech activity detection on the first audio signal to identify speech segments and non-speech segments in the first audio signal; In the speech segment, the target audio component in the target direction is enhanced by a preset beamforming algorithm; In the noise components of the non-speech segment or the speech segment, adaptive noise cancellation is performed based on a reference signal to obtain the denoised audio interaction signal, wherein the reference signal is determined based on the second audio signal.
9. The voice interaction method of claim 1, wherein, The method further includes: During the transmission of the audio interaction signal, the connection status of the communication link used to receive the second audio signal is monitored; In response to the interruption of the communication link, the uplink audio link is switched to a single-source mode that transmits only the first audio signal.
10. A wearable device, comprising: include: The audio acquisition module is used to acquire the first audio signal of the target object; An interactive input module is used to receive interactive instructions input by the target object; The communication module is connected to both the audio acquisition module and the interactive input module. The communication module includes a first communication unit and a second communication unit. The first communication unit is used to establish a first communication connection with the main control terminal, and the second communication unit is used to establish a second communication connection with another wearable device. The control module is connected to the audio acquisition module, the interactive input module, and the communication module. The control module is configured to implement the steps of the method according to any one of claims 1 to 9 when executed, so as to switch the communication link and perform voice interaction. 11.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-10 when the computer program is executed by the processor. When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 9.
12. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.
13. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 9.