Multi-stream processing of single-stream data

By detecting single-stream data in the processor of a portable communication device and generating multi-stream augmented data, processing multi-stream augmented data to generate multiple output channels, and then reducing these channels to generate single-stream output data, the problem of neural network performance improvements is solved, and signal processing performance improvements in low-power real-time applications are achieved.

CN120129906APending Publication Date: 2025-06-10QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380074778.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-31
Filing Date
2023-10-25
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In the prior art, the performance improvement of neural networks is limited by memory bandwidth and power constraints, especially in portable communication devices, which are difficult to improve signal processing performance in low-power real-time applications.

Method used

By detecting single stream data in the processor and generating multi-stream enhanced data, the multi-stream enhanced data is processed to generate multiple output channels and finally to reduce multiple output channels to generate single-stream output data.

Benefits of technology

This method improves the signal processing performance of the neural network without increasing the number of network weights, enhances device performance and user experience, especially on portable communication devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120129906A_ABST
    Figure CN120129906A_ABST
Patent Text Reader

Abstract

An apparatus includes one or more processors configured to detect single-stream data and generate multi-stream enhanced data including one or more modified versions of the single-stream data. The one or more processors are configured to process the multi-stream enhanced data to generate a plurality of output channels. The one or more processors are further configured to reduce the plurality of output channels to generate single-stream output data.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - Reference to Related Applications

[0002] This application claims the benefit of priority of co - owned Greek provisional patent application No. 20220100876, filed on October 31, 2022, the content of which is hereby incorporated by reference in its entirety. Technical Field

[0003] This disclosure generally relates to processing data streams. Background Art

[0004] Advances in technology have led to smaller and more powerful computing devices and an increase in the availability and consumption of media. For example, there are currently various portable personal computing devices, including wireless telephones such as mobile phones and smart phones, tablet computers, and laptop computers, which are small in size, light in weight, easy for users to carry, and capable of generating and consuming media content almost anywhere.

[0005] Advances in signal processing have led to improvements in applications that use input signals, such as voice call applications that can provide audio voice enhancement and noise reduction for input voice signals. In particular, compared to conventional techniques, signal processing using neural networks can provide enhanced performance. Improving the performance of such neural networks is typically achieved by increasing the size of the neural network, which requires the use of additional weight coefficients. However, in practice, the performance of neural networks is often limited by the amount of memory bandwidth available to transfer the weight coefficients from memory to the computing hardware used to execute the neural network. For example, transferring the weight coefficients may require more power than performing the computations using the weight coefficients. Improving the signal processing performance of neural networks in view of such memory bandwidth and power constraints associated with the transfer of weight coefficients would enhance device performance and the user experience, especially for low - power real - time applications on portable communication devices. Summary of the Invention

[0006] According to a particular aspect, a device includes a memory configured to store instructions. The device further includes one or more processors configured to detect single - stream data and generate multi - stream enhanced data including one or more modified versions of the single - stream data. The one or more processors are configured to process the multi - stream enhanced data to generate multiple output channels. The one or more processors are further configured to reduce the multiple output channels to produce single - stream output data.

[0007] According to a particular aspect, a method includes detecting single - stream data at one or more processors. The method includes generating multi - stream enhanced data including one or more modified versions of the single - stream data. The method includes processing the multi - stream enhanced data to generate multiple output channels. The method further includes reducing the multiple output channels to produce single - stream output data.

[0008] According to certain aspects, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to detect single-stream data. The instructions, when executed by the one or more processors, cause the one or more processors to generate multi-stream enhanced data including one or more modified versions of the single-stream data. The instructions, when executed by the one or more processors, cause the one or more processors to process the multi-stream enhanced data to generate multiple output channels. The instructions, when executed by the one or more processors, also cause the one or more processors to reduce the multiple output channels to produce single-stream output data.

[0009] According to certain aspects, an apparatus includes components for generating multi-stream enhanced data including one or more modified versions of single-stream data. The apparatus includes components for processing the multi-stream enhanced data to generate multiple output channels. The apparatus also includes components for reducing the multiple output channels to produce single-stream output data.

[0010] After reviewing the entire application, other aspects, advantages, and features of the present disclosure will become apparent, including the following sections: Brief Description of the Drawings, Detailed Description, and Claims. Brief Description of the Drawings

[0011] Figure 1 is a block diagram of certain illustrative aspects of a system operable to perform multi-stream processing of single-stream data in accordance with some examples of the present disclosure.

[0012] Figure 2 is in accordance with some examples of the present disclosure Figure 1 of a particular aspect of a system.

[0013] Figure 3 is in accordance with some examples of the present disclosure Figure 1 of a particular aspect of a system.

[0014] Figure 4 is in accordance with some examples of the present disclosure Figure 1 of a particular aspect of a system.

[0015] Figure 5 is in accordance with some examples of the present disclosure Figure 1 of a particular aspect of a system.

[0016] Figure 6 is in accordance with some examples of the present disclosure Figure 1 of a particular aspect of a system.

[0017] Figure 7 is in accordance with some examples of the present disclosure Figure 1 of a particular aspect of a system.

[0018] Figure 8 is a diagram of a specific aspect of a system in accordance with some examples of the present disclosure Figure 1 of the system

[0019] Figure 9 is a diagram of a specific aspect of a system in accordance with some examples of the present disclosure Figure 1 of the system

[0020] Figure 10 is a diagram of a specific aspect of a system in accordance with some examples of the present disclosure Figure 1 of the system

[0021] Figure 11 is a diagram of a specific aspect of a system in accordance with some examples of the present disclosure Figure 1 of the system

[0022] Figure 12 is a diagram of a specific aspect of a system in accordance with some examples of the present disclosure Figure 1 of the system

[0023] Figure 13 is a diagram showing a specific aspect of operations performed by a system in accordance with some examples of the present disclosure by Figure 1 the system

[0024] Figure 14 shows an example of an integrated circuit operable to perform multi-stream processing of single-stream data in accordance with some examples of the present disclosure

[0025] Figure 15 is a diagram of a mobile device operable to perform multi-stream processing of single-stream data in accordance with some examples of the present disclosure

[0026] Figure 16 is a diagram of a headset operable to perform multi-stream processing of single-stream data in accordance with some examples of the present disclosure

[0027] Figure 17 is a diagram of a wearable electronic device operable to perform multi-stream processing of single-stream data in accordance with some examples of the present disclosure

[0028] Figure 18 is a diagram of a voice-controlled speaker system operable to perform multi-stream processing of single-stream data in accordance with some examples of the present disclosure

[0029] Figure 19 is a diagram of a camera operable to perform multi-stream processing of single-stream data in accordance with some examples of the present disclosure

[0030] Figure 20 is a diagram of a headset (such as a virtual reality, mixed reality, or augmented reality headset) operable to perform multi-stream processing of single-stream data in accordance with some examples of the present disclosure

[0031] Figure 21 FIG. is of a first example of a vehicle operable to perform multi-stream processing of single-stream data according to some examples of the present disclosure.

[0032] Figure 22 FIG. is of a second example of a vehicle operable to perform multi-stream processing of single-stream data according to some examples of the present disclosure.

[0033] Figure 23 is according to some examples of the present disclosure that can be performed by Figure 1 FIG. of a specific implementation of a method for performing multi-stream processing of single-stream data by a device of.

[0034] Figure 24 FIG. is a block diagram of a specific illustrative example of a device operable to perform multi-stream processing of single-stream data according to some examples of the present disclosure. DETAILED DESCRIPTION

[0035] The performance of neural networks in processing real-time data (such as performing noise reduction in audio data during a voice call) is typically limited by the amount of memory bandwidth available to transfer weight coefficients from memory to the computing hardware used to execute the neural network. For example, the number of weight coefficients that can be sent to the computing hardware for processing an incoming audio data frame can be constrained by the available memory bandwidth and the frame rate of the incoming audio data. Additionally, the power consumption associated with sending the weight coefficients can exceed the power consumption of performing the computations associated with the weight coefficients.

[0036] Systems and methods for performing multi-stream processing of single-stream data are disclosed. For example, according to a particular aspect, single-stream data is used to generate multi-stream data using a process referred to herein as multi-stream enhancement. An example of single-stream data is single-channel audio, and multi-stream enhancement of single-channel audio can produce multiple different but related audio streams. However, single-stream data is not limited to single-channel audio, but can alternatively include dual-channel audio, multi-channel audio, or one or more other types of single-channel or multi-channel time series data.

[0037] According to some aspects, networks (such as recurrent neural networks) process each of multiple streams in parallel with one another by performing the same computations (e.g., reusing the same weights) on each of the multiple streams before reducing the multiple resulting processed streams into a single stream for output.

[0038] According to some aspects, multiple streams generated from a single stream via multi-stream enhancement are equivalent to each other but not identical. For illustration, multiple streams can be generated by performing one or more linear operations on a single stream, and the multiple streams can be numerically different from each other. As illustrative non-limiting examples, techniques that can be used to generate multiple streams include attenuation and / or amplification of single-stream data, time-domain shifting, frequency-domain phase shifting, and frequency-domain group phase shifting.

[0039] Since the multiple streams are equivalent to each other, the same neural network computation can be used to process them. Additionally, since the multiple streams are different from each other, features that may be lost in one stream can be picked up in another stream, resulting in a better output (e.g., improved speech retention for noise suppression) without increasing the number of weight coefficients compared to performing single-stream processing. Although processing more streams increases the amount of computation performed compared to processing a single stream, neural network accelerators typically have a sufficient amount of computational resources to accommodate the additional computation and are instead constrained by the memory bandwidth associated with loading weight coefficients. For illustration, components such as a neural processing unit (NPU) dedicated to neural network processing can provide dedicated circuitry to enable efficient parallel processing of very large datasets associated with machine learning models.

[0040] According to some aspects, multi-stream enhancement can be performed at runtime (e.g., during an inference operation) and applied to a recurrent network trained with only single-stream data. Alternatively, multi-stream enhancement can be performed both at training time and at runtime. Performing multi-stream enhancement during the training of a neural network enables the neural network to learn to process multi-stream enhanced data to achieve better results compared to training the neural network with single-stream training data.

[0041] Improving the signal processing performance of a neural network in view of the memory bandwidth and power constraints associated with the transfer of weight coefficients enhances device performance and improves the user experience, especially for low-power real-time applications on portable communication devices.

[0042] Certain aspects of the present disclosure are described below with reference to the accompanying drawings. In the specification, common features are denoted by common reference numerals. As used herein, various terms are for the purpose of describing particular embodiments only and are not intended to limit the embodiments. For example, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. Additionally, some features described herein are singular in some embodiments and plural in other embodiments. For illustration, Figure 1 depicting including one or more processors ( Figure 1Device 102 with a "processor" 104), which indicates that in some embodiments, device 102 includes a single processor 104, and in other embodiments, device 102 includes multiple processors 104. For ease of reference herein, unless aspects related to multiple features are described, these features are generally introduced as "one or more" features and are subsequently referred to in the singular or optional plural (as indicated by "(s)" in the feature name).

[0043] As used herein, the terms "comprise", "comprises" and "comprising" may be used interchangeably with "include", "includes" or "including". Additionally, the term "wherein" may be used interchangeably with "where". As used herein, "exemplary" indicates examples, embodiments and / or aspects and should not be construed as limiting or indicating a preference or preferred embodiment. As used herein, ordinal terms (e.g., "first", "second", "third", etc.) used to modify elements (such as structures, components, operations, etc.) do not themselves indicate any priority or order of the element relative to another element, but merely distinguish the element from another element with the same name (but using an ordinal term). As used herein, the term "set" refers to one or more of a particular element, and the term "plurality" refers to a plurality of a particular element (e.g., two or more).

[0044] As used herein, "coupled" may include "communicatively coupled", "electrically coupled" or "physically coupled", and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired network, wireless network or a combination thereof), etc. As an illustrative non-limiting example, two devices (or components) that are electrically coupled may be included in the same device or different devices and may be connected via electronics, one or more connectors or inductive coupling. In some embodiments, two devices (or components) that are communicatively coupled (e.g., electrically communicatively) may directly or indirectly send and receive signals (e.g., digital signals or analog signals) via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled or physically coupled) without an intermediate component.

[0045] In the present disclosure, terms such as "determine", "calculate", "estimate", "shift", "adjust", etc. may be used to describe how to perform one or more operations. It should be noted that these terms should not be construed as restrictive, and other techniques may be utilized to perform similar operations. Additionally, as mentioned herein, "generate", "calculate", "estimate", "use", "select", "access", and "determine" may be used interchangeably. For example, "generate", "calculate", "estimate", or "determine" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining a parameter (or signal) or may refer to using, selecting, or accessing a parameter (or signal) such as that generated by another component or device.

[0046] Referring to Figure 1 , certain illustrative aspects of a system 100 configured to perform multi-stream processing of single-stream data are shown. In Figure 1 the example shown, system 100 is configured to generate single-stream output data 140 based on multi-stream processing of single-stream data 120.

[0047] System 100 includes a device 102, and device 102 is coupled to or includes one or more sources 122 of media content of single-stream data 120. For example, sources 122 may include one or more microphones 126, one or more cameras 132, a communication channel 124, or a combination thereof. In Figure 1 the example shown, sources 122 are external to device 102 and are coupled to device 102 via an input interface 106; however, in other examples, one or more of sources 122 are components of device 102. By way of illustration, sources 122 may include a media engine (e.g., a game engine or an extended reality engine) of device 102, and the media engine generates single-stream data 120 based on instructions executed by one or more processors 104 of device 102.

[0048] Single-stream data 120 may include data representing speech 128 of a person 130. For example, when source 122 includes a microphone 126, microphone 126 may generate a signal based on the sound of speech 128 to provide single-stream data 120. When source 122 includes a camera 132, single-stream data 120 may alternatively or additionally include one or more images (e.g., video frames) depicting person 130. When source 122 includes a communication channel 124, single-stream data 120 may include transmitted data, such as multiple data packets encoding speech 128. Communication channel 124 may include or correspond to a wired connection between two or more devices, a wireless connection between two or more devices, or both. According to certain aspects, single-stream data 120 includes a sequence of data frames from source 122.

[0049] In Figure 1In [the figure], device 102 includes an input interface 106, an output interface 112, a processor 104, a memory 108, and a modem 110. The memory 108 is configured to store weight coefficients (shown as network weights 114) accessible to the operations of the processor 104 in conjunction with a network 170 (e.g., a recurrent network), as further described below. The input interface 106 is coupled to the processor 104 and is configured to be coupled to one or more of the sources 122. In an illustrative example, the input interface 106 is configured to receive a microphone output from a microphone 126 and provide the microphone output to the processor 104 as single-stream data 120.

[0050] The output interface 112 is coupled to the processor 104 and is configured to be coupled to one or more output devices, such as one or more speakers 142, one or more display devices 146, etc. The output interface 112 is configured to receive data representing single-stream output data 140 from the processor 104 and send the single-stream output data 140 to the output device. For illustration, in an embodiment where the single-stream output data 140 includes audio data, the speaker 142 is configured to output the audio of the single-stream output data 140. In an embodiment where the single-stream output data 140 includes video data, the display device 146 is configured to output the video of the single-stream output data 140.

[0051] The processor 104 is configured to receive the single-stream data 120 and generate single-stream output data 140 based on multi-stream processing of the single-stream data 120. In Figure 1 the example shown, the processor 104 includes a multi-stream enhanced data generator 160, a multi-stream data processing unit 164, and a channel reducer 168. Each of the multi-stream enhanced data generator 160, the multi-stream data processing unit 164, and the channel reducer 168 may include or correspond to dedicated hardware, instructions executable by the processor 104, or a combination thereof to perform the various operations described herein. In a particular example, the processor 104 includes, corresponds to, or is included in an NPU.

[0052] The processor 104 is configured to detect single-stream data 120 that can be received via the input interface 106 and provide the single-stream data 120 to the multi-stream enhanced data generator 160. The multi-stream enhanced data generator 160 is configured to generate multi-stream enhanced data 162 including one or more modified versions of the single-stream data 120. For example, the multi-stream enhanced data generator 160 is configured to apply one or more first operations to the single-stream data 120 to generate one or more modified versions of the single-stream data 120, as further described below. According to one aspect, the first operation produces a modified version of the single-stream data 120 that is equivalent but numerically different from the single-stream data 120. Examples of the first operation include frequency domain phase shift, frequency domain group phase shift, time domain phase shift, or application of a gain, each of which is further described in detail below.

[0053] The multi-stream data processing unit 164 is configured to process the multi-stream enhanced data 162 to generate a plurality of output channels 166. In some embodiments, the multi-stream data processing unit 164 includes one or more trained models depicted as a network 170 that processes each stream in the multi-stream enhanced data 162 in parallel and uses the same network weights 114 for each stream in the multi-stream enhanced data 162. Examples of trained models include machine learning models such as neural networks, adaptive neuro-fuzzy inference systems, support vector machines, decision trees, regression models, Bayesian models, or Boltzmann machines, or an ensemble, variant, or other combination thereof. Variants of decision trees include, for example, but are not limited to, random forests, boosted decision trees, etc. Variants of neural networks include, for example, but are not limited to, transformers, self-attention networks, convolutional neural networks, deep neural networks, deep belief networks, etc.

[0054] In some examples, the network 170 performs multi-stream processing at the multi-stream data processing unit 164 and can include, but is not limited to, a recurrent neural network (RNN) (e.g., a neural network having one or more recurrent layers, one or more long short-term memory (LSTM) layers, one or more gated recurrent unit (GRU) layers), a recurrent convolutional neural network (RCNN), a self-attention network (e.g., a transformer), other machine learning models suitable for processing time series data in a time-dynamic manner, or a variant, ensemble, or combination thereof.

[0055] The channel reducer 168 is configured to reduce a plurality of output channels 166 to generate single-stream output data 140. For example, to reduce the plurality of output channels 166, the channel reducer 168 may be configured to perform one or more second operations on the output channels 166 to generate adjusted output channel data, as further described below. The second operation corresponds to the inverse operation of the first operation applied by the multi-stream enhanced data generator 160. After performing the second operation, the channel reducer 168 combines the adjusted output channel data (e.g., averages the values from the plurality of adjusted channels) to generate single-stream output data 140.

[0056] During operation, in an illustrative embodiment, the single-stream data 120 includes audio data. In an example, the single-stream data 120 includes single-channel audio data captured by the microphone 126. In other examples, the single-stream data 120 includes dual-channel audio data (e.g., captured by two microphones 126) or multi-channel audio data (e.g., captured by more than two microphones 126). The multi-stream enhanced data generator 160 processes the single-stream data 120 to generate multi-stream enhanced data 162 based on the single-stream data 120.

[0057] Continuing with the above example, the multi-stream enhanced data 162 is input to the multi-stream data processing unit 164, and the multi-stream data processing unit 164 performs noise reduction on each stream of the multi-stream enhanced data 162 such that each output channel 166 corresponds to a noise-reduced version of the corresponding stream of the multi-stream enhanced data 162. The channel reducer 168 processes and combines the output channels 166 to generate single-stream output data 140. The single-stream output data 140 includes a noise-reduced version of the audio data.

[0058] In some embodiments, the modem 110 is configured to receive the single-stream data 120 from a second device 152 via wireless transmission over the communication channel 150. By way of illustration, the communication channel 150 may include or correspond to a wired connection between two or more devices, a wireless connection between two or more devices, or both. The single-stream data 120 may be received in conjunction with a federated learning network, as further described in reference Figure 13 and the processor 104 may also be configured to send the single-stream output data 140 to the second device 152 via the modem 110. In some embodiments, the single-stream output data 140 is provided to the modem 110 for transmission over the communication channel 150 to a device 152, such as for playback at one or more playback devices coupled to or included in the second device 152.

[0059] Although the above description has focused mainly on examples where the single - stream data 120 represents audio data, in some embodiments, the single - stream data 120 can include or correspond to image or video data, or can include or correspond to one or more types of non - media data, such as motion sensor data or any other type of time - series data. Although the above description describes the multi - stream data processing unit 164 performing noise reduction, in other embodiments, the multi - stream data processing unit 164 performs one or more other types of processing instead of or in addition to noise reduction.

[0060] In some embodiments, multi - stream augmented training data is used to train the network 170, such as during training operations performed at the device 102, at one or more other devices, or a combination thereof. For example, the processor 104 can be configured to use the multi - stream augmented training data to train the network 170, and after training, the processor 104 can use the trained network 170 to process the multi - stream augmented data 162 (e.g., process the single - stream data 120) during inference operations. In other embodiments, single - stream training data (e.g., bypassing the multi - stream augmented data generator 160 and the channel reducer 168) is used to train the network 170, such as during training operations performed at the device 102, at one or more other devices, or a combination thereof. For example, the processor 104 can be configured to use the single - stream augmented training data to train the network 170, and after training, the processor 104 can use the trained network 170 to process the multi - stream augmented data 162.

[0061] Thus, the system 100 facilitates the processing of the single - stream data 120 based on generating and processing the multi - stream augmented data 162. Compared to processing the single - stream data 120 without multi - stream augmentation, improved results are achieved at the single - stream output data 140 by increasing the number of streams processed at the recurrent network 170 but using the same set of network weights 114 for each stream, while substantially not increasing the processor bandwidth for transferring the network weights 114 from the memory 108 to the memory 104.

[0062] Figure 2 is of a specific aspect of a Figure 1 system according to some examples of the present disclosure. Specifically, Figure 2 highlights an example of the multi - stream augmented data generator 160, the multi - stream data processing unit 164, and the channel reducer 168 according to a specific embodiment.

[0063] In Figure 2In the example shown, the multi-stream enhanced data generator 160 generates multi-stream enhanced data 162 that includes one or more modified versions of the single-stream data 120, which are shown as a first modified version 210, a second modified version 212, and one or more other modified versions including a modified version 214. As shown, the multi-stream enhanced data generator 160 is configured to perform one or more first operations 202 on the single-stream data 120 to generate one or more modified versions 210-214 of the single-stream data 120. Examples of the first operations include frequency domain phase shift, frequency domain group phase shift, gain adjustment, and time domain phase shift, as further described below.

[0064] In some embodiments, the multi-stream enhanced data 162 includes the single-stream data 120. For example, the single-stream data 120 may bypass the first operation 202 as shown, or one or more of the first operations 202 that do not change the single-stream data 120 may be performed (e.g., by applying a gain of 1 or a delay of 0, etc. to the single-stream data 120). However, in other embodiments, the multi-stream enhanced data 162 may not include the single-stream data 120 (e.g., each stream of the multi-stream enhanced data 162 is different from the single-stream data 120).

[0065] The multi-stream data processing unit 164 processes the multi-stream enhanced data 162 to generate an output channel 166, and the channel reducer 168 processes the output channel 166 to generate the single-stream output data 140. To reduce the multiple output channels 166 into a single output stream, the channel reducer 168 is configured to perform one or more second operations 204 on at least one of the multiple output channels 166 to generate adjusted multi-channel output data 230, and perform a combining operation 206 on the channels of the adjusted multi-channel output data 230 to generate the single-stream output data 140.

[0066] One or more of the second operations 204 correspond to the inverse operations of one or more of the first operations 202. As used herein, "inverse operation" is used to reverse the changes performed by a previous operation. For example, if the first operation applies a gain of 2 to a signal, the inverse operation of that first operation applies a gain of 0.5. As another example, if the first operation applies a time shift or phase shift of 1 unit to a signal, then the inverse operation of that first operation applies a time shift or phase shift of -1 unit.

[0067] In a particular example, the combining operation 206 includes averaging the values of the channels of the adjusted multi-channel output data 230. For example, the combining operation 206 may perform an averaging operation (e.g., arithmetic mean) on the first sample or data unit of each channel of the adjusted multi-channel output data 230 to generate the first sample or data unit of the single-stream output data 140, perform an averaging operation on the second sample or data unit of each channel of the adjusted multi-channel output data 230 to generate the second sample or data unit of the single-stream output data 140, and so on.

[0068] Generating the multi-stream enhanced data 162 provides a diversity of equivalent but different data streams for processing by the network 170. As a result, one or more features or characteristics in the single-stream data 120 can be presented to the network 170 in various resolutions, time scales, etc. in the various streams of the multi-stream enhanced data 162, thereby achieving a more robust overall performance of the network 170 with respect to such features or characteristics. For example, processing of one or more of the multiple output channels 166 can have improved results (e.g., greater noise reduction) compared to processing the single-stream data 120. Performing the second operation 204 reverses the changes imposed by the first operation(s) 202 and restores the output channels 166 to a common condition (e.g., re-aligning in time, returning to the original gain level, etc.), which enables the combining operation 206 to combine the adjusted multi-channel output data 230 to form the single-stream output data 140.

[0069] In some embodiments, the multi-stream enhanced data 162 includes M streams, where M is an integer greater than 1. In experiments where the network 170 performs single-channel noise suppression on voice calls using different values of M (and without increasing the number of network weights 114), it has been observed that larger values of M result in increased noise reduction performance compared to smaller values of M. This result has been observed for cases where the network 170 is trained using single-stream training data, and to an even greater extent for cases where the network 170 is trained using multi-stream enhanced data. In one example, it has been observed that the noise reduction performance for M = 12 is substantially similar (e.g., within 1 - 2% in terms of the Perceptual Objective Listening Quality Analysis (POLQA) score for mobile voice call data) or better (e.g., significantly higher POLQA scores for hands-free voice call data) compared to processing single-stream audio data using a similar network with approximately twice the number of network weights. Thus, the multi-stream enhancement techniques described herein can provide similar or improved performance while using approximately half the number of weights.

[0070] Figure 3 is according to some examples of the present disclosureFigure 1 A diagram of a particular aspect of the system. In particular, Figure 3 Highlights a first example of a first operation 202 that can be performed by the multi-stream enhanced data generator 160 according to a particular embodiment.

[0071] In Figure 3 the example illustrated, the first operation 202 includes performing a frequency domain transform (illustrated as a Fast Fourier Transform (FFT) 302) and one or more frequency domain phase shifts 304. The FFT 302 processes the single-stream data 120 (represented as x(t)) to generate a frequency domain version 312 of the single-stream data 120. The frequency domain version 312 is represented as X(n,k), where n indicates the sequence index and k indicates the bin index.

[0072] The frequency domain phase shift 304 includes applying different phase shifts to the frequency domain version 312 to generate multiple sets of phase-shifted data. For example, a first phase shift 320 (e.g., a constant phase shift across all frequency bins) can be applied to the frequency domain version 312 via a multiplier 306 to generate a first phase-shifted data 330, represented as Y 1 (n,k). The first phase shift 320 can be applied as where j represents the square root of -1, and represents the constant phase shift. Other phase shifts (e.g., other values of ) can be applied to the frequency domain version 312 to generate other phase-shifted data, including a Mth phase shift 324 applied to the frequency domain version 312 via a multiplier 308 to generate Mth phase-shifted data 334, represented as Y M (n,k). In this example, the M sets of the resulting phase-shifted data 330 - 334 form the multi-stream enhanced data 162.

[0073] Figure 4 is a diagram of a particular aspect of the Figure 1 system according to some examples of the present disclosure. In particular, Figure 4 Highlights a first example of a second operation 204 that can be performed by the channel reducer 168 according to a particular embodiment.

[0074] In Figure 4 the example illustrated, the second operation 204 includes performing one or more inverse frequency domain phase shifts 404 on a single channel of the output channel 166, which reverse the Figure 3 frequency domain phase shift 304 illustrated in. For example, a first inverse phase shift 420 (e.g., a constant phase shift across all frequency bins) can be applied to the data of the first channel in the output channel 166, represented as Y 1 ′(n,k)410, via a multiplier 406 to generate a first adjusted data, represented as X1 ′(n,k)430. In the illustrated embodiment, Y 1 ′(n,k)410 is processed at the multi-stream data processing unit 164 Figure 3 of Y 1 (n,k)330, and the first inverse phase shift 420 can be applied as

[0075] Other inverse phase shifts can be applied to other channels of the output channel 166 to generate other adjusted data, including the Mth inverse phase shift 424 applied to the Mth channel of the output channel 166 (denoted as Y M ′(n,k)414) of the data to generate the Mth adjusted data denoted as X M ′(n,k)434. In the illustrated embodiment, Y M ′(n,k)414 is processed at the multi-stream data processing unit 164 Figure 3 of Y M (n,k)334, and the Mth inverse phase shift 424 can be applied to reverse the Mth phase shift 324.

[0076] In Figure 4 the example illustrated, the second operation 204 also includes performing an inverse transform, illustrated as an inverse FFT (IFFT) 402, on each of the frequency domain adjusted data 430 to 434 to generate time domain adjusted data 440 to 444. For example, processing X 1 ′(n,k)430 to generate the first time domain adjusted data x 1 ′(t)440, and X M ′(n,k)434 is processed to generate the Mth time domain adjusted data x M ′(t)444. In this example, the M sets of the resulting time domain adjusted data 440 - 444 form the adjusted multi-channel output data 230.

[0077] Figure 5 is a diagram of a particular aspect of a system in accordance with some examples of the present disclosure. In particular, Figure 1 highlights a second example of the first operation 202 that can be performed by the multi-stream enhanced data generator 160 according to a particular embodiment. Figure 5 In

[0078] the example illustrated, the first operation 202 includes performing a frequency domain transform (illustrated as FFT 302) and one or more frequency domain group phase shifts 504. The FFT 302 processes the single-stream data x(t)120 to generate a frequency domain version X(n,k)312 of the single-stream data 120. Figure 5 In

[0079] The frequency domain group phase shift 504 includes applying different sets of group phase shifts to the frequency domain version X(n,k) 312 to generate multiple sets of group phase shifted data. For example, a first group delay 520 can be applied to the frequency domain version X(n,k) 312 via a multiplier 506 to generate data Y 1 (n,k) 530 of the first group delay. The first group delay 520 can be applied to each frequency bin in the form of exp(j2πkτ / N), where exp() represents the exponential function, k represents the bin index, N is the FFT size, and τ is the group delay. According to some embodiments, the absolute group delay |τ| is much smaller than the window size. Other group delays (e.g., other values of τ) can be applied to the frequency domain version X(n,k) 312 to generate other group delay data, including applying a multiplier 508 to the frequency domain version X(n,k) 312 to generate data Y M (n,k) 534 of the Mth group delay 524. In this example, the M sets of the resulting group delay data 530 - 534 form the multi-stream enhanced data 162.

[0080] Figure 6 is according to some examples of the present disclosure Figure 1 of a particular aspect of a system. In particular, Figure 6 highlights a second example of a second operation 204 that can be performed by the channel reducer 168 according to a particular embodiment.

[0081] In Figure 6 the example shown, one or more second operations 204 include performing one or more inverse frequency domain group phase shifts 604 on a single channel of the output channel 166, which reverse Figure 5 the frequency domain group phase shift 504 shown. For example, a first inverse group delay 620 can be applied to the data of the first channel in the output channel 166, denoted as Y 1 ′(n,k) 610, to generate first adjusted data, denoted as X 1 ′(n,k) 630. In the embodiment shown, Y 1 ′(n,k) 610 corresponds to the result of Y Figure 5 processed at the multi-stream data processing unit 164 1 (n,k) 530, and the first inverse group delay 620 can be applied to each frequency bin in the form of exp(-j2πkτ / N).

[0082] Other inverse group delays can be applied to other channels of the output channel 166 to generate other adjusted data, including applying a multiplier 608 to the Mth channel of the output channel 166 (denoted as Y MThe M-th inverse group delay 624 of the data of ′(n,k)614) to generate what is denoted as X M The M-th adjusted data of ′(n,k)634. In the illustrated embodiment, Y M ′(n,k)614 is processed at the multi-stream data processing unit 164 Figure 5 with the Y M (n,k)534. The result is corresponding, and the M-th inverse group delay 624 can be applied to reverse the M-th group delay 524

[0083] In Figure 6 the example illustrated, the second operation 204 also includes performing an inverse transform, illustrated as IFFT 602, on each of the frequency domain adjusted data 630 - 634 to generate time domain adjusted data 640 - 644. For example, X 1 ′(n,k)630 is processed to generate first time domain adjusted data x 1 ′(t)640, and X M ′(n,k)634 is processed to generate the M-th time domain adjusted data x M ′(t)644. In this example, the M sets of the resulting time domain adjusted data 640 - 644 form the adjusted multi-channel output data 230

[0084] Figure 7 is a diagram of a particular aspect of a Figure 1 system according to some examples of the present disclosure. In particular Figure 7 highlights a third example of the first operation 202 that can be performed by the multi-stream enhanced data generator 160 according to a particular embodiment

[0085] In Figure 7 the example illustrated, the first operation 202 includes performing one or more gain adjustments 704. The gain adjustment 704 includes applying different gains to the single-stream data x(t)120 to generate multiple sets of gain adjusted data. For example, a first gain g 720 can be applied to the single-stream data x(t)120 via a multiplier 706 to generate first gain adjusted data y 1 (t)730. Other gains can be applied to the single-stream data x(t)120 to generate other gain adjusted data, including a M-th gain 724 applied to the single-stream data x(t)120 via a multiplier 708 to generate the M-th gain adjusted data y M (t)734. In this example, the M sets of the resulting gain adjusted data 730 - 734 form the multi-stream enhanced data 162

[0086] Figure 8 is a diagram of a particular aspect of a Figure 1Diagram of a specific aspect of the system. In particular, Figure 8 Highlights a third example of a second operation 204 that can be performed by the channel reducer 168 according to a specific embodiment.

[0087] In Figure 8 the example shown, one or more second operations 204 include performing one or more inverse gain adjustments 804 on a single channel of the output channel 166, which reverse Figure 7 the gain adjustment 704 shown. For example, a first inverse gain 820 can be applied to the data of the first channel of the output channel 166, denoted as y 1 ′(t)810, via a multiplier 806 to generate first adjusted data, denoted as x 1 ′(t)830. In the embodiment shown, y 1 ′(t)810 corresponds to the result of processing Figure 7 y 1 (t)730 at the multi-stream data processing unit 164, and the first inverse gain 820 can be applied in the form of 1 / g.

[0088] Other inverse gains can be applied to other channels of the output channel 166 to generate other adjusted data, including the Mth inverse gain 824 applied to the data of the Mth channel of the output channel 166 (denoted as y M ′(t)814) via a multiplier 808 to generate the Mth adjusted data denoted as x M ′(t)834. In the embodiment shown, y M ′(t)814 corresponds to the result of processing Figure 7 y M (t)734 at the multi-stream data processing unit 164, and the Mth inverse gain 824 can be applied as the inverse (e.g., reciprocal) of the Mth gain 724. In this example, the M sets of the resulting adjusted data 840 - 844 form the adjusted multi-channel output data 230.

[0089] Figure 9 is a diagram of a specific aspect of the Figure 1 system according to some examples of the present disclosure. In particular, Figure 9 highlights a fourth example of a first operation 202 that can be performed by the multi-stream enhanced data generator 160 according to a specific embodiment.

[0090] In Figure 9In the example illustrated, one or more first operations 202 include performing one or more time domain shifts 904. The time domain shift 904 includes applying different shifts (e.g., forward or backward) to the single-stream data x(t) 120 to generate multiple sets of shifted data. In frame-by-frame processing, this can be achieved by halving (or 1 / 3 or 1 / 4, etc.) the hop size while keeping the same window function. For example, the first figure 950 graphically shows a simplified example of a set of window functions associated with frame-by-frame processing, and the second figure 952 shows the set of window functions after applying the shift.

[0091] In the illustrated embodiment of the time domain shift 904, a first shift amount 920 can be applied to the single-stream data x(t) 120 via a shifter 906 to generate first shifted data y 1 (t) 930. Other shift amounts can be applied to the single-stream data x(t) 120 to generate other shifted data, including a first M shift amount 924 applied to the single-stream data x(t) 120 via a shifter 908 to generate Mth shifted data y M (t) 934. In this example, the M sets of the resulting shifted data 930 - 934 form multi-stream enhanced data 162.

[0092] Figure 10 is of a specific aspect of a Figure 1 system according to some examples of the present disclosure. In particular, Figure 10 highlights a fourth example of a second operation 204 that can be performed by a channel reducer 168 according to a specific embodiment.

[0093] In Figure 10 the example illustrated, one or more second operations 204 include performing inversion on a single channel of the output channel 166 Figure 9 and one or more inverse time domain shifts 1004 of the time domain shift 904 illustrated in 1 . For example, a first inverse shift amount 1020 can be applied to the data of the first channel in the output channel 166 (labeled as y 1 ′(t) 1010) via a shifter 1006 to generate first adjusted data (labeled as x 1 ′(t) 1030). In the illustrated embodiment, y Figure 9 ′(t) 1010 corresponds to the result of y 1 (t) 930 processed at the multi-stream data processing unit 164, and the first inverse shift amount 1020 can have the same magnitude as the first shift amount 920 but in the opposite direction.

[0094] Other inverse shifts can be applied to other channels in output channel 166 to generate other adjusted data, including the Mth inverse shift amount 1024 applied to the Mth channel in output channel 166 via shifter 1008 (represented as y M ′(t)1014), to generate the Mth adjusted data represented as x M ′(t)1034. In the illustrated embodiment, y M ′(t)1014 corresponds to the result of processing y Figure 9 at multi-stream data processing unit 164 M (t)934, and the Mth inverse shift amount 1024 can be applied as the inverse of the Mth shift amount 924 (e.g., equal magnitude, opposite direction). In this example, the M sets of the resulting adjusted data 1040 - 1044 form adjusted multi-channel output data 230.

[0095] Figure 11 is a diagram of a specific aspect of a Figure 1 system according to some examples of the present disclosure. Specifically, Figure 11 highlights a first example of network 170 implemented in NPU 1104 according to a specific embodiment.

[0096] In Figure 11 the example shown, NPU 1104 includes a multi-stream enhanced data generator 160, network 170, and channel reducer 168. For example, processor 104 can be included in NPU 1104. NPU 1104 is coupled to another processor, shown as digital signal processor (DSP) 1102. However, in other embodiments, NPU 1104 can be coupled to one or more other types of processors, by way of illustrative non-limiting examples, such as a central processing unit (CPU).

[0097] NPU 1104 is also coupled to memory 108 and is configured to access network weights 114 in conjunction with processing multi-stream enhanced data 162. However, the storage capacity in NPU 1104 (shown as random access memory (RAM) 1120) may not be sufficient to store the entire set of network weights 114 on-chip. As a result, NPU 1104 can sequentially access a first set 1110 of network weights 114 from memory 108 to perform a first part of processing multi-stream enhanced data 162, access a second set 1112 of network weights 114 to perform a second part of the processing, etc., until accessing a Kth set 1114 of network weights 114 to perform a Kth part of the processing of multi-stream enhanced data 162 (where K is an integer greater than 1).

[0098] For example, the first set 1110 may correspond to the weights of one or more first layers of the network 170. After the first frame of each stream of the multi-stream enhanced data 162 is processed in parallel at one or more first layers using the first set 1110 of network weights 114, the NPU 1104 may retrieve the second set 1112 from the memory 108 and store the second set 1112 in the RAM 1120, overwriting the first set 1110. The second set 1112 may correspond to the weights of one or more second layers of the network 170, which are used to continue the parallel processing of the first frame of each stream of the multi-stream enhanced data 162. The processing continues until the Kth set 1114 of one or more final layers of the network 170 has been stored in the RAM 1120 and is used to complete the processing of the first frame of each stream of the multi-stream enhanced data 162, resulting in the generation of the first frame of each output channel in the plurality of output channels 166. After the first frame of each of the plurality of output channels 166 is generated, the first set 1110 is loaded into the RAM 1120 again, and the NPU 1104 begins to process the second frame of each stream in the multi-stream enhanced data 162 in parallel at one or more first layers.

[0099] For real-time processing such as real-time audio noise reduction, the NPU 1104 has excess computing power, but the performance of the NPU 1104 may be constrained by the size of the network 170 in terms of the number of network weights 114, the memory bandwidth available to transfer the network weights 114 from the memory 108 to the NPU 1104, the power consumption associated with transferring the network weights 114, or a combination thereof. Although increasing the size of the RAM 1120 can reduce or eliminate the repeated transfer of the weight sets 1110 - 1114 for each sequential input frame of the multi-stream enhanced data 162, the size of the RAM 1120 can be constrained based on factors such as chip size, chip cost, and power consumption, especially when the NPU 1104 is implemented in a portable electronic device.

[0100] By using the multi-stream enhanced data 162, the performance of the network 170 can be enhanced by increasing the number of streams processed in parallel by the network 170 using the excess computing power of the NPU 1104 without increasing the number of network weights 114.

[0101] Figure 12 is of a specific aspect of a Figure 1 system according to some examples of the present disclosure. Specifically, Figure 12 highlights a second example of the network 170 implemented in the NPU 1104 according to a specific embodiment.

[0102] In Figure 12In the example shown, the multi-stream enhanced data generator 160 and the channel reducer 168 are implemented at the DSP 1102 rather than at the NPU 1104. The multi-stream enhanced data 162 is transferred from the DSP 1102 to the NPU 1104 and is processed as referenced Figure 11 as described. After processing one or more frames of the multi-stream enhanced data 162 at the NPU 1104 (e.g., after the first frame of each of the plurality of output channels 166 has been generated), one or more frames of the output channels 166 are transferred from the NPU 1104 to the channel reducer 168 at the DSP 1102, which generates corresponding frames of the single-stream output data 140.

[0103] Figure 13 is a diagram illustrating specific aspects of operations performed by a Figure 1 system in accordance with some examples of the present disclosure. Specifically, Figure 13 highlights an example of communicating between multiple devices using components of the system 100 in conjunction with the federated learning network 1304 in accordance with a particular implementation.

[0104] In Figure 13 the example illustrated, the federated learning network 1304 includes a primary device 1302 (e.g., a user device) and a plurality of other devices, illustrated as devices 1310, 1312, and one or more other devices including device 1314. In a particular implementation, one or more of the devices 1310-1314 correspond to edge devices, and the devices 1310-1314 can include various computing capabilities. In the example, one or more of the devices 1310-1314 correspond to servers, personal computers, portable electronic devices, or one or more other devices coupled to the device 1302 via one or more wired or wireless networks.

[0105] In a particular implementation, each of the devices 1310-1312 is configured to perform multi-stream enhancement and reduction functionality in a manner similar to that described for the device 102. For example, the device 1310 is configured to receive single-stream input data and perform enhancement 1320 (e.g., as described for the multi-stream enhanced data generator 160), network processing 1322 (e.g., performing inference, training, or both at the network 170), and de-enhancement 1324 (e.g., as described for the channel reducer 168) to generate output data 1326 (e.g., Figure 1For the single - stream output data 140), the device 1310 can send the output data to the device 1302 via a modem (e.g., the modem 110). Similarly, the device 1312 is configured to perform enhancement 1330, network processing 1332 (e.g., inference, training, or both) and de - enhancement 1334 to generate output data 1336, and the device 1314 is configured to perform enhancement 1340, network processing 1342 (e.g., inference, training, or both) and de - enhancement 1344 to generate output data 1346.

[0106] According to some embodiments, the devices 1310 - 1314 operate as a distributed computing network for performing signal processing. For example, the device 1302 can detect available nodes in the local network environment and send a copy of the single - stream data 120 to each available node (e.g., the devices 1310 - 1314). Each of the devices 1310 - 1314 uses the enhancement, network processing, and de - enhancement capabilities of that device to locally process the single - stream data 120 to generate respective output data sets 1326, 1336, and 1346. Each of the output data sets 1326, 1336, and 1346 includes a version of the single - stream output data 140 generated by the respective devices 1310, 1312, and 1314 based on the single - stream data 120. The output data sets 1326, 1336, and 1346 can be combined (e.g., reduced, such as via weighted or unweighted averaging) at the parameter averaging / reduction operation 1350 to generate the output 1352. The parameter averaging / reduction operation 1350 can be performed at the device 1302, at one or more of the devices 1310 - 1314, or at another device.

[0107] The device 1302 uses the output 1352 to generate the single - stream output data 140. In some embodiments, the device 1302 does not perform signal processing on the single - stream data 120, and the single - stream output data 140 matches the output 1352. In other embodiments, the device 1302 can be Figure 1 corresponding to the device 102 and can process the single - stream data 120 in parallel with the processing performed at the devices 1310 - 1314. For example, the device 1302 can include the output 1352 as an input to the combining operation 206 at the channel reducer 168. As another example, the device 1302 can combine the single - stream output data generated at the channel reducer 168 with the output 1352 to generate the single - stream output data 140.

[0108] In some embodiments, device 1302 may transmit enhancement parameters to each of devices 1310 - 1314 such that devices 1310 - 1314 do not perform the same calculations. For example, device 1302 may perform enhancement and reduction using gain adjustment, and may instruct device 1310 to use frequency domain phase shift, device 1312 to use frequency domain group phase shift, and device 1314 to use time domain phase shift. By distributing processing among multiple devices 1310 - 1314, device 1302 may obtain the benefits of various different types of enhancement and reduction techniques to generate single - stream output data 140.

[0109] In some embodiments, the federated learning network 1304 is configured to perform distributed training to determine or update parameters associated with enhanced multi - stream processing, such as network weights 114. For example, device 1310 may receive a copy of the parameters from device 1302 and may perform training operations on a local version of network 170 using locally stored data streams as training data to generate updated parameters. Similarly, device 1312 may receive a copy of the parameters and may perform training operations using locally stored data streams at device 1312 as training data to generate updated parameters, and device 1314 may receive a copy of the parameters and use locally stored data streams at device 1314 as training data to perform training operations to generate updated parameters.

[0110] The updated parameters generated by device 1310 may be included in output data 1326, the updated parameters generated by device 1312 may be included in output data 1336, and the updated parameters generated by device 1314 may be included in output data 1346. The updated parameters may be combined (e.g., averaged) at parameter averaging / reduction operation 1350 to generate a set of updated parameters included in output 1352 provided to device 1302. Since the data used as training data remains local to each of devices 1310 - 1314, a set of updated parameters may be generated based on a wide variety of data from multiple devices without compromising the privacy of any data used in training.

[0111] In some embodiments, devices 1310 - 1314 are clustered or grouped according to computational ability (such as by processor type). The clusters may be ranked and / or prioritized based on relative computational ability. For example, when combining updated parameters from various clusters at parameter averaging / reduction operation 1350, a weighted average may be used, where updates from clusters with stronger computational ability may be given more weight compared to updates from clusters with relatively less computational ability.

[0112] Figure 14Embodiment 1400 of device 102 is depicted as an integrated circuit 1402 that includes one or more processors 104. The integrated circuit 1402 also includes a signal input 1404, such as one or more bus interfaces, to enable receipt of single-stream data 120 for processing. The integrated circuit 1402 also includes a signal output 1406, such as a bus interface, to enable transmission of an output signal, such as single-stream output data 140. In Figure 14 the example shown, processor 104 includes a multi-stream enhancement engine 1410, which includes a multi-stream enhanced data generator 160, a multi-stream data processing unit 164, and a channel reducer 168. The integrated circuit 1402 enables an embodiment in operation to perform multi-stream processing of single-stream data as a component in a system that includes a microphone, such as Figure 15 a mobile phone or tablet as shown, Figure 16 a headset as shown, Figure 17 a wearable electronic device as shown, Figure 18 a voice-controlled speaker system as shown, Figure 19 a camera as shown, Figure 20 a virtual reality, mixed reality, or augmented reality headset as shown, or Figure 21 or Figure 22 a vehicle as shown.

[0113] As an illustrative non-limiting example, Figure 15 embodiment 1500 is depicted, in which device 102 includes a mobile device 1502, such as a phone or tablet. The mobile device 1502 includes a microphone 126, a camera 132, and a display screen 1504. The components of processor 104, including the multi-stream enhancement engine 1410, are integrated in the mobile device 1502 and are shown using dashed lines to indicate internal components that are generally not visible to the user of the mobile device 1502. In a particular example, the multi-stream enhancement engine 1410 operates to perform multi-stream processing of an input media stream. For example, the microphone 126 may capture the voice of a user of the mobile device 1502, and the multi-stream enhancement engine 1410 may process the captured voice to generate an output media stream corresponding to a noise-reduced version of the voice.

[0114] Figure 16Depicts Embodiment 1600, where device 102 includes a headset device 1602. The headset device 1602 includes a microphone 126. Components of the processor 104, including the multi-stream enhancement engine 1410, are integrated in the headset device 1602. In a particular example, the multi-stream enhancement engine 1410 operates to perform multi-stream processing of an input media stream. For example, the microphone 126 can capture the voice of a user of the headset device 1602, and the multi-stream enhancement engine 1410 can process the captured voice to generate an output media stream corresponding to a noise-reduced version of the voice. The noise-reduced version of the voice can be used to output the media stream from one or more speakers 142 of the headset device 1602, or can be sent to another device (e.g., a mobile device, a gaming console, a voice assistant, etc.) for playback of the output media stream.

[0115] Figure 17 Depicts Embodiment 1700, where device 102 includes a wearable electronic device 1702, shown as a "smartwatch", and the wearable electronic device 1702 includes a processor 104 and a display screen 1704. Components of the processor 104, including the multi-stream enhancement engine 1410, are integrated in the wearable electronic device 1702. In a particular example, the multi-stream enhancement engine 1410 operates to perform multi-stream processing of an input media stream. For example, the microphone 126 can capture the voice of a user of the wearable electronic device 1702, and the multi-stream enhancement engine 1410 can process the captured voice to generate an output media stream corresponding to a noise-reduced version of the voice. The noise-reduced version of the voice can be used to generate an output at the display screen 1704 of the wearable electronic device 1702, such as in conjunction with a voice interface, or can be sent to another device (e.g., a mobile device, a gaming console, a voice assistant, etc.) for playback of the output media stream.

[0116] Figure 18 Is Embodiment 1800, where device 102 includes a wireless speaker and a voice-activated device 1802. The wireless speaker and voice-activated device 1802 can have a wireless network connectivity and is configured to perform an assistive operation. Figure 18 The wireless speaker and voice-activated device 1802 includes a processor 104, which includes the multi-stream enhancement engine 1410. Additionally, the wireless speaker and voice-activated device 1802 includes a microphone 126 and a speaker 142. During operation, in response to receiving an input media stream including a user's voice, the multi-stream enhancement engine 1410 operates to perform multi-stream processing of the input media stream. For example, the microphone 126 can capture the voice of a user of the wireless speaker and voice-activated device 1802, and the multi-stream enhancement engine 1410 can process the captured voice to generate an output media stream corresponding to a noise-reduced version of the voice, which can be used in conjunction with a voice interface to provide instructions for an assistive operation.

[0117] Figure 19 depicts Embodiment 1900, in which device 102 is integrated into or includes a portable electronic device corresponding to camera 132. In Figure 19 it, camera 132 includes processor 104 and microphone 126. Processor 104 includes multi-stream enhancement engine 1410. During operation, camera 132, microphone 126, or both generate an input media stream, and multi-stream enhancement engine 1410 operates to perform multi-stream processing of the input media stream. For example, microphone 126 can capture the speech of a user of camera 132, and multi-stream enhancement engine 1410 can process the captured speech to generate an output media stream corresponding to a noise-reduced version of the speech, which can be used in conjunction with a speech interface to provide operation instructions to camera 132. In another embodiment, multi-stream enhancement engine 1410 is configured to perform processing of an image data stream corresponding to video captured by camera 132, such as performing jitter filtering, trailing filtering, or one or more other types of processing.

[0118] Figure 20 depicts Embodiment 2000, in which device 102 includes a portable electronic device corresponding to extended reality headset 2002 (e.g., a virtual reality headset, a mixed reality headset, an augmented reality headset, or a combination thereof). Extended reality headset 2002 includes microphone 126 and processor 104. In certain aspects, a visual interface device is positioned in front of the user's eyes to enable the display of augmented reality, mixed reality, or virtual reality images or scenes to the user when wearing extended reality headset 2002. In a particular example, the visual interface device is configured to display a notification indicating detected user speech in an audio signal from microphone 126. In a particular embodiment, processor 104 includes multi-stream enhancement engine 1410. During operation, microphone 126 can generate an input media stream including the speech of a user of extended reality headset 2002, and multi-stream enhancement engine 1410 can process the captured speech to generate an output media stream corresponding to a noise-reduced version of the speech. As an illustrative non-limiting example, the output media stream can be sent to an extended reality server or to other participants in a shared virtual environment, or can be used in conjunction with a speech interface to provide operation instructions to extended reality headset 2002.

[0119] Figure 21Depicts an embodiment 2100, where device 102 corresponds to or is integrated within vehicle 2102, which is illustrated as a manned or unmanned aerial device (e.g., a package delivery drone). Microphone 126 and processor 104 are integrated into vehicle 2102. In a particular embodiment, processor 104 includes a multi-stream enhancement engine 1410. During operation, microphone 126 can capture the speech of a person near vehicle 2102 (such as speech including delivery instructions from an authorized user of vehicle 2102), and multi-stream enhancement engine 1410 can process the captured speech to generate an output media stream corresponding to a noise-reduced version of the speech. As an illustrative non-limiting example, the output media stream can be sent to another device (e.g., a server device), or can be used in conjunction with a voice interface to provide operation instructions or queries to vehicle 2102.

[0120] Figure 22 Depicts another embodiment 2200, where device 102 corresponds to or is integrated within vehicle 2202, which is shown as an automobile. Vehicle 2202 includes processor 104, which includes multi-stream enhancement engine 1410. Vehicle 2202 also includes microphone 126, speaker 142, and display device 146. Microphone 126 is positioned to capture the words of an operator of vehicle 2202 or a passenger in vehicle 2202. During operation, microphone 126 can capture the speech of an operator or passenger of vehicle 2102, and multi-stream enhancement engine 1410 can process the captured speech to generate an output media stream corresponding to a noise-reduced version of the speech. As an illustrative non-limiting example, the output media stream can be sent to another device (e.g., a server device), or can be used in conjunction with a voice interface to provide operation instructions or queries to vehicle 2202.

[0121] Reference Figure 23 , shows a particular embodiment of a method 2300 for multi-stream processing of single-stream data. In a particular aspect, one or more operations of method 2300 are performed by Figure 1 at least one of multi-stream enhanced data generator 160, multi-stream data processing unit 164, channel reducer 168, processor 104, device 102, device 152, system 100, or a combination thereof.

[0122] Method 2300 includes: at block 2302, detecting single-stream data at one or more processors. For example, processor 104 can detect the reception of single-stream data 120 via input interface 106, via modem 110, or both.

[0123] Method 2300 includes: at block 2304, generating multi-stream enhanced data including one or more modified versions of single-stream data. For example, multi-stream enhanced data generator 160 generates multi-stream enhanced data 162 including one or more modified versions of single-stream data 120, such as by applying first operation 202.

[0124] Method 2300 includes: at block 2306, processing the multi-stream enhanced data to generate a plurality of output channels. For example, multi-stream data processing unit 164 processes multi-stream enhanced data 162 at network 170 to generate a plurality of output channels 166.

[0125] Method 2300 includes: at block 2308, reducing the plurality of output channels to generate single-stream output data. For example, channel reducer 168 processes output channels 166 to generate single-stream output data 140.

[0126] In some embodiments, method 2300 includes performing one or more first operations on the single-stream data to generate one or more modified versions of the single-stream data. For example, multi-stream enhanced data generator 160 performs one or more first operations 202, which may include frequency domain phase shift 304, frequency domain group phase shift 504, time domain phase shift 904, applying a gain (such as described by reference gain adjustment 704), or a combination thereof.

[0127] According to a particular aspect, reducing the plurality of output channels includes performing one or more second operations on at least one of the multi-output channels to generate adjusted multi-channel output data, wherein the one or more second operations correspond to inverse operations of the one or more first operations. For example, channel reducer 168 may perform one or more second operations 204, which may include inverse frequency domain phase shift 404, inverse frequency domain group phase shift 604, inverse time domain phase shift 1004, inverse gain adjustment 804, or a combination thereof. Reducing the plurality of output channels further includes combining the channels of the adjusted multi-channel output data to generate single-stream output data, such as described by reference combination operation 206.

[0128] Figure 23 Method 2300 can be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit such as an NPU, CPU, DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Figure 23 Method 2300 can be executed by a processor executing instructions, such as described by reference to Figure 24 as described.

[0129] Referring to Figure 24 , a block diagram of a particular illustrative embodiment of a device is depicted and is generally designated as 2400. In various embodiments, device 2400 may have more thanFigure 24 more or fewer components than those described in. In an illustrative embodiment, device 2400 may correspond to device 102 or device 152. In an illustrative embodiment, device 2400 may perform one or more operations referenced Figures 1 - 23 described.

[0130] In a particular embodiment, device 2400 includes a processor 2406 (e.g., a central processing unit (CPU)). Device 2400 may include one or more additional processors 2410 (e.g., one or more NPUs, one or more DSPs, or a combination thereof). In a particular aspect, Figure 1 processor 104 corresponds to processor 2406, processor 2410, or a combination thereof. Processor 2410 may include a voice and music codec (coder-decoder, CODEC) 2408, which includes a voice encoder (“vocoder”) encoder 2436, a vocoder decoder 2438, a multi-stream enhanced data generator 160, a multi-stream data processing unit 164, a channel reducer 168, or a combination thereof.

[0131] Device 2400 may include a memory 108 and a CODEC 2434. Memory 108 may include instructions 2456 executable by one or more additional processors 2410 (or processor 2406) to implement the functionality described with reference to the multi-stream enhanced data generator 160, the multi-stream data processing unit 164, the channel reducer 168, or a combination thereof. In Figure 24 the example shown, memory 108 also includes network weights 114.

[0132] In Figure 24 , device 2400 includes a modem 110 coupled to an antenna 2452 via a transceiver 2450. The modem 110, transceiver 2450, and antenna 2452 may be operable to receive an input media stream, transmit an output media stream, or a combination thereof.

[0133] Device 2400 may include a display device 146 coupled to a display controller 2426. A speaker 142 and a microphone 126 may be coupled to a CODEC 2434. The CODEC 2434 may include a digital-to-analog converter (DAC) 2402, an analog-to-digital converter (ADC) 2404, or both. In a particular embodiment, the CODEC 2434 may receive an analog signal from the microphone 126, convert the analog signal into a digital signal using the analog-to-digital converter 2404, and provide the digital signal to a voice and music codec 2408. The voice and music codec 2408 may process the digital signal, and the digital signal may be further processed by a multi-stream enhanced data generator 160, a multi-stream data processing unit 164, a channel reducer 168, or a combination thereof. In a particular embodiment, the voice and music codec 2408 may provide the digital signal to the CODEC 2434. The CODEC 2434 may convert the digital signal into an analog signal using the digital-to-analog converter 2402 and may provide the analog signal to the speaker 142.

[0134] In a particular embodiment, the device 2400 may be included in a system-in-package or system-on-chip device 2422. In a particular embodiment, the memory 108, the processor 2406, the processor 2410, the display controller 2426, the CODEC 2434, and the modem 110 are included in the system-in-package or system-on-chip device 2422. In a particular embodiment, an input device 2430 and a power supply 2444 are coupled to the system-in-package or system-on-chip device 2422. Additionally, in a particular embodiment, as Figure 24 shown, the display device 146, the input device 2430, the speaker 142, the microphone 126, the antenna 2452, and the power supply 2444 are external to the system-in-package or system-on-chip device 2422. In a particular embodiment, each of the display device 146, the input device 2430, the speaker 142, the microphone 126, the antenna 2452, and the power supply 2444 may be coupled to a component of the system-in-package or system-on-chip device 2422, such as an interface (e.g., an input interface 106 or an output interface 112) or a controller.

[0135] Device 2400 may include a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, headphones, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and a voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0136] In conjunction with the described embodiments, the apparatus includes components for generating multi-stream enhanced data including one or more modified versions of single-stream data. For example, the components for generating multi-stream enhanced data may correspond to processor 104, multi-stream enhanced data generator 160, multiplier 306, multiplier 308, multiplier 506, multiplier 508, multiplier 706, multiplier 708, shifter 906, shifter 908, NPU 1104, processor 2406, processor 2410, one or more other circuits or components configured to generate multi-stream enhanced data including one or more modified versions of single-stream data, or any combination thereof.

[0137] In conjunction with the described embodiments, the apparatus further includes components for processing the multi-stream enhanced data to generate a plurality of output channels. For example, the components for processing the multi-stream enhanced data to generate a plurality of output channels may correspond to processor 104, multi-stream data processing unit 164, network 170, NPU 1104, processor 2406, processor 2410, one or more other circuits or components configured to process the multi-stream enhanced data to generate a plurality of output channels, or any combination thereof.

[0138] In conjunction with the described implementation, the apparatus further includes components for reducing the plurality of output channels to produce single-stream output data. For example, the unit for reducing the plurality of output channels to produce single-stream output data may correspond to processor 104, channel reducer 168, multiplier 406, multiplier 408, multiplier 606, multiplier 608, multiplier 806, multiplier 808, shifter 1006, shifter 1008, NPU 1104, processor 2406, processor 2410, one or more other circuits or components configured to reduce the plurality of output channels to produce single-stream output data, or any combination thereof.

[0139] In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 108) stores instructions (e.g., instructions 2456) that, when executed by one or more processors (e.g., one or more processors 104, NPU 1104, one or more processors 2310, or processor 2406), cause the one or more processors to detect single-stream data (e.g., single-stream data 120); generate multi-stream enhanced data including one or more modified versions of the single-stream data (e.g., multi-stream enhanced data 162); process the multi-stream enhanced data to generate a plurality of output channels (e.g., output channels 166), and reduce the plurality of output channels to produce single-stream output data (e.g., single-stream output data 140).

[0140] The following describes specific aspects of the present disclosure in a set of interrelated examples:

[0141] According to Example 1, a device includes: a memory configured to store instructions; and one or more processors configured to: detect single-stream data; generate multi-stream enhanced data including one or more modified versions of the single-stream data; process the multi-stream enhanced data to generate a plurality of output channels; and reduce the plurality of output channels to produce single-stream output data.

[0142] Example 2 includes the device as described in Example 1, wherein the one or more processors are further configured to perform one or more first operations on the single-stream data to generate one or more modified versions of the single-stream data.

[0143] Example 3 includes the device as described in Example 2, wherein, in order to reduce the plurality of output channels, the one or more processors are further configured to: perform one or more second operations on at least one of the plurality of output channels to generate adjusted multi-channel output data, the one or more second operations corresponding to inverse operations of the one or more first operations; and perform a combination operation on the channels of the adjusted multi-channel output data to generate single-stream output data.

[0144] Example 4 includes the device as described in Example 3, wherein the combination operation includes averaging the values of the channels of the adjusted multi-channel output data.

[0145] Example 5 includes the device as described in any one of Examples 2 to 4, wherein the one or more first operations include frequency-domain phase shift.

[0146] Example 6 includes the device as described in any one of Examples 2 to 5, wherein the one or more first operations include frequency-domain group phase shift.

[0147] Example 7 includes the device as described in any one of Examples 2 to 6, wherein the one or more first operations include time-domain shift.

[0148] Example 8 includes the apparatus according to any one of Examples 2 to 7, wherein one or more of the first operations include applying a gain.

[0149] Example 9 includes the apparatus according to any one of Examples 1 to 8, wherein the multi-stream enhanced data further includes single-stream data.

[0150] Example 10 includes the apparatus according to any one of Examples 1 to 9, wherein one or more processors are configured to process the multi-stream enhanced data using a recurrent network that processes each stream in the multi-stream enhanced data in parallel and uses the same network weights for each stream in the multi-stream enhanced data.

[0151] Example 11 includes the apparatus as shown in Example 10, wherein the recurrent network is trained using the multi-stream enhanced training data.

[0152] Example 12 includes the apparatus as shown in Example 10, wherein the recurrent network is trained using the single-stream training data.

[0153] Example 13 includes the apparatus as shown in Example 10, wherein one or more processors are configured to: train the recurrent network using the multi-stream enhanced training data; and process the multi-stream enhanced data using the trained recurrent network during an inference operation.

[0154] Example 14 includes the apparatus of Example 10, wherein one or more processors are configured to: train the recurrent network using the single-stream training data; and process the multi-stream enhanced data using the trained recurrent network during an inference operation.

[0155] Example 15 includes the apparatus according to any one of Examples 1 to 14, wherein the single-stream data includes audio data, and wherein the single-stream output data includes a noise-reduced version of the audio data.

[0156] Example 16 includes the apparatus according to any one of Examples 1 to 15, wherein the single-stream data includes single-channel audio data.

[0157] Example 17 includes the apparatus according to any one of Examples 1 to 15, wherein the single-stream data includes dual-channel audio data.

[0158] Example 18 includes the apparatus according to any one of Examples 1 to 15, wherein the single-stream data includes multi-channel audio data.

[0159] Example 19 includes the apparatus according to any one of Examples 1 to 18, further comprising one or more speakers configured to output the audio of the single-stream output data.

[0160] Example 20 includes the device as described in any one of Examples 1 to 19, and further includes one or more microphones configured to provide single-stream data.

[0161] Example 21 includes the device as described in any one of Examples 1 to 20, and further includes a modem configured to receive single-stream data from a second device via wireless transmission.

[0162] Example 22 includes the device as described in Example 21, wherein the single-stream data is received in combination with a joint learning network, and wherein the one or more processors are further configured to send single-stream output data to the second device via the modem.

[0163] Example 23 includes the device as described in any one of Examples 1 to 22, wherein the one or more processors are included in a neural processing unit (NPU).

[0164] Example 24 includes the device as described in any one of Examples 1 to 23, wherein the memory and the one or more processors are included in a vehicle.

[0165] Example 25 includes the device as described in any one of Examples 1 to 23, wherein the memory and the one or more processors are included in an extended reality headset device.

[0166] According to Example 26, a method includes: detecting single-stream data at one or more processors; generating multi-stream enhanced data including one or more modified versions of the single-stream data; processing the multi-stream enhanced data to generate a plurality of output channels; and reducing the plurality of output channels to produce single-stream output data.

[0167] Example 27 includes the method as described in Example 26, and further includes performing one or more first operations on the single-stream data to generate one or more modified versions of the single-stream data.

[0168] Example 28 includes the method as described in Example 27, wherein reducing the plurality of output channels includes: performing one or more second operations on at least one of the plurality of output channels to generate adjusted multi-channel output data, the one or more second operations corresponding to the inverse operations of the one or more first operations; and combining the channels of the adjusted multi-channel output data to generate single-stream output data.

[0169] Example 29 includes the method as described in Example 27 or Example 28, wherein the one or more first operations include frequency-domain phase shift.

[0170] Example 30 includes the method as described in any one of Examples 27 to 29, wherein the one or more first operations include frequency-domain group phase shift.

[0171] Example 31 includes the method according to any one of Examples 27 to 30, wherein one or more first operations include time-domain shifting.

[0172] Example 32 includes the method according to any one of Examples 27 to 31, wherein one or more first operations include applying a gain.

[0173] Example 33 includes the method according to any one of Examples 26 to 32, wherein the multi-stream enhanced data further includes single-stream data.

[0174] Example 34 includes the method according to any one of Examples 26 to 33, wherein a recurrent network is used to process the multi-stream enhanced data, the recurrent network processes each stream in the multi-stream enhanced data in parallel, and the same network weights are used for each stream in the multi-stream enhanced data.

[0175] Example 35 includes the method according to Example 34, wherein the recurrent network is trained using the multi-stream enhanced training data.

[0176] Example 36 includes the method according to Example 34, wherein the recurrent network is trained using the single-stream training data.

[0177] Example 37 includes the method according to Example 34, further comprising: training the recurrent network using the multi-stream enhanced training data; and processing the multi-stream enhanced data using the trained recurrent network during the inference operation.

[0178] Example 38 includes the method according to Example 34, further comprising: training the recurrent network using the single-stream training data; and processing the multi-stream enhanced data using the trained recurrent network during the inference operation.

[0179] Example 39 includes the method according to any one of Examples 26 to 39, wherein the single-stream data includes audio data, and wherein the single-stream output data includes a noise-reduced version of the audio data.

[0180] Example 40 includes the method according to any one of Examples 26 to 39, wherein the single-stream data includes single-channel audio data.

[0181] Example 41 includes the method according to any one of Examples 26 to 39, wherein the single-stream data includes dual-channel audio data.

[0182] Example 42 includes the method according to any one of Examples 26 to 39, wherein the single-stream data includes multi-channel audio data.

[0183] Example 43 includes the method as described in any one of Examples 26 to 42, and further includes outputting audio of single - stream output data at one or more speakers.

[0184] Example 44 includes the method as described in any one of Examples 26 to 43, wherein the single - stream data is provided by one or more microphones.

[0185] Example 45 includes the method as described in any one of Examples 26 to 43, wherein the single - stream data is single - stream data received from a second device via wireless transmission.

[0186] Example 46 includes the method as described in Example 45, wherein the single - stream data is received in combination with a joint learning network, and further includes sending the single - stream output data to a second device via a modem.

[0187] Example 47 includes the method as described in any one of Examples 26 to 46, and is executed in a neural processing unit (NPU).

[0188] Example 48 includes the method as described in any one of Examples 26 to 47, and is executed at one or more processors included in a vehicle.

[0189] Example 49 includes the method as described in any one of Examples 26 to 47, which is executed at one or more processors, and the method is included in an extended reality headset device.

[0190] According to Example 50, a device includes: a memory configured to store instructions; and a processor configured to execute the instructions to perform the method as described in any one of Examples 26 to 49.

[0191] According to Example 51, a computer - readable medium stores instructions that can be executed by a processor to cause the processor to perform the method as described in any one of Examples 26 to 49.

[0192] According to Example 52, a device includes components for performing the method as described in any one of Examples 26 to 49.

[0193] According to Example 53, a non - transitory computer - readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to: detect single - stream data; generate multi - stream enhanced data including one or more modified versions of the single - stream data; process the multi - stream enhanced data to generate multiple output channels; and reduce the multiple output channels to produce single - stream output data.

[0194] According to Example 54, an apparatus includes: components for generating multi-stream enhanced data including one or more modified versions of single-stream data; components for processing the multi-stream enhanced data to generate a plurality of output channels; and components for reducing the plurality of output channels to produce single-stream output data.

[0195] Those skilled in the art will further appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software executed by a processor, or combinations of both. The various illustrative components, blocks, configurations, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends upon the particular application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, and such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure.

[0196] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. The software module may reside in random access memory (RAM), flash memory, read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), registers, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or a user terminal.

[0197] The previous description of the disclosed aspects is provided to enable a person skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features as defined by the appended claims.

Claims

1. A device, comprising: a memory configured to store instructions; and one or more processors configured to: detect single - stream data; generate multi - stream enhanced data including one or more modified versions of the single - stream data; process the multi - stream enhanced data to generate multiple output channels; and reduce the multiple output channels to produce single - stream output data.

2. The device according to claim 1, wherein the one or more processors are further configured to perform one or more first operations on the single - stream data to generate the one or more modified versions of the single - stream data.

3. The device according to claim 2, wherein to reduce the multiple output channels, the one or more processors are further configured to: perform one or more second operations on at least one of the multiple output channels to generate adjusted multi - channel output data, the one or more second operations corresponding to the inverse operations of the one or more first operations; and perform a combination operation on the channels of the adjusted multi - channel output data to generate single - stream output data.

4. The device according to claim 3, wherein the combination operation includes averaging the values of the channels of the adjusted multi - channel output data.

5. The device according to claim 2, wherein the one or more first operations include frequency - domain phase shift.

6. The device according to claim 2, wherein the one or more first operations include frequency - domain group phase shift.

7. The device according to claim 2, wherein the one or more first operations include time - domain shift.

8. The device according to claim 2, wherein the one or more first operations include applying a gain.

9. The device according to claim 1, wherein the multi - stream enhanced data further includes the single - stream data.

10. The device according to claim 1, wherein the one or more processors are configured to use a recurrent network to process the multi - stream enhanced data, the recurrent network processing each stream in the multi - stream enhanced data in parallel and using the same network weights for each stream in the multi - stream enhanced data.

11. The device according to claim 10, wherein the recurrent network is trained using multi - stream enhanced training data.

12. The device according to claim 10, wherein the recurrent network is trained using single - stream training data.

13. The device according to claim 10, wherein the one or more processors are configured to: train the recurrent network using multi - stream enhanced training data; and process the multi - stream enhanced data using the trained recurrent network during inference operations.

14. The device according to claim 10, wherein the one or more processors are configured to: train a recurrent network using single - stream training data; and process the multi - stream enhanced data using the trained recurrent network during inference operations.

15. The device according to claim 1, wherein the single - stream data includes audio data, and wherein the single - stream output data includes a noise - reduced version of the audio data.

16. The device according to claim 1, wherein, the single-stream data includes single-channel audio data.

17. The device according to claim 1, wherein, the single-stream data includes dual-channel audio data.

18. The device according to claim 1, wherein, the single-stream data includes multi-channel audio data.

19. The device according to claim 1, further comprising one or more speakers configured to output the audio of the single-stream output data.

20. The device according to claim 1, further comprising one or more microphones configured to provide the single-stream data.

21. The device according to claim 1, further comprising a modem configured to receive the single-stream data from a second device via wireless transmission.

22. The device according to claim 21, wherein, the single-stream data is received in combination with a federated learning network, and wherein the one or more processors are further configured to send the single-stream output data to the second device via the modem.

23. The device according to claim 1, wherein, the one or more processors are included in a neural processing unit (NPU).

24. The device according to claim 1, wherein, the memory and the one or more processors are included in a vehicle.

25. The device according to claim 1, wherein, the memory and the one or more processors are included in an extended reality headset device.

26. A method, comprising: detecting single-stream data at one or more processors; generating multi-stream enhanced data including one or more modified versions of the single-stream data; processing the multi-stream enhanced data to generate multiple output channels; and reducing the multiple output channels to produce single-stream output data.

27. The method according to claim 26, further comprising: performing one or more first operations on the single-stream data to generate the one or more modified versions of the single-stream data.

28. The method according to claim 27, wherein, reducing the multiple output channels includes: performing one or more second operations on at least one of the multiple output channels to generate adjusted multi-channel output data, the one or more second operations corresponding to the inverse operations of the one or more first operations; and combining the channels of the adjusted multi-channel output data to generate the single-stream output data.

29. A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: detect single-stream data; generate multi-stream enhanced data including one or more modified versions of the single-stream data; process the multi-stream enhanced data to generate multiple output channels; and reduce the multiple output channels to produce single-stream output data.

30. An apparatus, comprising: means for generating multi-stream enhanced data including one or more modified versions of the single-stream data; means for processing the multi-stream enhanced data to generate multiple output channels; and A component for reducing the plurality of output channels to produce single-stream output data.