Media segment prediction for media generation
The challenge of natural speech and facial expressions in media generation is solved by segmenting the input media stream and generating media output clip identifiers using feature extractors, discourse classifiers, and clip matchers, and improve transmission efficiency.
Patent Information
- Application Number
- CN202380072601.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-18
- Filing Date
- 2023-10-04
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art faces challenges in media generation, such as generating natural speech and facial expressions, and the transmission of media content requires a large amount of communication bandwidth.
Efficient media generation and transmission are achieved by segmenting the input media stream into fragments and generating media output fragment identifiers using feature extractors, discourse classifiers, and fragment matchers.
This method can generate natural voice and facial expressions, reduces the need for communication bandwidth and improves the efficiency of media generation.
Smart Images

Figure CN119998869A_ABST
Abstract
Description
CROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from commonly owned U.S. non-provisional patent application No. 18 / 047,572, filed on October 18, 2022, the entire contents of which are expressly incorporated herein by reference. Technical Field
[0002] The present disclosure generally relates to media segmentation and prediction to facilitate media generation. Background Art
[0003] Advances in technology have made computing devices smaller and more powerful, while also increasing the availability and consumption of media. For example, there are now a wide variety of portable personal computing devices, including wireless phones (such as mobile and smart phones, tablet computers, and laptop computers) that are small, lightweight, and easy for users to carry, and can generate and consume media content almost anywhere.
[0004] Although the above-mentioned technological advances have been devoted to improving the transmission of media content, such communications still face challenges. For example, the transmission of media content typically uses a large amount of communication bandwidth.
[0005] The above-mentioned technological advances also include efforts to improve media content generation. However, many challenges associated with media generation still exist. For the purpose of illustration, it may be challenging to generate speech that sounds like natural human speech. Similarly, it is challenging to generate naturally occurring representations of human facial expressions during speech. Summary of the invention
[0006] According to certain aspects, a device includes one or more processors configured to input one or more segments of an input media stream into a feature extractor. The one or more processors are further configured to pass an output of the feature extractor to an utterance classifier to generate at least one representation of at least one utterance category of a plurality of utterance categories. The one or more processors are further configured to pass the output of the feature extractor and the at least one representation to a segment matcher to generate a media output segment identifier.
[0007] According to certain aspects, a method includes inputting one or more segments of an input media stream into a feature extractor. The method also includes passing an output of the feature extractor into an utterance classifier to generate at least one representation of at least one utterance category of a plurality of utterance categories. The method also includes passing the output of the feature extractor and the at least one representation into a segment matcher to generate a media output segment identifier.
[0008] According to certain aspects, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to: input one or more segments of an input media stream into a feature extractor. The instructions, when executed by the one or more processors, also cause the one or more processors to pass an output of the feature extractor into an utterance classifier to generate at least one representation of at least one utterance category in a plurality of utterance categories. The instructions, when executed by the one or more processors, also cause the one or more processors to pass the output of the feature extractor and the at least one representation into a segment comparator to generate a media output segment identifier.
[0009] According to certain aspects, an apparatus includes means for inputting one or more segments of an input media stream into a feature extractor. The apparatus also includes means for passing an output of the feature extractor to an utterance classifier to generate at least one representation of at least one utterance category of a plurality of utterance categories. The apparatus also includes means for passing the output of the feature extractor and the at least one representation to a segment matcher to generate a media output segment identifier.
[0010] Other aspects, advantages, and features of the present disclosure will become apparent upon review of the entire application, including the Brief Description of the Drawings, Detailed Description, and Claims. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] Figure 1 is a block diagram of certain illustrative aspects of a system operable to generate media output segment identifiers based on an input media stream according to some examples of the present disclosure.
[0012] Figure 2 According to some examples of the present disclosure Figure 1 An illustration of certain aspects of a system.
[0013] Figure 3 According to some examples of the present disclosure Figure 1 An illustration of certain aspects of a system.
[0014] Figure 4 According to some examples of the present disclosure Figure 1 An illustration of certain aspects of a system.
[0015] Figure 5 According to some examples of the present disclosure Figure 1 An illustration of certain aspects of a system.
[0016] Figure 6 According to some examples of the present disclosure Figure 1 An illustration of certain aspects of a system.
[0017] Figure 7Some examples of the present disclosure are illustrated by Figure 1 An illustration of particular aspects of the operations performed by a segment matcher of a system to generate and use media output segment identifiers based on an input media stream.
[0018] Figure 8 Some examples of the present disclosure are illustrated by Figure 1 An illustration of particular aspects of the operations performed by a segment matcher of a system to generate and use media output segment identifiers based on an input media stream.
[0019] Fig. 9 Some examples of the present disclosure are illustrated by Figure 1 An illustration of certain aspects of operations performed by a system to generate and use media output segment identifiers based on an input media stream.
[0020] Fig.10 Some examples of the present disclosure are illustrated by Figure 1 An illustration of certain aspects of operations performed by a system to generate and use media output segment identifiers based on an input media stream.
[0021] Fig.11 Some examples of the present disclosure are illustrated by Figure 1 An illustration of certain aspects of operations performed by a system to generate and use media output segment identifiers based on an input media stream.
[0022] Fig.12 Some examples of the present disclosure are illustrated by Figure 1 An illustration of certain aspects of operations performed by a system to generate and use media output segment identifiers based on an input media stream.
[0023] Fig.13 Some examples of the present disclosure are illustrated by Figure 1 An illustration of certain aspects of operations performed by a system to generate and use media output segment identifiers based on an input media stream.
[0024] Fig.14 Some examples of the present disclosure are illustrated by Figure 1 An illustration of certain aspects of operations performed by a system to generate and use media output segment identifiers based on an input media stream.
[0025] Fig.15 Some examples of the present disclosure are illustrated by Figure 1 An illustration of certain aspects of operations performed by a system to generate and use media output segment identifiers based on an input media stream.
[0026] Fig.16Examples of integrated circuits operable to generate and / or use media output segment identifiers based on input media streams in accordance with some examples of the present disclosure are illustrated.
[0027] Fig.17 is an illustration of a mobile device operable to generate and / or use media output segment identifiers based on input media streams according to some examples of the present disclosure.
[0028] Fig.18 is a diagram of a headset operable to generate and / or use media output segment identifiers based on input media streams according to some examples of the present disclosure.
[0029] Fig.19 is a diagram of a wearable electronic device operable to generate and / or use media output segment identifiers based on input media streams according to some examples of the present disclosure.
[0030] Fig. 20 is a diagram of a voice-controlled speaker system operable to generate and / or use media output segment identifiers based on input media streams according to some examples of the present disclosure.
[0031] Fig.21 is a diagram of a camera operable to generate and / or use media output segment identifiers based on input media streams according to some examples of the present disclosure.
[0032] Fig. 22 is a diagram of a headset (such as a virtual reality, mixed reality, or augmented reality headset) operable to generate and / or use media output segment identifiers based on input media streams according to some examples of the present disclosure.
[0033] Fig.23 is an illustration of a first example of a vehicle operable to generate and / or use media output segment identifiers based on input media streams according to some examples of the present disclosure.
[0034] Fig.24 is a diagram of a second example of a vehicle operable to generate and / or use media output segment identifiers based on input media streams according to some examples of the present disclosure.
[0035] Fig.25 According to some examples of the present disclosure, Figure 1 A diagram of a specific implementation of a method for generating media output segment identifiers based on an input media stream performed by a device.
[0036] Fig.26 is a block diagram of a specific illustrative example of a device operable to generate and / or use media output segment identifiers based on input media streams according to some examples of the present disclosure. DETAILED DESCRIPTION
[0037] Humans are particularly good at recognizing the facial expressions and voice sounds of others. Nearly realistic depictions of humans in the media (such as in some computer-generated graphics) that are slightly unnatural can lead to the so-called "uncanny valley" effect, which can be disturbing or even offensive to people consuming such media. Even small differences between human-like representations in the media (e.g., faces or voices) and natural (i.e., real) depictions of humans can cause such discomfort. In addition, computer-generated voices can be more difficult to understand due to unnatural rhythms and emphasis, as well as a lack of the natural variations in sound present in human speech.
[0038] Systems and methods for media segmentation and prediction for facilitating media generation are disclosed. For example, according to certain aspects, at least audio content of a media stream is segmented and processed to determine media output segment identifiers. The media output segment identifiers include data sufficient to select specific media segments for output. In certain aspects, the media segments correspond to segments of pre-recorded natural human speech. In some implementations, the media segments include sounds representing one or more utterances. In some implementations, the media segments also include or alternatively include one or more images depicting human facial movements associated with the generation of the one or more utterances.
[0039] As a simple non-limiting example, an input media stream may include a speech sound of a first human speaker. In this example, each group of one or more speech sounds (e.g., each phoneme, multiple groups of phonemes, or other segments of speech) is processed to generate a corresponding media output segment identifier. In this example, the media output segment identifier may include an array or vector of values identifying one or more media segments, which correspond to pre-recorded media associated with another human speaker generating the same or similar speech sound. The media output segment identifier may be used to generate an output media stream based on pre-recorded media. For example, the media output segment identifier may be sent from a first device receiving the input media stream to a second device. The second device may generate an output media stream containing realistic speech sounds based on the speech sounds of the input audio stream. Additionally or alternatively, the output media stream generated by the second device may include a realistic depiction of facial movements (e.g., lip movements) of a person emitting speech sounds based on the speech sounds of the input audio stream.
[0040] The pre-recorded media used to generate the output media stream may depict the same person speaking to generate the input media stream; however, this is not required. For example, during a telephone conversation in which several people are communicating and expect to hear a familiar voice, it may be useful to use pre-recorded media of the same person speaking to generate the input media stream. However, it may also be useful to use pre-recorded media of a different person than the person speaking to generate the input media stream. For example, storing data associated with the pre-recorded media used to generate the output media stream uses a large amount of memory resources. Therefore, for use cases in which having a familiar voice is not important, pre-recorded media content from a small number of people (e.g., one, two, five, or another number of people) can be stored and used to generate output media streams for any number of people providing input media streams. For the purpose of illustration, a set of pre-recorded media content can be stored and used to generate an output media stream, regardless of who provided the input media stream.
[0041] Another example where using pre-recorded media of a first person to generate an output media stream based on the speech of a second person may be beneficial is where it is desired to change speech characteristics between an input media stream and an output media stream. For purposes of illustration, if the second person's accent is difficult to understand for some listeners, the second person's speech may be processed and used to generate an output media stream including the speech sounds of the first person, where the first person's accent is different and more easily understood. Other examples where using pre-recorded media of a first person to generate an output media stream based on the speech of a second person may be beneficial include voice alteration or anonymization.
[0042] In a communication situation, the media output segment identifier may be determined on the transmitter side or the receiver side of the communication channel. For example, a sending device may receive an input media stream and determine a media output segment identifier for each segment of the input media stream. In this example, the sending device may transmit data representing the media output segment identifier to a receiving device, and the receiving device may generate an output media stream based on the media output segment identifier. Generally speaking, the media output segment identifier is significantly smaller than the media content processed to determine the media output segment 166 (e.g., represented by fewer bits than the media content). Thus, communication resources (e.g., bandwidth, channel time, etc.) may be saved by sending the media output segment identifier instead of the input media stream it represents.
[0043] As another example, a sending device may transmit data representing an input media stream using any available communication scheme (e.g., using voice over Internet Protocol). In this example, the input media stream may be sent in a sequence of packets. In this example, a receiving device generates an output media stream based on the received packets; however, occasionally, the receiving device may not receive one or more packets, or one or more packets may be damaged when received. In this example, the missing or damaged packets cause gaps in the output media stream. The receiving device may fill the gaps by predicting one or more media output segment identifiers corresponding to the portions of the input media stream associated with the missing or damaged packets. Filling the gaps in the output media stream based on the predicted media output segment identifiers may generate a more natural sounding output than other packet loss concealment processes.
[0044] Certain aspects of the present disclosure are described below with reference to the accompanying drawings. In this specification, common features are designated by common reference numerals. As used herein, various terms are used only for the purpose of describing a particular implementation and are not intended to limit the implementation. For example, the singular forms "a", "an", and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. In addition, some features described herein are singular in some implementations and plural in other implementations. For the purpose of illustration, Figure 1 Depicts a system including one or more processors ( Figure 1 104), indicating that in some implementations, the device 102 includes a single processor 104, while in other implementations, the device 102 includes multiple processors 104. For ease of reference herein, such features are generally introduced as "one or more" features and are subsequently referred to in the singular or optionally in the plural (as indicated by "(s)" in the feature name) unless aspects are being described that relate to multiple of the features.
[0045] As used herein, the term "comprise" may be used interchangeably with "include". Additionally, the term "wherein" may be used interchangeably with "wherein". As used herein, "exemplary" indicates an example, a specific implementation, and / or an aspect, and should not be interpreted as limiting or indicating a preferred or preferred specific implementation. As used herein, ordinal terms (e.g., "first", "second", "third", etc.) used to modify elements (such as structures, components, operations, etc.) do not themselves indicate any priority or order of the element relative to another element, but only distinguish the element from another element with the same name (but using ordinal terms). As used herein, the term "set" refers to one or more specific elements in a specific element, and the term "plurality" refers to a plurality of (e.g., two or more) specific elements.
[0046] As used herein, "coupling" may include "communicatively coupled", "electrically coupled" or "physically coupled", and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. As illustrative, non-limiting examples, two electrically coupled devices (or components) may be included in the same device, or in different devices, and may be connected via electronic devices, one or more connectors, or inductive coupling. In some specific implementations, two devices (or components) that are communicatively coupled (such as electrically connected) may directly or indirectly transmit and receive signals (e.g., digital signals or analog signals) via one or more wires, buses, networks, etc. As used herein, "direct coupling" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled or physically coupled) without intermediate components.
[0047] In the present disclosure, terms such as "determine", "calculate", "estimate", "shift", "adjust", etc. may be used to describe how to perform one or more operations. It should be noted that such terms should not be interpreted as limiting, and other technologies may be used to perform similar operations. In addition, as mentioned herein, "generate", "calculate", "estimate", "use", "select", "access", and "determine" may be used interchangeably. For example, "generating", "calculating", "estimating" or "determining" a parameter (or signal) may refer to actively generating, estimating, calculating or determining the parameter (or signal), or may refer to using, selecting or accessing a parameter (or signal) that has already been generated (such as by another component or device).
[0048] refer to Figure 1 , illustrates certain illustrative aspects of a system 100 configured to generate one or more media output segment identifiers 162 based on an input media stream 120. Additionally, in Figure 1 In the example illustrated in , system 100 is configured to generate output media stream 140 and / or output media stream 180 based on media output segment identifier 162 .
[0049] System 100 includes a device 102 that is coupled to or includes one or more sources 122 of media content of an input media stream 120. For example, source 122 may include microphone 126, camera 132, communication channel 124, or a combination thereof. Figure 1, the sources 122 are external to the device 102 and coupled to the device 102 via the input interface 106; however, in other examples, one or more of the sources 122 are components of the device 102. For purposes of illustration, the sources 122 may include a media engine (e.g., a game engine or an extended reality engine) of the device 102 that generates the input media stream 120 based on instructions executed by one or more processors 104 of the device 102.
[0050] The input media stream 120 includes at least data representing the speech 128 of the person 130. For example, when the source 122 includes the microphone 126, the microphone 126 may generate a signal based on the sound of the speech 128. When the source 122 includes the camera 132, the input media stream 120 may also include one or more images (e.g., video frames) depicting the person 130. When the source 122 includes the communication channel 124, the input media stream 120 may include transmission data representing the speech 128, such as a plurality of data packets used to encode the speech 128. The communication channel 124 may include or correspond to a wired connection between two or more devices, a wireless connection between two or more devices, or both.
[0051] exist Figure 1 In the example of , the device 102 is configured to process the input media stream 120 to determine the media output segment identifiers 162. Each media output segment identifier 162 indicates one or more media segments to be included in the output media stream 140 to represent one or more corresponding segments of the input media stream 120. As an example, the input media stream 120 can be parsed into segments, each segment corresponding to one or more phonemes or other utterance segments of speech. In this example, each media output segment identifier 162 corresponds to a media segment that includes one or more phonemes or other utterance segments of speech that are similar to the one or more phonemes or other utterance segments of the speech 128 of the input media stream 120, as further explained below.
[0052] exist Figure 1 , the device 102 includes an input interface 106, an output interface 112, a processor 104, a memory 108, and a modem 110. The input interface 106 is coupled to the processor 104 and is configured to be coupled to one or more of the sources 122. For example, the input interface 106 is configured to receive a microphone output from a microphone 126 and provide the microphone output as an input media stream 120 to the processor 104.
[0053] The output interface 112 is coupled to the processor 104 and is configured to couple to one or more output devices, such as one or more speakers 142, one or more display devices 146, etc. The output interface 112 is configured to receive data representing an output media stream 140 from the processor 104 and transmit the output media stream 140 corresponding to the data to the output device.
[0054] The processor 104 is configured to receive the input media stream 120 and determine the media output segment identifier 162 based on the input media stream 120. Figure 1 In the example illustrated in , the processor 104 includes a media segment identifier 160, a segment mapper 164, and a media stream assembler 168. The media segment identifier is configured to process the input media stream 120 and determine a media output segment identifier 162, the segment mapper 164 is configured to determine one or more media output segments 166 based on the media output identifier 162, and the media stream assembler 168 is configured to generate an output media stream based on the media output segments 166, as described in further detail below. Each of the media segment identifier 160, the segment mapper 164, and the media stream assembler 168 includes or corresponds to instructions that can be executed by the processor 104 to perform various operations described herein.
[0055] In some implementations, the media segment identifier 160 includes one or more trained models. Examples of trained models include machine learning models such as neural networks, adaptive neuro-fuzzy inference systems, support vector machines, decision trees, regression models, Bayesian models, Boltzmann machines, or sets, variants, or other combinations thereof. Variants of decision trees include, for example, but not limited to, random forests, boosted decision trees, and the like. Variants of neural networks include, for example, but not limited to, transformers, self-attention networks, convolutional neural networks, deep neural networks, deep belief networks, and the like.
[0056] In a particular implementation, the media segment identifier 160 is configured to parse the input media stream 120 into segments. Each segment represents a portion of the input media stream 120 that can be mapped to a media output segment 166. As a non-limiting example, each segment can include silence, background noise, one or more phonemes or other utterances in the speech 128, etc. In some implementations, the parsing of the input media stream is based on the content of the speech 128 (e.g., the sounds in the speech). In such implementations, different segments can represent different durations of the input media stream 120. For purposes of illustration, a first segment can correspond to 50 milliseconds of the input media stream 120, and a second segment can correspond to 160 milliseconds of the input media stream 120. In one experiment, a sample input media stream 120 of English speech with a total duration of approximately 2.5 hours was processed to generate approximately 96,000 segments, and the average segment represented approximately 100 milliseconds of the input media stream 120. The specified duration represented by a segment may vary from implementation to implementation based on, for example, the content of speech 128, the language of speech 128, and how one or more models of media segment identifier 160 are trained. Furthermore, although segments of variable duration are described herein, in some implementations, segments of fixed duration may be used.
[0057] In certain implementations, after the media segment identifier 160 determines a segment, the segment (and optionally one or more nearby segments) may be input to a feature extractor of the media segment identifier 160. The feature extractor is configured to generate feature data (e.g., a feature vector, a feature array, a feature map, a set of values representing speech parameters, etc.) representing various aspects of the segment. In some implementations, the feature extractor is a time-dynamic feature extractor. For example, feature data associated with a particular segment may be affected by the content of one or more segments preceding the particular segment in the input media stream 120, or may be affected by the content of one or more segments following the particular segment in the input media stream 120, or both. Examples of trained models that may be used to perform time-dynamic feature extraction include (but are not limited to) recurrent neural networks (RNNs) (e.g., neural networks having one or more recurrent layers, one or more long short-term memory (LSTM) layers, one or more gated recurrent unit (GRU) layers, etc.), recurrent convolutional neural networks (RCNNs), self-attention networks (e.g., transformers), other machine learning models suitable for processing time series data in a time-dynamic manner, or variants, collections, or combinations thereof.
[0058] According to certain aspects, the input media stream 120 includes a sequence of data frames of content from a source 122, and the media segment identifier 160 generates a sequence of media output segment identifiers 162 based on the input media stream 120. Each media output segment identifier 162 is generated based on one or more data frames of the input media stream 120. In addition, the number of data frames of the input media stream 120 used to generate a single media output content identifier 162 may vary from one media output segment identifier 162 to another. As a non-limiting example, each data frame may represent 25 milliseconds of audio content, and each media output segment identifier 162 may represent 25 milliseconds to several hundred milliseconds of audio content. Therefore, the feature extractor can be viewed as a sequence-to-sequence feature extractor that is configured to generate a sequence of multiple sets of feature data (e.g., feature vectors, feature arrays, feature maps, or multiple sets of values of speech parameters) based on the content sequence of the input media stream 120. From this perspective, the sequence-to-sequence feature extractor acquires data at a first rate (e.g., one data frame every 25 milliseconds) and outputs data at a second rate (e.g., one media output segment identifier 162 every 25 milliseconds to N×25 milliseconds, where N is an integer greater than or equal to 1), where on average, the first rate is not equal to the second rate.
[0059] As an example, the feature extractor may generate a media output segment identifier 162 for each phoneme, group of phonemes, or some other unit of speech 128. As used herein, the term "phoneme" is broadly used to refer to a unit of sound that distinguishes one word from another in a particular language. Although various sources have proposed or agreed upon specific lists of "phonemes" that are useful for academic purposes, no such list is specifically referenced by the term phoneme as used herein. In a specific implementation of using a trained model to segment the speech 128, the specific phonemes or other units of speech used to distinguish the segments may be based on the training of the model. As an example, in the experiment mentioned above, in which approximately 2.5 hours of speech were segmented into approximately 96,000 segments, the trained model performing the segmentation was trained to group diphones into one segment, where a diphone refers to a pair of consecutive phonemes.
[0060] In addition to the feature extractor, the media segment identifier 160 also includes other components configured to determine the media output segment identifier 162 based on the feature data output by the feature extractor. In some implementations, the media output segment identifier 162 includes a plurality of elements of a vector, an array, or another data structure, and the plurality of elements includes one element for each media output segment 166 that can be indicated by the media output segment identifier 162. For example, if the segment mapper 164 is able to access or generate 2000 different media output segments 166, then in this implementation, the media output segment identifier 162 may include a vector having 2000 elements, each of the 2000 elements corresponding to a respective one of the 2000 media output segments 166.
[0061] In some such implementations, each media output segment identifier 162 is a one-hot vector or a one-hot array (or an encoded version of a one-hot vector or a one-hot array). For purposes of illustration, continuing with the above example, if segment mapper 164 is able to access or generate 2000 different media output segments 166 and media output segment identifiers 162 are one-hot vectors, then 1999 elements in media output segment identifiers 162 will have a first value (e.g., a value of 0) indicating that the media output segments 166 corresponding to those elements are not indicated by media output segment identifiers 162, and 1 element in media output segment identifiers 162 will have a second value (e.g., a value of 1) indicating that the media output segment 166 corresponding to that element is indicated by media output segment identifier 162.
[0062] In some implementations, the media output segment identifier 162 is not a one-hot vector. For example, in some such implementations, the media output segment identifier 162 is a vector or array including a plurality of elements having non-zero values. For purposes of illustration, the media output segment identifier 162 may include a likelihood value for each element of the array or vector, indicating the likelihood that the corresponding media output segment 166 corresponds to a segment of the input media stream 120 represented by feature data from the feature extractor. In some such implementations, the media output segment identifier 162 does not include a likelihood value for each element. For example, one or more thresholds may be used to filter the likelihood values so that only particularly relevant likelihood values are included in the media output segment identifier 162, and other likelihood values are cleared. For purposes of illustration, the media output segment identifier 162 may include the highest likelihood values of the first two, first three, first five, or other number, and the remaining elements of the media output segment identifier 162 may include zero values. As another illustrative example, media output segment identifiers 162 may include each likelihood value that exceeds a threshold (eg, likelihood values of 0.1, 0.2, 0.5, or some other value), and the remaining elements of media output segment identifiers 162 may include zero values.
[0063] In several of the above examples, the segment mapper 164 is described as being able to access or generate 2000 different media output segments 166. These examples are provided for illustrative purposes only and not for limitation. The specific number of different media output segments 166 that the segment mapper 164 can access or generate may vary from implementation to implementation, depending on the configuration of the media segment identifier 160, the segment mapper 164, or other factors. As an illustrative example, in the experiment mentioned above, in which approximately 2.5 hours of audio are processed, each identified segment (e.g., each identified diphone) in the entire 2.5 hours of audio is stored as a different media output segment 166. Therefore, the segment mapper 164 in this experiment is able to access or generate approximately 96,000 media output segments 166, and the media output segment identifier 162 is a vector comprising approximately 96,000 values. Since there are many common sounds in English, it is expected that many media output segments 166 of this group of media output segments are very similar. Thus, the number of media output segments in the set (and, accordingly, the dimensionality of media output segment identifiers 162 ) may be reduced by performing additional processing to identify duplicate or near-duplicate media output segments 166 .
[0064] In some embodiments, the segment mapper 164 is capable of generating or accessing a set of media output segments 166, which includes each common phoneme in a specific language, each common continuous phoneme group in the specific language, each phoneme actually used in the specific language, or each continuous phoneme group actually used in the specific language. In some embodiments, the group of media output segments 166 includes at least a representative set of common phonemes or continuous groups of phonemes in the specific language. For example, the group of media output segments 166 can be obtained from a recording that is considered to have sufficient duration and diversity to correspond to a representative set of media content. For illustrative purposes, in the experiment mentioned above, a 2.5-hour English speech recording is considered sufficient to provide a feasible set of representative media output segments 166.
[0065] Although retaining media output segments that are very similar to one another (rather than deduplicating the set of media output segments) increases the dimensionality of media output segment identifiers 162, retaining at least some of the media output segments that are similar to one another can facilitate generating more natural sounding speech in output media stream 140. Therefore, for implementations where reducing computing resources used to generate output media stream 140 (such as memory required to store output segment data 114 representing the media output segments, processing time and power to calculate media output segment identifiers 162, etc.) takes priority over optimizing the natural sounding quality of the output speech, the set of media output segments can be processed to reduce the number of duplicate or nearly duplicate media output segments. Conversely, for implementations where generating natural sounding speech in output media stream 140 has a higher priority, a set of media output segments that includes some nearly duplicate media output segments can be used to generate speech with more natural pronunciation variations.
[0066] In embodiments where media segment identifiers 160 include one or more models trained to determine media output segment identifiers 162 and media output segment identifiers 162 are high-dimensional (e.g., having thousands or tens of thousands of elements), training the one or more models may be challenging. For example, in the experiments mentioned above, media output segment identifiers 162 had approximately 96,000 elements. The trained model used to generate media output segment identifiers 162 may be considered a classifier that indicates a category corresponding to one or more of media output segments 166. For purposes of illustration, when media output segment identifiers 162 are one-hot vectors, a single non-zero value of media output segment identifiers 162 indicates a category corresponding to media output segment 166. Training a classifier to reliably select a single element from approximately 96,000 categories, where some of the categories may be similar or nearly identical, is a challenging training scenario.
[0067] This training challenge can be reduced by dividing the inference process into hierarchical stages using a multi-stage model. For example, the 96,000 categories are grouped into supersets ("utterance categories"). In this example, the utterance classifier of the trained model determines the utterance category associated with a particular set of feature data representing a segment of the input media stream 120. The utterance category and the feature data are provided as input to the segment matcher of the trained model to generate a media output segment identifier 162. In this hierarchical approach, the utterance category is provided to the segment matcher along with the feature data so that the analysis performed by the segment matcher is biased towards (e.g., the analysis performed by the segment matcher is weighted) a result that favors assigning a media output segment identifier 162 to indicate a media output segment 166 that is within the indicated utterance category.
[0068] exist Figure 1 14, one or more of the media output segment identifiers 162 are provided to the segment mapper 164. Additionally or alternatively, one or more of the media output segment identifiers 162 are provided to the modem 110 for transmission to one or more other devices (e.g., the device 152). For example, in the case where the device 102 is to generate an output media stream 140, the media output segment identifiers 162 are provided to the segment mapper 164, and processing at the segment mapper 164 and the media stream assembler 168 results in the generation of data representing the output media stream 140. For purposes of illustration, when the input media stream 120 is received from the communication channel 124, the device 102 may provide the output media stream 140 to the speaker 142, the display device 146, or both. In another example, when the media output segment identifiers 162 are provided to the segment mapper 164 is when the device 102 is receiving speech 128 from the microphone 126 and the device 102 is to perform noise reduction or voice modification operations (such as changing the accent). In this example, data representing output media stream 140 may be generated by media stream assembler 168 and provided to modem 110 for transmission to device 152 via communication channel 150. In the event that device 152 is to generate output media stream 180 based on input media stream 120, media output segment identifier 162 is provided to modem 110 for transmission to device 152 via communication channel 150. Communication channel 150 may include or correspond to a wired connection between two or more devices, a wireless connection between two or more devices, or both.
[0069] When the media output segment identifiers 162 are provided to the segment mapper 164, the segment mapper 164 generates or retrieves the media output segments 166 corresponding to each media output segment identifier 162 and provides the media output segments 166 to the media stream assembler 168. For example, for a particular media output segment identifier 162, the segment mapper 164 may access the corresponding output segment data 114 from the memory 108. The output segment data 114 may include the corresponding media output segment 166, or the output segment data 114 may include data used by the segment mapper 164 to generate the corresponding media output segment 166.
[0070] In some implementations, segment mapper 164 retrieves media output segments 166 from a database or other data structure in memory 108 (e.g., from output segment data 114). In such implementations, each media output segment identifier 162 includes or corresponds to a unique identifier of a corresponding media output segment stored in memory 108, or segment mapper 164 determines the unique identifier of the corresponding media output segment stored in memory 108 based on the media output segment identifier 162.
[0071] In some implementations, the segment mapper 164 generates a media output segment 166 corresponding to a particular media output segment identifier 162. For example, in such implementations, the output segment data 114 may include a set of fixed weights (e.g., link weights) that are used by the segment mapper 164 to generate the media output segment 166. In this example, the output of the segment mapper 164 includes a set of elements corresponding to media parameters. For example, the output of the segment mapper 164 may include a set of elements representing pulse code modulation (PCM) sample values, and each media output segment identifier 162 is a one-hot vector or a one-hot array. In some such implementations, the layer of the trained model that generates the media output segment identifier 162 may be considered as an embedding layer connected to the output layer represented by the segment mapper 164. The link weights between the nodes of the embedding layer and the nodes of the output layer are predetermined (e.g., before training the model) and are configured to cause the output layer to generate media parameters representing the media output segment 166 corresponding to the media output segment identifier 162.
[0072] Media stream assembler 168 assembles media output segments 166 from segment mapper 164 to generate data representing output media stream 140. For purposes of illustration, media stream assembler 168 concatenates or otherwise arranges media output segments 166 to form an ordered sequence of media output segments 166 for playback. In some examples, the data representing output media stream 140 is provided to output interface 112 for playback at speaker 142 as voice 144, to display device 146 for playback as video 148, or both. In the same example or a different example, this data representing output media stream 140 may be provided to modem 110 for transmission to device 152 via communication channel 150.
[0073] When media output segment identifier 162 (rather than output media stream 140) is provided to modem 110 for transmission to device 152, device 152 may generate output media stream 180 based on media output segment identifier 162. For example, in Figure 1 1 , device 152 includes a modem 170, a segment mapper 172, a media stream assembler 174, and an output interface 176. When device 152 receives a media output segment identifier 162, modem 170 of device 152 may provide the media output segment identifier 162 to segment mapper 172. Segment mapper 172 operates in the same manner as segment mapper 164 of device 102. For example, segment mapper 172 generates or accesses media output segments 166 corresponding to media output segment identifiers 162. Media stream assembler 174 assembles the media output segments 166 from segment mapper 172 in the same manner as described for media stream assembler 168 to generate data representing an output media stream 180, and the resulting output media stream 180 is output by device 152 via output interface 176.
[0074] In some implementations, the media output segments 166 available to the segment mapper 172 of the device 152 are different from the media output segments 166 available to the segment mapper 164 of the device 102. For purposes of illustration, the segment mapper 172 may have access to a set of media output segments 166 representing speech of a first speaker (e.g., a male), and the segment mapper 172 may have access to a set of media output segments 166 representing speech of a second speaker (e.g., a female). Regardless of whether the segment mappers 164, 172 have access to the same set of media output segments 166, the media output segments 166 available to each segment mapper 164, 172 are mapped so that the same phoneme or other unit of speech corresponds to the same media output segment identifier 162. For example, a particular media output segment identifier 162 may correspond to the sound of "ah", and both segment mappers 164, 172 map the particular media output segment identifier 162 to the sound of "ah" of their respective available media output segments 166.
[0075] In some implementations, the media segment identifier 160 can be used to predict a media output segment identifier 162 for a portion of the input media stream 120 that is not available. For example, when receiving data via the communication channel 124, sometimes a packet or other data unit may be lost or damaged. In such a case, the content (e.g., media) of the packet or other data unit is not available in the input media stream 120. Since the media segment identifier 160 is configured to generate a stream of media output segment identifiers 162 based on the input media stream 120, the media segment identifier 160 can be used to predict the media output segment identifier 162 corresponding to the missing content. During playback of the output media stream 140, the predicted media output segment identifier 162 can be used to replace the missing content.
[0076] Although the above description focuses primarily on examples where media output segments 166 represent audio data, in some implementations, media output segments 166 may include or correspond to image or video data. For example, media output segments 166 may include one or more images depicting the face of a person uttering a particular sound (e.g., one or more phonemes or other utterances). In this example, each media output segment identifier 162 maps to a corresponding set of one or more images (e.g., mapped to a corresponding media output segment 166). When the input media stream 120 represents a particular sound, the media segment identifier 160 generates a media output segment identifier 162 that maps to a media output segment 166 representing the person uttering the particular sound. A set of one or more images of the media output segment 166 may be combined with other images corresponding to other media output segments 166 to generate a sequence of image frames of the output media stream 140. The sequence of image frames provides a realistic depiction of a person speaking a series of sounds corresponding to the input media stream 120. Because the sequence of image frames is composed of actual pre-recorded images of a person making similar sounds (although possibly in a different order because different words may have been spoken), the sequence of image frames of output media stream 140 avoids the uncanny valley problem associated with entirely computer-generated video.
[0077] Thus, the system 100 facilitates the generation of audio, video, or both of media including human speech in a manner that is natural in sound, appearance, or both. The system 100 also facilitates low bit rate voice data communications and outputs natural sounding speech on a recipient device. The system 100 is also capable of modifying the audio characteristics of an input media stream, such as noise reduction, voice modification, anonymization, etc.
[0078] Figure 2 According to some examples of the present disclosure Figure 1 100. Specifically, Figure 2 Examples of input media stream 120, media segment identifier 160, media output segment identifier 162, segment mapper 164, media stream assembler 168, and output media stream 140 are illustrated.
[0079] exist Figure 2 In the example illustrated in , the media segment identifier 160 includes a feature extractor 202, an utterance classifier 206, and a segment matcher 210. The feature extractor 202 is coupled to the utterance classifier 206 and the segment matcher 210, and is configured to generate feature data 204 representing a portion of the input media stream 120. For example, one or more segments of the input media stream 120 may be input into the feature extractor 202 to generate feature data 204 corresponding to the one or more segments of the input media stream 120. As shown in FIG. Figure 1As explained, input media stream 120 may include data representing speech, and optionally may include data representing other content such as other sounds and / or images.
[0080] In some implementations, the feature extractor 202 is a time-dynamic feature extractor. In such implementations, each set of feature data 204 generated by the feature extractor 202 represents the content of one or more segments of the input media stream 120 in its temporal context relative to one or more other segments of the input media stream 120. For purposes of illustration, a particular set of feature data 204 may represent the content of two segments of the input media stream 120 in the context of one or more previous segments of the input media stream 120, in the context of one or more subsequent segments of the input media stream 120, or in the context of both. In some such implementations, the time-dynamic feature extractor is a segment-to-segment feature extractor. In such implementations, each set of feature data 204 represents a variable number of segments of the input media stream 120.
[0081] In some implementations, feature extractor 202 includes or corresponds to an audio processor configured to extract values of specific speech parameters. In such implementations, feature data 204 includes data representing speech parameters. Non-limiting examples of such speech parameters include speech energy, pitch, duration, spectral representation, spectral envelope, one or more other measurable or quantifiable parameters representing speech, or combinations thereof. In some implementations, Figure 6 The feature extractor of includes one or more trained models and an audio processor. In such a specific implementation, the feature data 204 may include feature vectors and speech parameters.
[0082] In certain implementations, the feature data 204 includes a vector or array that encodes information about a corresponding segment of the input media stream 120. When the feature data 204 corresponds to a feature vector or feature array, the feature data 204 typically does not include human interpretable content. For example, the feature data 204 may encode a segment of the input media stream 120 into a high-dimensional feature space for further processing by the utterance classifier 206, the segment matcher 210, or both of the media segment identifier 160.
[0083] exist Figure 2In the example illustrated in , the feature data 204 is passed to a speech classifier 206. The speech classifier 206 is configured (and trained) to determine a speech class 208 for a segment of the input media stream 120 based on the feature data 204. Optionally, the speech class 208 may also be determined based on feedback from classifications associated with a previous set of feature data 204. The speech class 208 associated with the segment is one of a plurality of speech classes, the specific number and nature of which depends on the configuration of the media segment identifier 160. As an example, each speech class may represent a phoneme or a group of phonemes. In this example, the particular speech class 208 indicated by the speech classifier 206 is an estimate of which phonemes are represented by the segment of the input media stream 120 represented by the feature data 204. In other examples, the speech class represents a diphone or other artificial sound grouping.
[0084] exist Figure 2 In the example illustrated in , the feature data 204 and the utterance category 208 associated with the feature data 204 are passed to the segment matcher 210. The segment matcher 210 is configured to determine the media output segment identifier 162 based on the feature data 204 and the utterance category 208. As described above, in some implementations, the media output segment identifier 162 includes a high-dimensional vector or array (e.g., a vector including thousands or tens of thousands of elements). In such implementations, passing the utterance category 208 along with the feature data 204 to the segment matcher 210 improves the accuracy with which the segment matcher 210 generates the media output segment identifier 162. For example, the utterance category 208 may weight the analysis of the segment matcher 210 to favor the selection of the media output segment identifier 162 associated with the utterance category 208. As another example, the search space examined by the segment matcher 210 may be limited based on the utterance category 208. For purposes of illustration, in some implementations, the segment matcher 210 compares the speech parameters of the feature data 204 with the speech parameters of the candidate segments to determine the media output segment identifier 162. In such implementations, the candidate segments can be selected based on the utterance category 208 such that the candidate segments represent a subset of available output segments that include sounds corresponding to the utterance category 208.
[0085] In certain implementations, the feature extractor 202, the speech classifier 206, and the fragment matcher 210 each include or correspond to a trained model. In some such implementations, the feature extractor 202, the speech classifier 206, and the fragment matcher 210 are trained together. For example, the feature extractor 202, the speech classifier 206, and the fragment matcher 210 may be trained together at the same time. As another example, the feature extractor 202, the speech classifier 206, and the fragment matcher 210 may be trained in an iterative, sequential manner. For purposes of illustration, in a first iteration, one or more first models are trained to generate outputs based on inputs, and in a second iteration, one or more second models are trained to generate inputs or outputs using one or more first models. In yet another example, the feature extractor 202 and the speech classifier 206 may be trained independently of the fragment matcher 210. In this example, the feature extractor 202 and the utterance classifier 206 are trained to generate feature data and utterance categories 208, and are used to train one or more different segment matchers 210 (such as two segment matchers 210), which are configured to generate different media output segment identifiers 162 that map to different groups of media output segments. For illustrative purposes, one segment matcher 210 may map to high-dimensional media output segment identifiers associated with a large set of media output segments, and another segment matcher 210 may map to low-dimensional media output segment identifiers. In this illustrative example, the high-dimensional media output segment identifiers may be used in situations where it is desired to generate a high-fidelity and natural-sounding output media stream 140, and the low-dimensional media output segment identifiers may be used in situations where it is desired to save processing resources or communication resources.
[0086] The segment matcher 210 is configured to generate the media output segment identifier 162, as shown in FIG. Figure 1 As depicted, the media output segment identifier is provided to segment mapper 164. Segment mapper 164 uses media output segment identifier 162 to access or generate one or more media output segments, which are provided to media stream assembler 168 to generate output media stream 140.
[0087] like Figure 2 As illustrated in , one benefit of using a multi-stage approach to generate media output segment identifiers 162 is that training of one or more models of media segment identifiers 160 may be simplified and / or improved. For example, in some implementations, media output segment identifiers 162 are high-dimensional (e.g., thousands or tens of thousands of elements) one-hot vectors, where several elements may correspond to similar media output segments. In such implementations, monolithically training two or more models of media segment identifiers 160 may be challenging. Figure 2The phased approach illustrated in enables the model of the media segment identifier 160 to be trained separately.
[0088] Additionally, the multi-stage approach facilitates the use of a transfer learning approach, wherein a model of the media segment identifier 160 may be reused to train a new segment matcher 210 for use with a different set of media output data. For purposes of illustration, after initially training a first instance of the feature extractor 202, the utterance classifier 206, and the segment matcher 210, the feature extractor 202 and the utterance classifier 206 may be used to facilitate the training of a second instance of the segment matcher 210. In this illustrative example, the first instance of the segment matcher 210 may be configured to generate media output segment identifiers 162 that map to a first set of media output segments, and the second instance of the segment matcher 210 may be configured to generate media output segment identifiers 162 that map to a second set of media output segments. Thus, computational resources are conserved when training a model of the media segment identifier 160 based on different sets of media output segments.
[0089] In some implementations, the media segment identifier 160, the segment mapper 164, or both are configured to receive one or more constraints 250 as input and cause the output media stream 140 to be generated based on the constraints 250. The constraints 250 may be user-specified (e.g., based on Figure 1 For purposes of illustration, constraints 250 may indicate a particular type of voice or audio modification to be performed to generate output media stream 140. As one example, constraints 250 may indicate desired speech characteristics of the output speech (e.g., gender, accent, voiced or unvoiced speech, etc.), in which case media segment identifier 160 and / or segment mapper 164 operate to select media output segments 166 to have the desired speech characteristics. As another example, constraints 250 may indicate the voice of a particular speaker of the output speech (e.g., the voice of a particular person), in which case media segment identifier 160 and / or segment mapper 164 operate to select media output segments 166 that include speech content of the particular person.
[0090] Figure 3 According to some examples of the present disclosure Figure 1 Specifically, Figure 3 Examples of input media stream 120, media segment identifier 160, media output segment identifier 162, segment mapper 164, media stream assembler 168, and output media stream 140 are illustrated. Figure 3 Details of the feature extractor 202, utterance classifier 206, and segment matcher 210 are further illustrated according to specific implementations.
[0091] exist Figure 3 In the example illustrated in , the feature extractor 202 includes one or more convolutional networks 302 and one or more self-attention networks 304. As an example, the convolutional network 302 includes or corresponds to one or more two-dimensional (2D) convolutional layers. In a specific implementation, the self-attention network 304 includes one or more multi-head (e.g., multi-head (e.g., ×6) attention networks, such as one or more transformer networks.
[0092] exist Figure 3 In the example illustrated in , the utterance classifier 206 includes one or more embedding networks 306 and one or more self-attention networks 308. As an example, the embedding network 306 includes or corresponds to an embedding network typically used in a transformer network. In a specific implementation, the self-attention network 308 includes one or more multi-head (e.g., multi-head (e.g., ×12) attention networks, such as one or more transformer networks. Figure 3 In the example of , the output of the utterance classifier 206 is fed back into the utterance classifier 206 during subsequent iterations. For example, the first set of feature data 204 is passed to the utterance classifier 206 to generate a first utterance category 208 for the first set of feature data 204. In this example, the first utterance category 208 is passed to the segment matcher 210. Furthermore, in this example, the first utterance category 208 is provided as input to the utterance classifier 206 along with the second set of feature data 204 to generate a second utterance category 208 for the second set of feature data 204.
[0093] exist Figure 3 In the example illustrated in , the fragment matcher 210 includes one or more embedding networks 310 and one or more self-attention networks 312. As an example, the embedding network 310 includes or corresponds to an embedding network typically used in a transformer network. In a specific implementation, the self-attention network 312 includes one or more multi-head (e.g., multi-head (e.g., ×12) attention networks, such as one or more transformer networks. Figure 3 In the example of , the output of the segment matcher 210 is fed back into the segment matcher 210 during subsequent iterations. For example, the first utterance category 208 and the first set of feature data 204 are passed to the segment matcher 210 to generate the first media output segment identifier 162 of the first set of feature data 204. In this example, the first media output segment identifier 162 is output by the segment matcher 210. Furthermore, in this example, the first media output segment identifier 162 is provided as input to the segment matcher 210 along with the second utterance category 208 and the second set of feature data 204 to generate the second media output segment identifier 162 of the second set of feature data 204.
[0094] As described above, the media output segment identifiers 162 from the segment matcher 210 are provided to the segment mapper 164. The segment mapper 164 uses the media output segment identifiers 162 (and optionally the constraints 250) to access or generate one or more media output segments, which are provided to the media stream assembler 168 to generate the output media stream 140.
[0095] Figure 4 According to some examples of the present disclosure Figure 1 Specifically, Figure 4 A first example of the operation of the fragment mapper 164 according to a particular implementation is highlighted.
[0096] Figure 4 Also illustrated are examples of a media segment identifier 160 and a media stream assembler 168, each of which may include information about Figures 1 to 3 Any of the features described in and / or performing the Figures 1 to 3 For example, in Figure 4 In , the media segment identifier 160 includes a feature extractor 202 configured to generate feature data 204 based on one or more segments of the input media stream 120. Additionally, in Figure 4 In the embodiment, the media segment identifier 160 includes an utterance classifier 206 configured to receive the feature data 204 and generate an utterance category 208 corresponding to the feature data 204. Figure 4 , the media segment identifier 160 includes a segment matcher 210 configured to receive the feature data 204 and the corresponding utterance category 208 , and generate a media output segment identifier 162 corresponding to the feature data 204 .
[0097] exist Figure 4 , segment mapper 164 is configured to access media segment database 402 to retrieve media output segments 166 corresponding to media output segment identifiers 162. In this example, media segment database 402 includes output segment data 114, which in this example includes segments of recorded media content, such as speech, video of facial movements during speech, etc. Each of the segments of recorded media content represents a recording of a human speaker uttering a particular sound (e.g., a phoneme, a set of phonemes, or some other segment of speech). For example, the segments of recorded media content may be derived from one or more recordings of a particular speaker, and the segments of recorded media content may be selectively rearranged to generate an output media stream 140 representing speech content that is different from the original recording of the particular speaker.
[0098] Figure 4 The media output segment identifier 162 in the example corresponds to or can be used (e.g., in conjunction with the constraint 250) to determine a unique identifier of a particular media output segment 166 from the output segment data 114. For example, the segment mapper 164 uses the media output segment identifier 162 to perform a database lookup to retrieve the media output segment 166.
[0099] Although in Figure 4 1, but in some implementations, segment mapper 164 may be configured to retrieve media output segments from two or more different media segment databases 402. Additionally or alternatively, each media segment database 402 may store segments of recorded media content representing records of two or more persons (e.g., talker_1 through talker_M, where M is an integer greater than 1). In such implementations, segment mapper 164 may access media output segments associated with a particular person based on information indicated in media output segment identifier 162. Optionally, in some implementations, segment mapper 164 may access media output segments associated with a particular person based on constraint 250. For example, constraint 250 may indicate whether output media stream 140 should include speech from a first person (e.g., a male speaker) or from a second person (e.g., a female speaker). As another example, constraint 250 may indicate that speech from multiple speakers is to be mixed to generate output media stream 140. In other examples, other voice modifications or voice selection preferences may be indicated via constraints 250 .
[0100] Figure 5 According to some examples of the present disclosure Figure 1 Specifically, Figure 5 A second example of the operation of the fragment mapper 164 according to a particular implementation is highlighted.
[0101] Figure 5 Also illustrated are examples of a media segment identifier 160 and a media stream assembler 168, each of which may include information about Figures 1 to 3 Any of the features described in and / or performing the Figures 1 to 3 For example, in Figure 5 In , the media segment identifier 160 includes a feature extractor 202 configured to generate feature data 204 based on one or more segments of the input media stream 120. Additionally, in Figure 5 In the embodiment, the media segment identifier 160 includes an utterance classifier 206 configured to receive the feature data 204 and generate an utterance category 208 corresponding to the feature data 204. Figure 5, the media segment identifier 160 includes a segment matcher 210 configured to receive the feature data 204 and the corresponding utterance category 208, and generate a media output segment identifier 162 corresponding to the feature data 204. Optionally, the media segment identifier 160 may also receive a constraint 250, in which case the media output segment identifier 162 is selected based on the constraint 250.
[0102] exist Figure 5 In the example illustrated in , segment mapper 164 is configured to generate media segment data 510 for media output segments 166 corresponding to media output segment identifiers 162. In this example, segment mapper 164 corresponds to or includes one or more layers of a neural network. For example, media output segment identifiers 162 are indicated by activation nodes of embedding layer 504. In this example, embedding layer 504 is a one-hot encoding layer; therefore, media output segment identifiers 162 are one-hot vectors or one-hot arrays. Embedding layer 504 may correspond to an output layer of segment matcher 210, or embedding layer 504 may be coupled to an output layer of segment matcher 210.
[0103] Segment mapper 164 also includes an output layer 506 coupled to embedding layer 504 via one or more links. For purposes of illustration, output layer 506 may be fully connected to embedding layer 504. In this illustrative example, each node of output layer 506 is connected to each node of embedding layer 504 via a corresponding link, and each link between embedding layer 504 and output layer 506 is associated with a corresponding link weight. Figure 5 , a set of link weights (weights 508) between one node of embedding layer 504 and a node of output layer 506 is illustrated; however, each other node of embedding layer 504 is also connected to output layer 506 via a link associated with a corresponding link weight.
[0104] Figure 5 Each set of link weights in the example of corresponds to a parameter of a corresponding media output segment. Figure 5 , weight 508 is a parameter of the media output segment 166 indicated by the media output segment identifier 162. As an example, weight 508 indicates a PCM parameter value of the media output segment 166.
[0105] In a particular example, during operation, segment mapper 164 calculates a value of media segment data 510 for each node of output layer 506, where the value of a particular node corresponds to one (1) times the weight associated with the activated node of embedding layer 504 plus zero (0) times the weight associated with each other node of embedding layer 504. Thus, in this example, each value of media segment data 510 corresponds to a value of weight 508.
[0106] In this particular example, because weights 508 correspond to values of media segment data 510, each set of weights for segment mapper 164 can be considered as a memory cell, and media output segment identifier 162 can be considered as a cell index that identifies a particular memory cell of segment mapper 164.
[0107] Figure 6 According to some examples of the present disclosure Figure 1 Specifically, Figure 6 Examples of input media stream 120, media segment identifier 160, media output segment identifier 162, segment mapper 164, media stream assembler 168, and output media stream 140 are illustrated.
[0108] exist Figure 6 In the example illustrated in , the media segment identifier 160 includes a feature extractor 202, an utterance classifier 206, and a segment matcher 210. The feature extractor 202 is coupled to the utterance classifier 206 and the segment matcher 210, and is configured to generate feature data 204 representing a portion of the input media stream 120. For example, one or more segments of the input media stream 120 may be input into the feature extractor 202 to generate feature data 204 corresponding to the one or more segments of the input media stream 120. As shown in FIG. Figure 1 As explained, input media stream 120 may include data representing speech, and optionally may include data representing other content such as other sounds and / or images.
[0109] In some specific implementations, Figure 6 The feature extractor 202 includes or corresponds to one or more trained models (e.g., machine learning models such as Figure 3 The convolutional network 302 and the self-attention network 304 are as described above. In other specific implementations, Figure 6 The feature extractor of includes or corresponds to an audio processor configured to extract speech parameters from the audio content. In such specific implementations, the feature data 204 includes data representing speech parameters. Non-limiting examples of such speech parameters include speech energy, pitch, duration, spectral representation, speaker identity, speaker gender, spectral envelope, one or more other measurable or quantifiable parameters representing speech, or a combination thereof. In other specific implementations, Figure 6 The feature extractor of includes one or more trained models and an audio processor. In such a specific implementation, the feature data 204 may include feature vectors and speech parameters.
[0110] exist Figure 6, the feature data 204 is passed to an utterance classifier 206. The utterance classifier 206 is configured to determine an utterance category 208 associated with the segment of the input media stream 120 represented by the feature data 204. The utterance category 208 associated with the segment is one of a plurality of utterance categories, the specific number and nature of which depends on the configuration of the media segment identifier 160. As an example, each utterance category may include a phoneme label that identifies a phoneme or a group of phonemes identified in the input segment of the input media stream 120. In this example, the particular utterance category 208 indicated by the utterance classifier 206 is an estimate of which phonemes are represented by the segment of the input media stream 120 represented by the feature data 204. In other examples, the utterance category represents a diphone or other artificial sound grouping.
[0111] exist Figure 6 In the example illustrated in , the feature data 204 and the utterance category 208 associated with the feature data 204 are passed to the segment matcher 210. In some implementations, additional information is provided to the segment matcher 210. For example, in Figure 6 , data representing a set of candidate frames 604 is provided to the segment matcher 210. In this example, each candidate frame 604 corresponds to a portion of a corresponding media output segment of the output segment data 114. Optionally, data specifying the constraints 250 may also be provided to the segment matcher 210.
[0112] The segment matcher 210 is configured to determine the media output segment identifier 162 based on the feature data 204 and the utterance category 208. Figure 6 In the example illustrated in , the segment matcher 210 includes a frame comparator 602 and a segment comparator 608. In this example, the frame comparator 602 and the segment comparator 608 are configured to operate in cooperation to determine the media output segment identifier 162.
[0113] The frame comparator 602 is configured to compare the one or more candidate frames 604 with data representing one or more input frames of the input media stream 120 to determine one or more candidate segments 606. In some implementations, the segment matcher 210 is configured to obtain data representing one or more candidate frames 604 of the media output segments based on the utterance category 208 and the utterance category of the output segment data 114. For purposes of illustration, the frame comparator 602 may query the output segment data 114 to identify media output segments that include frames associated with the utterance category 208. In this exemplary implementation, the frame comparator 602 compares the data representing the candidate frames 604 with the feature data 204 representing one or more frames of the input media stream 120 to determine which candidate frames are most similar to the frames of the input media stream 120. One or more output media segments associated with the most similar candidate frames 604 are provided to the segment comparator 608 as candidate segments 606.
[0114] In some implementations, as further described below, the feature data 204 representing an input frame of the input media stream 120 includes a value of a speech parameter of the input frame. In such implementations, the frame comparator 602 compares the value of the speech parameter of the input frame with the value of the speech parameter of the candidate frame 604 to select the candidate frame that is most similar to the frame of the input media stream 120. In some implementations, as further described below, the feature data 204 representing an input frame of the input media stream 120 includes an embedding (e.g., a feature vector, a feature array, or a feature map) representing the input frame. In such implementations, the frame comparator 602 compares the embedding representing the input frame with the embedding representing the candidate frame 604 to select the candidate frame that is most similar to the frame of the input media stream 120.
[0115] In a particular implementation where the feature data 204 includes speech parameter values for the input frame, the feature data 204 for the input frame may include a plurality of speech parameter values, such as a pitch of the frame, a spectral representation of the frame, an energy of the frame, a duration of the frame, etc. In this implementation, the media segment database 402 may include a corresponding speech parameter value for each output frame of each output media segment of the output segment data 114, and the frame comparator 602 obtains the speech parameter values for the output frames from the media segment database 402. Alternatively, if the media segment database 402 does not include a corresponding speech parameter value for each output frame, the frame comparator 602 or the feature extractor 202 may determine the corresponding speech parameter value for the output frame.
[0116] In this particular implementation, the frame comparator 602 compares the speech parameter values of the input frame with the speech parameter values of the candidate frame 604. In this particular implementation, the frame comparator 602 determines a frame matching score for each candidate frame 604, wherein the frame matching score indicates the similarity of the candidate frame 604 to the input frame based on the value of the speech parameter. For the purpose of illustration, the frame comparator 602 can determine the difference of each speech parameter in two or more speech parameters, and these differences can be summed up to generate the frame matching score of the specific candidate frame 604. In some examples, the difference is scaled, normalized, weighted or otherwise transformed before summing up. For the purpose of illustration, the difference can be adjusted so that each difference falls within the range of 0 to 1, and the average value or weighted average value of the adjusted difference can be calculated to sum up the adjusted difference to generate the frame matching score. In other examples, other similarity statistics are calculated to generate the frame matching score of each candidate frame 604.
[0117] In this particular implementation, the candidate frames 604 may include every output frame of the media segment database 402, or may include a subset of the output frames of the media segment database 402, such as only those output frames that satisfy the constraints 250. For example, the frame comparator 602 may determine the frame matching score based on the speech parameter value of each output frame of the media segment database 402 (or each output frame that satisfies the constraints 250), regardless of the utterance category 208 of the input frame and the utterance category of the output frame. In this example, during the determination of the media output segment identifier 162, the segment comparator 608 compares the utterance category 208 of one or more input frames with the utterance category of the candidate segment 606. In another example, the frame comparator 602 may determine the frame matching score based on the speech parameter value of each output frame of the media segment database 402 that matches the utterance category 208 of the input frame. Optionally, the frame comparator 602 may determine the frame matching score based on the speech parameter value of each output frame of the media segment database 402 that both matches the utterance category 208 of the input frame and satisfies the constraints 250.
[0118] In certain implementations where the feature data 204 includes an embedding representing an input frame, the frame comparator 602 may include a trained machine learning model configured to compare the embedding representing the input frame with the embedding representing each candidate frame 604 to assign a frame matching score to the candidate frame 604. For example, the segment matcher 210 may pass the output of the feature extractor 202 (e.g., the feature data 204) and data representing a particular candidate frame 604 to the trained machine learning model so that the trained machine learning model outputs a frame matching score for the particular candidate frame 604. The embedding representing the candidate frame 604 may be pre-computed and stored in the media segment database 402, or the candidate frame 604 may be provided to the feature extractor 202 to generate the embedding at runtime.
[0119] In this implementation, the candidate frames 604 may include every output frame of the media segment database 402, or may include a subset of the output frames of the media segment database 402, such as only those output frames that satisfy the constraints 250. For example, the frame comparator 602 may determine a frame match score based on an embedding representing every output frame of the media segment database 402 (or every output frame that satisfies the constraints 250), regardless of the utterance category 208 of the input frame and the utterance category of the output frame. In this example, during the determination of the media output segment identifier 162, the segment comparator 608 compares the utterance category 208 of one or more input frames with the utterance category of the candidate segment 606. In another example, the frame comparator 602 may determine a frame match score based on an embedding representing every output frame of the media segment database 402 that matches the utterance category 208 of the input frame. Optionally, the frame comparator 602 may determine a frame match score for every output frame of the media segment database 402 that both matches the utterance category 208 of the input frame and satisfies the constraints 250.
[0120] The frame matching score of each candidate frame 604 (whether based on the value of the speech parameter or based on the embedded comparison) indicates an estimate of the similarity of the candidate frame 604 to the input frame. The frame comparator 602 determines the candidate segments 606 based on which output segments of the media segment database 402 are associated with the candidate frame 604 having the largest frame matching score. In some implementations, the frame comparator 602 can also apply the constraints 250 when determining the candidate segments 606.
[0121] exist Figure 6 In the example illustrated in , the segment comparator 608 is configured to determine a segment match score for each of the one or more candidate segments 606. The segment match score for a particular candidate segment 606 indicates an estimate of the similarity of the particular candidate segment to an input segment represented by the output of the feature extractor 202 (e.g., the feature data 204). The segment matcher 210 is configured to determine the media output segment identifier 162 based at least in part on the segment match score. For example, in some implementations, the segment matcher 210 determines the media output segment identifier 162 based on the segment match score and the constraint 250. In other implementations, the segment matcher 210 determines the media output segment identifier 162 based solely on the segment match score. For purposes of illustration, the media output segment identifier 162 may identify the media output segment associated with the maximum segment match score among the candidate segments 606.
[0122] exist Figure 7In the specific implementation illustrated in , the segment comparator 608 is configured to determine the segment match score for the particular candidate segment based on dynamic time warping of the data representing the particular candidate segment and the data representing the input segment (e.g., the feature data 204). Figure 7 , the segment comparator 608 includes a dynamic time warping calculator 702. The dynamic time warping calculator 702 is configured to determine a segment match score for a particular candidate segment 606 based on a dynamic time warping calculation that compares data representing an input segment with data representing a particular candidate segment 606. In this example, the segment comparator 608 selects a best matching output segment 710 from the candidate segments 606 and outputs an identifier of the best matching output segment 710 as the media output segment identifier 162. Using dynamic time warping, the best matching output segment 710 corresponds to the candidate segment 606 that has a minimum cost (e.g., shortest total distance) warping path relative to the input segment.
[0123] exist Figure 8 In another specific implementation illustrated in FIG. 6 , the segment comparator 608 is configured to compare the segment matches based on a set of output frames (in Figure 8 The best matching output segment 710 is determined by identifying the candidate segment 820 associated with the output segment 710. For the purpose of illustration, Figure 8 In the example illustrated in , the input fragment includes a sequence of input frames (in Figure 8 abbreviated as "IF"), the input frame sequence includes IF t-2 IF t-1、 IF t IF t+1 and IF t+2 In this example, the segment comparator 608 accumulates a list 806 of M best matching candidate frames 604 for each input frame 802 (e.g., stored in a buffer), where M is an integer greater than 1 and can be a default or configurable parameter. In some implementations, the list 806 associated with each input frame 802 is sorted based on the frame matching score associated with each candidate frame 604.
[0124] Periodically or occasionally (e.g., when the end of an input segment is reached), the segment comparator 608 performs a weighted sorting 810 of the frame matching scores of the list 806. For example, the frame matching scores may be weighted based on feedback from previous outputs of the segment comparator 608, a comparison of the utterance category 208 associated with the input frame 802 and the utterance category associated with the candidate frame 604, or both. In some implementations, the frame matching scores are also weighted based on the memory location where the candidate frame 604 is stored in the media segment database 402. For example, the output frames may be stored in the media segment database 402 in the order in which they were recorded. For purposes of illustration, each output frame is associated with a memory location index indicating the storage location of the output frame in the media segment database 402, and the memory location index increases monotonically and sequentially for output frames representing a particular sequence of captured audio data. In this example, the weighted ordering 810 weights the frame matching scores of the candidate frames 604 to favor a sequence of candidate frames 604, wherein the sequence of candidate frames 604 includes candidate frames 604 associated with monotonically increasing memory cell indices. In some examples, the weighting may also take into account the spacing between memory cell indices. For purposes of illustration, memory cell indices having a spacing greater than a threshold may be considered non-sequential and, therefore, not weighted favorably.
[0125] In certain aspects, the best matching output segment 710 corresponds to the highest ranked set of candidate segments 820 after weighted sorting 810. In some implementations, the highest ranked set of candidate segments 820 may not represent a complete sequence of output frames. Figure 8 In the example, the output frames with gray shading correspond to the sequence, and for the input frame IF t+1 The highest ranked output frame is not part of the sequence. In this case, the segment comparator 608 may perform an optional padding operation 822 to complete the output frame sequence. For purposes of illustration, in the event that the candidate segment 820 is identified as the highest ranked segment, the segment comparator 608 may select the entire candidate segment 820 as the best matching output segment 710. In this illustrative example, the best matching output segment 710 may include an output frame that is not among the candidate frames 604 (e.g., a memory cell location ( Figure 8 OF 337 ) and the memory cell location associated with the final output frame of the best matching output segment 710 ( Figure 8 OF 345 ) in the output frame of the memory cell locations between ).
[0126] Return to Figure 6, the segment comparator 608 outputs an identifier (e.g., a memory unit index or another identifier) of the best matching output segment 710 as the media output segment identifier 162. The segment mapper 164 uses the media output segment identifier 162 to access or generate one or more media output segments, which are provided to the media stream assembler 168 to generate the output media stream 140. For example, the segment mapper 164 uses the media output segment identifier 162 (and optionally, the constraints 250) to perform a database lookup to retrieve the media output segment 166 from the media segment database 402.
[0127] Fig. 9 Some examples of the present disclosure are illustrated by Figure 1 1. A diagram of certain aspects of the operations performed by a system to generate and use media output segment identifiers based on input media streams. Specifically, Fig. 9 A first example of voice modification using components of system 100 according to some implementations is highlighted.
[0128] Fig. 9 Examples of media segment identifier 160, segment mapper 164, and media stream assembler 168 are illustrated, each of which may include information about Figures 1 to 8 Any of the features described in and / or performing the Figures 1 to 8 For example, in Fig. 9 , the media segment identifier 160 includes a feature extractor 202, an utterance classifier 206, and a segment matcher 210. Additionally, the media segment identifier 160 generates data (eg, media output segment identifiers 162) used by the segment mapper 164 and the media stream assembler 168 to generate the output media stream 140.
[0129] exist Fig. 9 , the input media stream 120 is received via microphone 126 and includes speech 902 of a first person (e.g., person 130). In this example, media segment identifier 160 parses the input media stream 120 into segments and provides the segments as input to feature extractor 202. The input media stream 120 may be parsed based on time (e.g., such that each segment represents a portion of the input media stream having a specified duration) or based on the content of the first person's speech 902 (e.g., based on sounds present in the first person's speech 902).
[0130] exist Fig. 9In the example illustrated in , the feature extractor 202 generates feature data 204 based on one or more segments of the input media stream 120, and passes the feature data 204 to the utterance classifier 206 and the segment matcher 210. The utterance classifier 206 determines an utterance category 208 associated with the feature data 204, and passes the utterance category 208 to the segment matcher 210. The segment matcher 210 generates a media output segment identifier 162 based on the utterance category 208 and the feature data 204.
[0131] Optionally, in Fig. 9 In the example illustrated in , the input media stream 120 may be processed to determine one or more constraints 250. For example, in Fig. 9 , constraints 250 include a speaker identifier (“speaker id.”) 912 that indicates a name or other identifier of first person 130 . Fig. 9 The speaker id 912 of the first person 130 may be automatically generated by the speaker identification process 910. As one example, the speaker identification process 910 may generate the speaker id 912 based on the speech characteristics of the first person 130. As another example, the speaker identification process 910 may generate the speaker id 912 based on an image of the first person 130 captured by one or more cameras (e.g., camera 132). In yet another example, the speaker identification process 910 may generate the speaker id 912 based on an image of the first person 130 captured by one or more cameras (e.g., camera 132). Figure 1 102) to generate speaker id. 912. Depending on the settings associated with media segment identifiers 160 and / or segment mapper 164 (e.g., other constraints 250), speaker id. 912 can be used to select output media segment identifiers 162 or media output segments 166 corresponding to the speech of the second person who sounds like the first person 130, or to select output media segment identifiers 162 or media output segments 166 corresponding to the speech of the second person who sounds different from the first person 130.
[0132] Segment mapper 164 retrieves or generates media output segments 166 based on media output segment identifiers 162. Fig. 9 In the example illustrated in , media output segment 166 corresponds to a segment of pre-recorded speech of another person (e.g., a second person different from person 130) that utters one or more sounds corresponding to the sounds of the segment of input media stream 120 represented by media output segment identifier 162. Media output segment 166 is provided to media stream assembler 168 along with other media segments to generate output media stream 140. Fig. 9In the example of FIG. 1 , the output media stream 140 is played by the speaker 142 as an output representing the second person's voice 904. Thus, the first person's voice 902 is used to generate the second person's corresponding voice 904 (eg, to perform voice conversion).
[0133] Fig.10 Some examples of the present disclosure are illustrated by Figure 1 1. A diagram of certain aspects of the operations performed by a system to generate and use media output segment identifiers based on input media streams. Specifically, Fig.10 A second example of using components of system 100 for voice modification according to some implementations is highlighted.
[0134] Fig.10 Examples of media segment identifier 160, segment mapper 164, and media stream assembler 168 are illustrated, each of which may include information about Figures 1 to 8 Any of the features described in and / or performing the Figures 1 to 8 For example, in Fig.10 , the media segment identifier 160 includes a feature extractor 202, an utterance classifier 206, and a segment matcher 210. Additionally, the media segment identifier 160 generates data (eg, media output segment identifiers 162) used by the segment mapper 164 and the media stream assembler 168 to generate the output media stream 140.
[0135] exist Fig.10 , the input media stream 120 is received via a microphone 126 and includes speaker recognizable speech 1002 of a person 130. In this example, the media segment identifier 160 parses the input media stream 120 into segments and provides the segments as input to the feature extractor 202. The input media stream 120 may be parsed based on time (e.g., such that each segment represents a portion of the input media stream having a specified duration) or based on the content of the speaker recognizable speech 1002 (e.g., based on the sounds present in the speaker recognizable speech 1002).
[0136] exist Fig.10In the example illustrated in , the feature extractor 202 generates feature data 204 based on one or more segments of the input media stream 120, and passes the feature data 204 to the utterance classifier 206 and the segment matcher 210. The utterance classifier 206 determines the utterance category 208 associated with the feature data 204, and passes the utterance category 208 to the segment matcher 210. The segment matcher 210 generates the media output segment identifier 162 based on the utterance category 208 and the feature data 204. Optionally, the segment matcher 210 can also generate the media output segment identifier 162 based on the constraint 250 (e.g., based on the speaker id 912 of the person 130).
[0137] Segment mapper 164 retrieves or generates media output segments 166 based on media output segment identifiers 162 (and optionally, constraints 250). Fig.10 In the example illustrated in , the media segment identifier 160 generates a series of media output segment identifiers 162 based on a series of segments of the input media stream 120, and the segment mapper 164 generates or retrieves media output segments 166 associated with different speakers for different media output segment identifiers 162 in the series of media output segment identifiers 162. For example, the segment mapper 164 can be configured to generate or retrieve a first media output segment 166 corresponding to a pre-recorded speech of a second person (e.g., a second person different from the person 130), and to generate or retrieve a second media output segment 166 corresponding to a pre-recorded speech of a third person (e.g., a third person different from the person 130 and different from the second person). In this example, the segment mapper 164 can alternate between the first media output segment 166 and the second media output segment 166. In this example, the output media stream 140, when played by the speaker 142, represents anonymized speech 1004 that cannot be identified as the speech of the person 130, nor as the speech of the second person, nor as the speech of the third person.
[0138] While the above example involves alternating between first media output segments 166 associated with a second person's voice and second media output segments 166 associated with a third person's voice, in other implementations, segment mapper 164 may change between multiple sets of media output segments 166 associated with more than two different people. Additionally or alternatively, segment mapper 164 may change between multiple sets of media output segments 166 in a pattern different from alternating with each media output segment identifier 162. For example, segment mapper 164 may randomly select a set of specific media output segments from among multiple sets of media output segments associated with different people, and retrieve or generate media output segments 166 from the set of specific media output segments based on media output segment identifier 162. For purposes of illustration, when media segment identifier 160 outputs media output segment identifier 162, segment mapper 164 (or Figure 1 The segment mapper 164 selects a speaker (e.g., a person whose pre-recorded speech is to be used) based on the output segment data 114 ( Figure 1 ) to generate or retrieve media output fragment 166.
[0139] Fig.11 Some examples of the present disclosure are illustrated by Figure 1 1. A diagram of certain aspects of the operations performed by a system to generate and use media output segment identifiers based on input media streams. Specifically, Fig.11 A third example of using components of system 100 for voice modification according to some implementations is highlighted.
[0140] Fig.11 Examples of media segment identifier 160, segment mapper 164, and media stream assembler 168 are illustrated, each of which may include information about Figures 1 to 8 Any of the features described in and / or performing the Figures 1 to 8 For example, in Fig.11 , the media segment identifier 160 includes a feature extractor 202, an utterance classifier 206, and a segment matcher 210. Additionally, the media segment identifier 160 generates data (eg, media output segment identifiers 162) used by the segment mapper 164 and the media stream assembler 168 to generate the output media stream 140.
[0141] exist Fig.11, the input media stream 120 is received via the microphone 126 and includes speech 1102 with a first accent (e.g., speech of a person 130, where the person 130 speaks with the first accent). In this example, the media segment identifier 160 parses the input media stream 120 into segments and provides the segments as input to the feature extractor 202. The input media stream 120 may be parsed based on time (e.g., such that each segment represents a portion of the input media stream having a specified duration) or based on the content of the speech 1102 with the first accent (e.g., based on sounds present in the speech 1102 with the first accent).
[0142] exist Fig.11 In the example illustrated in , the feature extractor 202 generates feature data 204 based on one or more segments of the input media stream 120, and passes the feature data 204 to the utterance classifier 206 and the segment matcher 210. The utterance classifier 206 determines an utterance category 208 associated with the feature data 204, and passes the utterance category 208 to the segment matcher 210. The segment matcher 210 generates a media output segment identifier 162 based on the utterance category 208 and the feature data 204. Optionally, the segment matcher 210 may further determine the media output segment identifier 162 based on the constraint 250.
[0143] Segment mapper 164 retrieves or generates media output segments 166 based on media output segment identifiers 162. Fig.11 In the example illustrated in , media output segment 166 corresponds to a segment of pre-recorded speech of a person speaking with a second accent (e.g., a second person different from person 130 or a second person different from person 130 who does not speak with the first accent) that makes one or more sounds corresponding to the sounds of the segment of input media stream 120 represented by media output segment identifier 162. Media output segment 166 is provided to media stream assembler 168 along with other media segments to generate output media stream 140. Fig.11 In the example of , the output media stream 140 is played by the speaker 142 as an output representing the speech 1104 having the second accent. Thus, the speech 1102 having the first accent is used to generate the corresponding speech 1104 having the second accent.
[0144] Fig.12 Some examples of the present disclosure are illustrated by Figure 1 1. A diagram of certain aspects of the operations performed by a system to generate and use media output segment identifiers based on input media streams. Specifically, Fig.12 A fourth example of using components of system 100 for voice modification according to some implementations is highlighted.
[0145] Fig.12 Examples of media segment identifier 160, segment mapper 164, and media stream assembler 168 are illustrated, each of which may include information about Figures 1 to 8 Any of the features described in and / or performing the Figures 1 to 8 For example, in Fig.12 , the media segment identifier 160 includes a feature extractor 202, an utterance classifier 206, and a segment matcher 210. Additionally, the media segment identifier 160 generates data (eg, media output segment identifiers 162) used by the segment mapper 164 and the media stream assembler 168 to generate the output media stream 140.
[0146] exist Fig.12 , the input media stream 120 is received via the microphone 126 and includes unvoiced speech 1202 (e.g., the low voice of the person 130). In this example, the media segment identifier 160 parses the input media stream 120 into segments and provides the segments as input to the feature extractor 202. The input media stream 120 may be parsed based on time (e.g., such that each segment represents a portion of the input media stream having a specified duration) or based on the content of the unvoiced speech 1202 (e.g., based on the sounds present in the unvoiced speech 1202).
[0147] exist Fig.12 In the example illustrated in , the feature extractor 202 generates feature data 204 based on one or more segments of the input media stream 120, and passes the feature data 204 to the utterance classifier 206 and the segment matcher 210. The utterance classifier 206 determines an utterance category 208 associated with the feature data 204, and passes the utterance category 208 to the segment matcher 210. The segment matcher 210 generates a media output segment identifier 162 based on the utterance category 208 and the feature data 204. Optionally, the segment matcher 210 may further determine the media output segment identifier 162 based on the constraint 250.
[0148] Segment mapper 164 retrieves or generates media output segments 166 based on media output segment identifiers 162 (and optionally, constraints 250). Fig.12 In the example illustrated in , media output segment 166 corresponds to a segment of pre-recorded voiced speech that includes one or more sounds corresponding to the sounds of the segment of input media stream 120 represented by media output segment identifier 162. The voiced speech may be pre-recorded by person 130 or another person. Media output segment 166 is provided to media stream assembler 168 along with other media segments to generate output media stream 140. Fig.12In the example of , the output media stream 140 is played by the speaker 142 as output representing the voiced speech 1204 based on the unvoiced speech 1202 in the input media stream 120. Fig.12 It is illustrated that the voiced speech 1204 is generated based on the unvoiced speech 1202 , but in other specific implementations, the unvoiced speech 1202 may be generated based on the voiced speech 1204 .
[0149] Fig.13 Some examples of the present disclosure are illustrated by Figure 1 1. A diagram of certain aspects of the operations performed by a system to generate and use media output segment identifiers based on input media streams. Specifically, Fig.13 Examples of noise reduction using components of system 100 according to some implementations are highlighted.
[0150] Fig.13 Examples of media segment identifier 160, segment mapper 164, and media stream assembler 168 are illustrated, each of which may include information about Figures 1 to 8 Any of the features described in and / or performing the Figures 1 to 8 For example, in Fig.13 , the media segment identifier 160 includes a feature extractor 202, an utterance classifier 206, and a segment matcher 210. Additionally, the media segment identifier 160 generates data (eg, media output segment identifiers 162) used by the segment mapper 164 and the media stream assembler 168 to generate the output media stream 140.
[0151] exist Fig.13 In the example illustrated in , the input media stream 120 is received via the microphone 126 and includes a speech with a first noise 1302. In this example, the media segment identifier 160 parses the input media stream 120 into segments and provides the segments as input to the feature extractor 202. The input media stream 120 may be parsed based on time (e.g., so that each segment represents a portion of the input media stream 120 having a specified duration) or based on the content of the input media stream 120 (e.g., based on the sounds present in the speech with the first noise 1302).
[0152] exist Fig.13In the example illustrated in , the feature extractor 202 generates feature data 204 based on one or more segments of the input media stream 120, and passes the feature data 204 to the utterance classifier 206 and the segment matcher 210. The utterance classifier 206 determines an utterance category 208 associated with the feature data 204, and passes the utterance category 208 to the segment matcher 210. The segment matcher 210 generates a media output segment identifier 162 based on the utterance category 208 and the feature data 204. Optionally, the segment matcher 210 may further determine the media output segment identifier 162 based on the constraint 250.
[0153] Segment mapper 164 retrieves or generates media output segments 166 based on media output segment identifiers 162 (and optionally, constraints 250). Fig.13 In the example illustrated in , media output segment 166 corresponds to a segment of pre-recorded speech without a first noise, the pre-recorded segment of speech without noise includes one or more sounds corresponding to the sounds of the segment of input media stream 120 represented by media output segment identifier 162. The speech may be pre-recorded by person 130 or another person. Media output segment 166 is provided to media stream assembler 168 along with other media segments to generate output media stream 140. In Fig.13 In the example of FIG. 1 , the output media stream 140 is played by the speaker 142 as output representing the speech 1304 without the first noise.
[0154] Fig.14 Some examples of the present disclosure are illustrated by Figure 1 1. A diagram of certain aspects of the operations performed by a system to generate and use media output segment identifiers based on input media streams. Specifically, Fig.14 A first example of communication between two devices using components of system 100 according to some implementations is highlighted.
[0155] Fig.14 Examples of media segment identifier 160, segment mapper 172, and media stream assembler 174 are illustrated, each of which may include information about Figures 1 to 8 Any of the features described in and / or performing the Figures 1 to 8 For example, in Fig.14 , the media segment identifier 160 includes a feature extractor 202, an utterance classifier 206, and a segment matcher 210. Additionally, the media segment identifier 160 generates data (eg, media output segment identifiers 162) used by the segment mapper 172 and the media stream assembler 174 to generate the output media stream 180.
[0156] exist Fig.14, the input media stream 120 is received via a microphone 126 and includes speech 128 of a person 130. In this example, the media segment identifier 160 parses the input media stream 120 into segments and provides the segments as input to the feature extractor 202. The input media stream 120 may be parsed based on time (e.g., such that each segment represents a portion of the input media stream having a specified duration) or based on the content of the speech 128 (e.g., based on the sounds present in the speech 128).
[0157] exist Fig.14 In the example illustrated in , the feature extractor 202 generates feature data 204 based on one or more segments of the input media stream 120, and passes the feature data 204 to the utterance classifier 206 and the segment matcher 210. The utterance classifier 206 determines an utterance category 208 associated with the feature data 204, and passes the utterance category 208 to the segment matcher 210. The segment matcher 210 generates a media output segment identifier 162 based on the utterance category 208 and the feature data 204. Optionally, the segment matcher 210 may further determine the media output segment identifier 162 based on the constraint 250.
[0158] exist Fig.14In the example illustrated in , the media output segment identifier 162 is provided to the modem 110, which transmits data representing the media output segment identifier 162 to the device 152 via the communication channel 150. Optionally, data representing the constraints 250 may also be provided to the modem 110 for transmission. The media output segment identifier 162 may be viewed as a very compressed representation of specific features of the input media stream 120. For example, the input media stream 120 may include information representing every aspect of the speech 128, such as timing, pitch, volume, and other sounds detected by the microphone (e.g., noise or other speakers). In contrast, the media output segment identifier 162 represents specific features extracted from the input media stream, such as phonemes, diphones, or other speech segments. Therefore, the media output segment identifier 162 may be transmitted using much fewer communication resources (e.g., power, channel time, bits) than would be used to transmit information representing the entire waveform of the input media stream 120. For purposes of illustration, in the experiment described above, the media output segment identifier 162 included approximately 96,000 elements in the one-hot encoded vector, and data representing real-time one-way voice communication can be transmitted at approximately 100 bits per second. As noted above, the media output segment identifier 162 used in this experiment is larger than necessary; therefore, by reducing the dimensionality of the media output segment identifier 162, even lower bit rates can be achieved. For example, the media output segment identifier 162 can be reduced to include one element per phoneme or one element per diphone in the particular language being transmitted (e.g., by removing duplicates or near-duplicates as described above), which will further reduce the communication resources used to transmit the data representing the media output segment identifier 162.
[0159] exist Fig.14 , modem 170 receives information sent via communication channel 150 and provides media output segment identifier 162 to segment mapper 172. Segment mapper 172 retrieves or generates media output segment 166 based on media output segment identifier 162 (and optionally, based on constraint 250). For example, device 152 may include output segment data 1402. In this example, output segment data 1402 is Figures 1 to 6 , or a similar set of data. In this example, segment mapper 172 retrieves media output segments 166 from output segment data 1402 based on media output segment identifier 162 (and optionally, based on constraint 250).
[0160] exist Fig.14In the example illustrated in , media output segment 166 corresponds to a segment of pre-recorded speech of person 130 or another person (e.g., a second person different from person 130) that utters one or more sounds corresponding to the sounds of the segment of input media stream 120 represented by media output segment identifier 162. Media output segment 166 is provided to media stream assembler 174 along with other media segments to generate output media stream 180. Fig.14 In the example of FIG. 1 , output media stream 180 is played by speaker 1404 as output speech 1406 representing speech 128 .
[0161] Fig.15 Some examples of the present disclosure are illustrated by Figure 1 1. A diagram of certain aspects of the operations performed by a system to generate and use media output segment identifiers based on input media streams. Specifically, Fig.15 A second example of communication between two devices using components of system 100 according to some implementations is highlighted.
[0162] Fig.15 Examples of media segment identifier 160, segment mapper 164, and media stream assembler 168 are illustrated, each of which may include information about Figures 1 to 8 Any of the features described in and / or performing the Figures 1 to 8 For example, in Fig.15 , the media segment identifier 160 includes a feature extractor 202, an utterance classifier 206, and a segment matcher 210. Additionally, the media segment identifier 160 generates data (eg, media output segment identifiers 162) used by the segment mapper 164 and the media stream assembler 168 to generate the output media stream 140.
[0163] exist Fig.15 In the example illustrated in , input media stream 120 is received via communication channel 124. For example, microphone 1502 of device 152 receives audio data (e.g., an audio waveform) that includes speech 128 of person 130 and may also include other sounds (e.g., background noise, etc.). Device 152 encodes the audio data (e.g., the entire audio waveform or one or more sub-bands of the waveform) using a speech or audio codec (such as an Internet Protocol voice codec), and modem 170 of device 152 sends data packets including the encoded audio data to device 102.
[0164] exist Fig.15In the example illustrated in , device 102 is configured to receive audio data sent by device 152 as input media stream 120 and generate output media stream 140 based on the received audio data. For example, the audio data received from device 152 may be provided to media stream assembler 168 as received media fragment 1506.
[0165] In some cases, a portion of the input media stream 120 may be interrupted. For example, one or more data packets sent by the device 152 may be lost or damaged, leaving a gap 1504 in the input media stream 120. In such cases, the media segment identifier 160 and the segment mapper 164 may be used together to generate an estimated media segment 1508 to fill the missing audio data associated with the gap 1504.
[0166] For example, the input media stream 120 may be provided as input to the media segment identifier 160. The media segment identifier 160 parses the input media stream 120 into segments and provides the segments as input to the feature extractor 202. The input media stream 120 may be parsed based on time (e.g., such that each segment represents a portion of the input media stream having a specified duration) or based on the content of the speech 128 (e.g., based on the sounds present in the speech 128).
[0167] exist Fig.15 In the example illustrated in , the feature extractor 202 generates feature data 204 based on one or more segments of the input media stream 120, and passes the feature data 204 to the utterance classifier 206 and the segment matcher 210. The utterance classifier 206 determines an utterance category 208 associated with the feature data 204, and passes the utterance category 208 to the segment matcher 210. The segment matcher 210 generates a media output segment identifier 162 based on the utterance category 208 and the feature data 204. Optionally, the segment matcher 210 also generates the media output segment identifier 162 based on the constraint 250 (e.g., based on the speaker id 912 of the person 130).
[0168] exist Fig.15In the example illustrated in , one or more models of the media segment identifier 160 (e.g., the feature extractor 202, the utterance classifier 206, the segment matcher 210, or a combination thereof) are temporal dynamic models that generate predictions as outputs based on inputs that include temporal context. For example, although the audio content associated with the gap 1504 is not available, because the model of the media segment identifier 160 has been trained using normal speech, the media segment identifier 160 can generate an estimate of the media output segment identifier 162 associated with the audio of the gap 1504. In some specific implementations, the estimate of the media output segment identifier 162 generated by the media segment identifier 160 can be improved by training one or more models of the media segment identifier 160 using training data that includes occasional gaps similar to the gap 1504.
[0169] exist Fig.15 In the example illustrated in , the estimate of the media output segment identifier 162 (and optionally the constraint 250) is provided to the segment mapper 164 to generate or retrieve the estimated media segment 1508 corresponding to the gap 1504. In this example, the media stream assembler 168 uses the received media segments 1506 to generate the output media stream 140, and uses the estimated media segments 1508 to fill any gaps (e.g., gap 1504). The output media stream 140 with one or more filled gaps 1510 can be provided to the speaker 142 to generate an output including a representation of the speech 144.
[0170] Thus, system 100 may be used to improve the audio quality of the output of a recipient device in the event that one or more packets of audio data cannot be played due to packet loss or packet corruption.
[0171] Fig.16 The device 102 is depicted as an implementation 1600 of an integrated circuit 1602 that includes one or more processors 104. The integrated circuit 1602 also includes an audio input section 1604 (such as one or more bus interfaces) to enable receiving input media streams 120 for processing. The integrated circuit 1602 also includes a signal output section 1606 (such as a bus interface) to enable transmission of output signals (such as output media streams 140 or media output segment identifiers 162). Fig.16 In the example illustrated in , the processor 104 includes a media segment identifier 160 and optionally includes a segment mapper 164 and a media stream assembler 168. The integrated circuit 1602 is capable of implementing operations to generate media output segment identifiers and use them as a component in a system including a microphone, such as in Fig.14 A mobile phone or tablet computer as depicted in Fig.15 The headphones depicted in Fig.16 Wearable electronic devices such as those depicted in Fig.17 The audio-controlled loudspeaker system described in Fig.18 The camera depicted in Fig.19 A virtual reality, mixed reality, or augmented reality headset as depicted in Fig. 20 or Fig.21 The means of transport depicted in .
[0172] As an illustrative, non-limiting example, Fig.17 A specific implementation 1700 is depicted in which the device 102 includes a mobile device 1702, such as a phone or tablet. The mobile device 1702 includes a microphone 126, a camera 132, and a display screen 1704. Components of the processor 104, including the media segment identifier 160 and optionally the segment mapper 164 and the media stream assembler 168, are integrated into the mobile device 1702 and are illustrated using dashed lines to indicate internal components that are generally not visible to a user of the mobile device 1702. In a specific example, the media segment identifier 160 operates to generate media output segment identifiers corresponding to segments of an input media stream. For example, the microphone 126 may capture the speech of a user of the mobile device 1702, and the media segment identifier 160 may generate media output segment identifiers representing phonemes or other utterance segments of the speech. The media output segment identifiers may be used at the mobile device 1702 or sent to another mobile device to generate an output media stream. Additionally or alternatively, in some examples, mobile device 1702 may receive an input media stream from another device, and media segment identifier 160 may be operable to generate estimated media segments to fill gaps in the output media stream due to packet loss or corruption of the input media stream.
[0173] Fig.18A specific implementation 1800 is depicted in which the device 102 includes a headphone device 1802. The headphone device 1802 includes a microphone 126. Components of the processor 104, including the media segment identifier 160 and optionally the segment mapper 164 and the media stream assembler 168, are integrated into the headphone device 1802. In a specific example, the media segment identifier 160 operates to generate media output segment identifiers corresponding to segments of an input media stream. For example, the microphone 126 may capture the speech of a user of the headphone device 1802, and may generate media output segment identifiers representing phonemes or other utterance segments of the speech. The media output segment identifiers may be used to generate an output media stream from one or more speakers 142 of the headphone device 1802, or the media output segment identifiers may be sent to another device (e.g., a mobile device, a game console, a voice assistant, etc.) to generate an output media stream. Additionally or alternatively, in some examples, the headphone device 1802 may receive an input media stream from another device, and the media segment identifier 160 may be operable to generate estimated media segments to fill gaps in the output media stream due to packet loss or corruption of the input media stream.
[0174] Fig.19 A specific implementation 1900 is depicted in which the device 102 includes a wearable electronic device 1902 (illustrated as a "smart watch"). The wearable electronic device 1902 includes a processor 104 and a display screen 1904. Components of the processor 104 (including the media segment identifier 160 and optionally the segment mapper 164 and the media stream assembler 168) are integrated in the wearable electronic device 1902. In a specific example, the media segment identifier 160 operates to generate media output segment identifiers corresponding to segments of an input media stream. For example, the microphone 126 can capture the speech of a user of the wearable electronic device 1902, and a media output segment identifier representing a phoneme or other utterance segment of the speech can be generated. The media output segment identifier can be used to generate an output media stream at the display screen 1904 of the wearable electronic device 1902, or the media output segment identifier can be sent to another device (e.g., a mobile device, a game console, a voice assistant, etc.) to generate an output media stream. Additionally or alternatively, in some examples, wearable electronic device 1902 may receive an input media stream from another device, and media segment identifier 160 may be operable to generate estimated media segments to fill gaps in the output media stream caused by packet loss or corruption of the input media stream.
[0175] Fig. 202000 is an implementation in which the device 102 includes a wireless speaker and a voice activated device 2002. The wireless speaker and the voice activated device 2002 may have wireless network connectivity and be configured to perform auxiliary operations. Fig. 20 The wireless speaker and voice activated device 2002 includes a processor 104 that includes a media segment identifier 160 (and optionally a segment mapper 164 and a media stream assembler 168). Additionally, the wireless speaker and voice activated device 2002 includes a microphone 126 and a speaker 142. During operation, in response to receiving an input media stream including user speech, the media segment identifier 160 operates to generate media output segment identifiers corresponding to segments of the input media stream. For example, the microphone 126 may capture the speech of a user of the wireless speaker and voice activated device 2002, and may generate media output segment identifiers representing phonemes or other utterance segments of the speech. The media output segment identifiers may be sent to another device (e.g., a mobile device, a game console, a voice assistant, etc.) to generate an output media stream. In some examples, the wireless speaker and voice activated device 2002 may receive an input media stream from another device, and the media segment identifier 160 may operate to generate estimated media segments to fill gaps in the output media stream caused by packet loss or corruption of the input media stream.
[0176] Fig.21 An implementation 2100 is depicted in which the device 102 is integrated into or includes a portable electronic device corresponding to the camera 132. Fig.21 , camera 132 includes processor 104 and microphone 126. Processor includes media segment identifier 160 and optionally also includes segment mapper 164 and media stream assembler 168. During operation, camera 132, microphone 126, or both generate an input media stream, and media segment identifier 160 segments the input media stream to generate media output segment identifiers corresponding to segments of the input media stream. For example, microphone 126 may capture the speech of a user of camera 132, and may generate media output segment identifiers representing phonemes or other utterance segments of the speech. The media output segment identifiers may be sent to another device (e.g., a mobile device, a game console, a voice assistant, etc.) to generate an output media stream.
[0177] Fig. 22An implementation 2200 is depicted in which the device 102 includes a portable electronic device corresponding to an extended reality headset 2202 (e.g., a virtual reality headset, a mixed reality headset, an augmented reality headset, or a combination thereof). The extended reality headset 2202 includes a microphone 126 and a processor 104. In a particular aspect, a visual interface device is positioned in front of the user's eyes to enable display of augmented reality, mixed reality, or virtual reality images or scenes to the user while wearing the extended reality headset 2202. In a particular example, the visual interface device is configured to display a notification indicating a user's voice detected in an audio signal from the microphone 126. In a particular implementation, the processor 104 includes a media segment identifier 160 and optionally also includes a segment mapper 164 and a media stream assembler 168. During operation, the microphone 126 may generate an input media stream, and the media segment identifier 160 segments the input media stream to generate media output segment identifiers corresponding to segments of the input media stream. For example, microphone 126 may capture the voice of a user of extended reality headset 2202, and media segment identifier 160 may generate media output segment identifiers representing phonemes or other utterance segments of the voice. The media output segment identifiers may be sent to another device (e.g., a mobile device, a game console, a voice assistant, etc.) to generate an output media stream.
[0178] Fig.23 An implementation 2300 is depicted in which the device 102 corresponds to or is integrated within a vehicle 2302 (illustrated as a manned or unmanned aerial device (e.g., a package delivery drone)). The microphone 126 and the processor 104 are integrated into the vehicle 2302. In a particular implementation, the processor 104 includes a media segment identifier 160 and optionally also includes a segment mapper 164 and a media stream assembler 168. During operation, the microphone 126 may generate an input media stream, and the media segment identifier 160 segments the input media stream to generate media output segment identifiers corresponding to segments of the input media stream. For example, the microphone 126 may capture speech of a person near the vehicle 2302 (such as speech including delivery instructions from an authorized user of the vehicle 2302), and the media segment identifier 160 may generate media output segment identifiers representing phonemes or other speech segments of the speech. The media output segment identifiers may be sent to another device (e.g., a server device, etc.) to generate an output media stream or store the media output segment identifiers (e.g., as evidence of the delivery instructions).
[0179] Fig.24Another embodiment 2400 is depicted in which the device 102 corresponds to or is integrated within a vehicle 2402 (illustrated as a car). The vehicle 2402 includes a processor 104, which includes a media segment identifier 160 and optionally includes a segment mapper 164 and a media stream assembler 168. The vehicle 2402 also includes a microphone 126, a speaker 142, and a display device 146. The microphone 126 is positioned to capture speech of an operator of the vehicle 2402 or a passenger of the vehicle 2402. During operation, the microphone 126 can generate an input media stream, and the media segment identifier 160 segments the input media stream to generate media output segment identifiers corresponding to segments of the input media stream. For example, the microphone 126 can capture the speech of the operator of the vehicle 2402, and the media segment identifier 160 can generate media output segment identifiers representing phonemes or other speech segments of the speech. The media output segment identifiers can be sent to another device (e.g., another vehicle, a mobile phone, etc.) to generate an output media stream. Additionally or alternatively, in some examples, vehicle 2402 may receive an input media stream from another device, and media segment identifier 160 may be operable to generate estimated media segments to fill gaps in the output media stream due to packet loss or corruption of the input media stream.
[0180] refer to Fig.25 , shows a specific implementation of a method 2500 for generating a media output segment identifier based on an input media stream. In certain aspects, one or more operations of the method 2500 are performed by the media segment identifier 160, the processor 104, the device 102, Figure 1 It is executed by at least one of the system 100, or a combination thereof.
[0181] The method 2500 includes inputting one or more segments of an input media stream into a feature extractor at block 2502. For example, the media segment identifier 160 parses the input media stream into media segments, which are provided as input to the feature extractor. Figure 2 The feature extractor 202 of FIG. For purposes of illustration, the input media stream may be parsed to generate segments including one or more phonemes or one or more other speech segments. The input media stream may be received via a microphone, via a camera, or via a communication channel. In certain aspects, the input media stream includes audio representing the speech of at least one person.
[0182] The method 2500 includes passing the output of the feature extractor to a speech classifier at block 2504 to generate at least one representation of at least one speech category among a plurality of speech categories. Figure 2The feature extractor 202 outputs feature data 204, which is provided as input to an utterance classifier 206. In this example, the utterance classifier 206 generates as output data indicating an utterance category 208 associated with the feature data 204.
[0183] Method 2500 includes passing the output of the feature extractor and at least one representation to a segment matcher to generate a media output segment identifier at block 2506. For example, Figure 2 The utterance category 208 and the feature data 204 are provided as input to the segment matcher 210. In this example, the segment matcher 210 generates the media output segment identifier 162 as output.
[0184] In some implementations, method 2500 further includes retrieving or generating one or more media output segments based on the media output segment identifier. In some such implementations, the media output segment identifier 162 may be passed to one or more memory cells of the segment mapper. In such implementations, each of the memory cells of the segment mapper includes a set of weights representing the corresponding media segment. For example, Figure 5 The segment mapper 164 includes an output layer 506 coupled to the embedding layer 504 via a plurality of links. In this example, each node of the embedding layer 504 can be envisioned as a memory unit that includes a corresponding set of weights 508 that are used by the output layer 506 to generate media segment data 510 representing a particular media output segment 166.
[0185] In some implementations, a media segment (eg, a media output segment) corresponding to a media output segment identifier can be retrieved from a database. Figure 4 The segment mapper 164 accesses the media segment database 402 to retrieve the media output segments 166.
[0186] In some implementations, method 2500 includes sending data indicating a media output segment identifier to another device. Figure 1 and Fig.11 Modem 110 of device 152 may send data indicative of media output segment identifier 162 to device 102 to enable device 152 to generate output media stream 180 .
[0187] Fig.25 The method 2500 may be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. As an example, Fig.25The method 2500 may be performed by a processor executing instructions, such as reference Fig.26 described.
[0188] refer to Fig.26 , depicts a block diagram of a particular exemplary implementation of a device, and generally designated as 2600. In various implementations, the device 2600 may have more Fig.26 More or fewer components may be illustrated. In an exemplary implementation, device 2600 may correspond to device 102 or device 152. In an exemplary implementation, device 2600 may execute the reference Figures 1 to 25 One or more operations described.
[0189] In certain implementations, device 2600 includes a processor 2606 (e.g., a central processing unit (CPU)). Device 2600 may include one or more additional processors 2610 (e.g., one or more DSPs). In certain aspects, Figure 1 The processor 104 may correspond to the processor 2606, the processor 2610, or a combination thereof. The processor 2610 may include a speech and music codec (CODEC) 2608, which includes a speech decoder ("vocoder") encoder 2636, a vocoder decoder 2638, a media segment identifier 160, a segment mapper 164, a media stream assembler 168, or a combination thereof.
[0190] The device 2600 may include memory 108 and CODEC 2634. The memory 108 may include instructions 2656 that are executable by one or more additional processors 2610 (or processor 2606) to implement the functionality described with reference to the media segment identifier 160, the segment mapper 164, the media stream assembler 168, or a combination thereof. Fig.26 In the example illustrated in , the memory 108 also includes output segment data 114 .
[0191] exist Fig.26 , device 2600 includes a modem 110 coupled to an antenna 2652 via a transceiver 2650. The modem 110, transceiver 2650, and antenna 2652 may be operable to receive an input media stream, send an output media stream, receive one or more media output segment identifiers, send one or more media output segment identifiers, or a combination thereof.
[0192] The device 2600 may include a display device 146 coupled to a display controller 2626. The speaker 142 and the microphone 126 may be coupled to a CODEC 2634. The CODEC 2634 may include a digital-to-analog converter (DAC) 2602, an analog-to-digital converter (ADC) 2604, or both. In a particular implementation, the CODEC 2634 may receive an analog signal from the microphone 126, convert the analog signal to a digital signal using the analog-to-digital converter 2604, and provide the digital signal to a voice and music codec 2608. The voice and music codec 2608 may process the digital signal, and the digital signal may be further processed by the media segment identifier 160. In a particular implementation, the voice and music codec 2608 may provide the digital signal to the CODEC 2634. The ADC 2634 may convert the digital signal to an analog signal using the digital-to-analog converter 2602, and may provide the analog signal to the speaker 142.
[0193] In a particular implementation, the device 2600 can be included in a system-in-package or system-on-chip device 2622. In a particular implementation, the memory 108, the processor 2606, the processor 2610, the display controller 2626, the decoder 2634, and the modem 110 are included in the system-in-package or system-on-chip device 2622. In a particular implementation, the input device 2630 and the power source 2644 are coupled to the system-in-package or system-on-chip device 2622. In addition, in a particular implementation, as shown in FIG. Fig.26 As illustrated, the display device 146, the input device 2630, the speaker 142, the microphone 126, the antenna 2652, and the power source 2644 are external to the system-in-package or system-on-chip device 2622. In a particular implementation, each of the display device 146, the input device 2630, the speaker 142, the microphone 126, the antenna 2652, and the power source 2644 may be coupled to a component of the system-in-package or system-on-chip device 2622, such as an interface (e.g., the input interface 106 or the output interface 112) or a controller.
[0194] Device 2600 may include a smart speaker, a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop, a computer, a tablet, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.
[0195] In conjunction with the described implementations, an apparatus includes means for inputting one or more segments of an input media stream into a feature extractor. For example, means for inputting one or more segments of an input media stream into a feature extractor may correspond to microphone 126, camera 132, communication channel 124, input interface 106, processor 104, media segment identifier 160, processor 2606, processor 2610, codec 2634, one or more other circuits or components configured to input one or more segments of an input media stream into a feature extractor, or any combination thereof.
[0196] In conjunction with the described implementation, the apparatus further includes means for passing the output of the feature extractor to an utterance classifier to generate at least one representation of at least one utterance category of the plurality of utterance categories. For example, means for passing the output of the feature extractor to the utterance classifier may correspond to the feature extractor 202, the processor 104, the media segment identifier 160, the processor 2606, the processor 2610, one or more other circuits or components configured to provide input to the utterance classifier, or any combination thereof.
[0197] In conjunction with the described implementation, the apparatus further includes means for passing the output of the feature extractor and at least one representation of at least one utterance category to the segment matcher to generate a media output segment identifier. For example, the means for passing the output of the feature extractor and at least one representation of at least one utterance category to the segment matcher may correspond to the feature extractor 202, the utterance classifier 206, the processor 104, the media segment identifier 160, the processor 2606, the processor 2610, one or more other circuits or components configured to provide input to the segment matcher, or any combination thereof.
[0198] In some embodiments, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 108) includes instructions (e.g., instructions 2656) that, when executed by one or more processors (e.g., one or more processors 104, one or more processors 2610, or processor 2606), cause the one or more processors to: input one or more segments of an input media stream into a feature extractor; pass the output of the feature extractor to a utterance classifier to generate at least one representation of at least one utterance category of a plurality of utterance categories; and pass the output of the feature extractor and the at least one representation to a segment matcher to generate a media output segment identifier.
[0199] Specific aspects of the present disclosure are described below in various sets of related embodiments:
[0200] According to embodiment 1, a device includes: one or more processors, wherein the one or more processors are configured to: input one or more segments of an input media stream into a feature extractor; pass the output of the feature extractor to a discourse classifier to generate at least one representation of at least one discourse category among multiple discourse categories; and pass the output of the feature extractor and the at least one representation to a segment matcher to generate a media output segment identifier.
[0201] Embodiment 2 includes the apparatus of embodiment 1, wherein the media output segment identifier comprises a unit index.
[0202] Embodiment 3 includes a device according to embodiment 1 or embodiment 2, wherein the segment matcher is configured to obtain data representing one or more candidate frames of the media output segment based on the at least one representation, and perform a comparison of the data representing the one or more candidate frames with the output of the feature extractor, and wherein the media output segment identifier is determined based on the result of the comparison.
[0203] Embodiment 4 includes a device according to any one of embodiments 1 to 3, wherein the segment matcher is configured to: obtain data representing multiple candidate frames, each of the multiple candidate frames corresponding to a portion of a corresponding media output segment; and determine frame matching scores for the multiple candidate frames, wherein the frame matching score of a particular candidate frame indicates an estimate of the similarity between the particular candidate frame and an input frame represented by the output of the feature extractor, wherein the media output segment identifier is determined at least in part based on the frame matching score.
[0204] Embodiment 5 includes a device according to embodiment 4, wherein the fragment matcher is configured to determine the frame matching score of the specific candidate frame by passing the output of the feature extractor and the data representing the specific candidate frame into a trained machine learning model so that the trained machine learning model outputs the frame matching score of the specific candidate frame.
[0205] Embodiment 6 includes a device according to embodiment 4, wherein the output of the feature extractor includes one or more speech parameter values of the input frame, and wherein the fragment matcher is configured to determine the frame matching score of the specific candidate frame based on a comparison of a speech parameter value among the one or more speech parameter values of the input frame with one or more corresponding speech parameter values of the specific candidate frame.
[0206] Embodiment 8 includes a device according to embodiment 4, wherein the segment matcher is configured to: determine one or more candidate segments based on the one or more frame matching scores; and for each of the one or more candidate segments, determine a segment matching score, wherein the segment matching score of a particular candidate segment indicates an estimate of the similarity of the particular candidate segment to the input segment represented by the output of the feature extractor, wherein the media output segment identifier is determined at least in part based on the one or more segment matching scores.
[0207] Embodiment 9 includes the apparatus of Embodiment 8, wherein the media output segment identifier identifies the media output segment associated with a maximum segment matching score among the one or more candidate segments.
[0208] Embodiment 10 includes an apparatus according to embodiment 8, wherein the fragment matcher is configured to determine the fragment matching score of the specific candidate fragment based on dynamic time warping of data representing the specific candidate fragment and data representing the input fragment.
[0209] Embodiment 11 comprises the apparatus according to embodiment 8, wherein the fragment matcher is configured to: for each input frame of the input fragment, determine a plurality of frame matching scores with respect to the input frame, wherein the input fragment comprises a plurality of input frames, wherein the fragment matching score of the particular candidate fragment is based on the frame matching scores of the candidate frame compared with different input frames in the plurality of input frames, and also based on a memory location associated with the candidate frame.
[0210] Embodiment 12 includes the apparatus of embodiment 8, wherein the segment matching score for the particular candidate segment is further based on the at least one representation of at least one utterance category.
[0211] Embodiment 13 includes a device according to any one of embodiments 1 to 12, wherein the segment matcher is configured to: obtain data representing one or more candidate frames of a media output segment; perform a comparison of the data representing the one or more candidate frames with the output of the feature extractor to identify one or more candidate segments; and determine the best matching media output segment among the one or more candidate segments based on the at least one representation, and wherein the media output segment identifier identifies the best matching media output segment.
[0212] Embodiment 14 includes the apparatus of any one of Embodiments 1 to 13, wherein the one or more processors are further configured to pass one or more constraints to the segment matcher to determine the media output segment identifier.
[0213] Embodiment 15 includes the apparatus of embodiment 14, wherein the one or more constraints include a speaker identifier.
[0214] Embodiment 16 includes the apparatus of any one of Embodiments 1 to 15, wherein the media output segment identifier identifies a recorded media segment corresponding to at least one phoneme.
[0215] Embodiment 17 includes a device according to any one of embodiments 1 to 16, wherein the one or more processors are further configured to pass the media output segment identifier to one or more memory units, wherein each of the one or more memory units includes a set of weights representing the corresponding media segment.
[0216] Embodiment 18 includes the apparatus of embodiment 16, wherein the one or more memory units generate a speech representation based on the media output segment identifier.
[0217] Embodiment 19 includes the device of any one of Embodiments 1 to 18, further comprising a modem coupled to the one or more processors, the modem configured to send data indicative of the media output segment identifier to another device.
[0218] Embodiment 20 includes the device of embodiment 19, wherein the input media stream includes audio representing speech, and the sending of the data indicating the media output segment identifier enables the generation of speech output in real time at other devices.
[0219] Embodiment 21 includes the apparatus of any one of Embodiments 1-16, wherein the one or more processors are further configured to retrieve a media segment corresponding to the media output segment identifier from a database.
[0220] Embodiment 22 includes the apparatus of any one of Embodiments 1 to 21, further comprising one or more receivers configured to receive an input media stream via a communication channel.
[0221] Embodiment 23 includes a device according to any one of embodiments 1 to 22, wherein the one or more processors are further configured to determine that a particular media segment of the input media stream is not available for playback, and wherein the media output segment identifier corresponds to an estimate of the particular media segment.
[0222] Embodiment 24 includes the apparatus of embodiment 23, wherein the one or more processors are further configured to concatenate the estimate of the particular media segment with one or more media segments of the input media stream to generate an audio stream.
[0223] Embodiment 25 includes an apparatus according to any one of Embodiments 1 to 24, wherein the output of the feature extractor includes feature data.
[0224] Embodiment 26 includes a device according to any one of Embodiments 1 to 25, wherein the one or more processors are further configured to pass the at least one representation of the at least one utterance category from the first iteration of the utterance classifier as input to the utterance classifier during a subsequent iteration of the utterance classifier.
[0225] Embodiment 27 includes a device according to any one of embodiments 1 to 26, wherein the one or more processors are further configured to pass the media output segment identifier from the first iteration of the segment matcher as input to the segment matcher during a subsequent iteration of the segment matcher.
[0226] Embodiment 28 includes an apparatus as described in any of Embodiments 1 to 27, wherein the input media stream includes audio representing the speech of at least one first person, and the media output segment identifier enables output of the corresponding speech of at least one second person.
[0227] Embodiment 29 includes an apparatus as described in any of Embodiments 1 to 28, wherein the input media stream includes audio representing a first speech having a first accent, and the media output segment identifier enables output of a corresponding second speech having a second accent.
[0228] Embodiment 30 includes the apparatus of any one of Embodiments 1 to 29, wherein the input media stream includes audio representing speech and a first noise, and the media output segment identifier enables output of the corresponding speech without the first noise.
[0229] Embodiment 31 includes the apparatus of any one of Embodiments 1 to 30, wherein the input media stream includes audio representing speech of at least one first person, and the media output segment identifier enables output of corresponding anonymized speech.
[0230] Embodiment 32 includes a device according to any one of Embodiments 1 to 31, wherein the one or more processors are further configured to concatenate the media segment associated with the media output segment identifier with one or more additional media segments to generate an audio stream.
[0231] Embodiment 33 includes a device according to any one of embodiments 1 to 32, wherein the device also includes one or more microphones, the one or more microphones are coupled to the one or more processors, and the one or more microphones are configured to receive audio data and generate the input media stream based on the audio data.
[0232] Embodiment 34 includes a device according to any one of embodiments 1 to 33, wherein the one or more processors are integrated in at least one of the following: a mobile phone, a tablet computer device, a wearable electronic device, a camera device, a virtual reality headset, a mixed reality headset, or an augmented reality headset.
[0233] Embodiment 35 includes a device according to any one of embodiments 1 to 34, wherein the device also includes one or more speakers, the one or more speakers are coupled to the one or more processors, and the one or more speakers are configured to output sound based on the media segment identified by the media output segment identifier.
[0234] According to embodiment 36, a method includes: inputting one or more segments of an input media stream into a feature extractor; passing the output of the feature extractor to a discourse classifier to generate at least one representation of at least one discourse category among a plurality of discourse categories; and passing the output of the feature extractor and the at least one representation to a segment matcher to generate a media output segment identifier.
[0235] Embodiment 37 includes the method according to embodiment 36, wherein the media output segment identifier includes a unit index.
[0236] Embodiment 38 includes the method of embodiment 36 or embodiment 37, wherein the media output segment identifier identifies a recorded media segment corresponding to at least one phoneme.
[0237] Embodiment 39 includes a method according to any one of embodiments 36 to 38, the method also comprising passing the media output segment identifier to one or more memory units, wherein each of the one or more memory units includes a set of weights representing the corresponding media segment.
[0238] Embodiment 40 includes the method of embodiment 39, wherein the one or more memory units generate a speech representation based on the media output segment identifier.
[0239] Embodiment 41 includes the method of any one of Embodiments 36 to 40, further comprising sending data indicative of the media output segment identifier to another device.
[0240] Embodiment 42 comprises a method according to embodiment 41, wherein the input media stream includes audio representing speech, and sending the data indicating the media output segment identifier enables real-time generation of speech output at other devices.
[0241] Embodiment 43 includes the method according to any one of embodiments 36 to 38, the method further comprising retrieving a media segment corresponding to the media output segment identifier from a database.
[0242] Embodiment 44 includes the method of any one of Embodiments 36 to 43, further comprising receiving the input media stream via a communication channel.
[0243] Embodiment 45 includes a method according to any one of Embodiments 36 to 44, the method further comprising determining that a particular media segment of the input media stream is not available for playback, and wherein the media output segment identifier corresponds to an estimate of the particular media segment.
[0244] Embodiment 46 includes the method according to embodiment 45, further comprising concatenating the estimate of the specific media segment with one or more media segments of the input media stream to generate an audio stream.
[0245] Embodiment 47 comprises a method according to any one of Embodiments 36 to 46, wherein the output of the feature extractor comprises feature data.
[0246] Embodiment 48 includes a method according to any one of Embodiments 36 to 47, further comprising passing the at least one representation of the at least one utterance category from the first iteration of the utterance classifier as input to the utterance classifier during a subsequent iteration of the utterance classifier.
[0247] Embodiment 49 includes the method of any one of Embodiments 36 to 48, further comprising passing the media output segment identifier from a first iteration of the segment matcher as input to the segment matcher during a subsequent iteration of the segment matcher.
[0248] Embodiment 50 comprises a method according to any one of Embodiments 36 to 49, wherein the input media stream includes audio representing the speech of at least one first person, and the media output segment identifier enables output of the corresponding speech of at least one second person.
[0249] Embodiment 51 comprises a method according to any one of Embodiments 36 to 50, wherein the input media stream includes audio representing a first speech having a first accent, and the media output segment identifier enables output of a corresponding second speech having a second accent.
[0250] Embodiment 52 includes a method according to any one of Embodiments 36 to 51, wherein the input media stream includes audio representing speech and a first noise, and the media output segment identifier enables output of the corresponding speech without the first noise.
[0251] Embodiment 53 includes the method of any one of Embodiments 36 to 52, wherein the input media stream includes audio representing speech of at least one first person, and the media output segment identifier enables output of corresponding anonymized speech.
[0252] Embodiment 54 includes the method of any one of Embodiments 36 to 53, further comprising concatenating the media segment associated with the media output segment identifier with one or more additional media segments to generate an audio stream.
[0253] According to embodiment 55, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to: input one or more segments of an input media stream into a feature extractor; pass the output of the feature extractor to a discourse classifier to generate at least one representation of at least one discourse category among a plurality of discourse categories; and pass the output of the feature extractor and the at least one representation to a segment matcher to generate a media output segment identifier.
[0254] Embodiment 56 includes the non-transitory computer-readable medium of embodiment 55, wherein the media output segment identifier comprises a unit index.
[0255] Embodiment 57 includes a non-transitory computer-readable medium according to embodiment 55 or embodiment 56, wherein the media output segment identifier identifies a recorded media segment corresponding to at least one phoneme.
[0256] Embodiment 58 comprises a non-transitory computer-readable medium according to any one of embodiments 55 to 57, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to pass the media output segment identifier to one or more memory units, wherein each of the one or more memory units comprises a set of weights representing the corresponding media segment.
[0257] Embodiment 59 includes the non-transitory computer-readable medium of embodiment 58, wherein the one or more memory units generate a speech representation based on the media output segment identifier.
[0258] Embodiment 60 includes a non-transitory computer-readable medium according to any one of Embodiments 55 to 59, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to send data indicating the media output segment identifier to another device.
[0259] Embodiment 61 comprises a non-transitory computer-readable medium according to embodiment 60, wherein the input media stream comprises audio representing speech, and the sending of the data indicating the media output segment identifier enables the generation of speech output in real time at other devices.
[0260] Embodiment 62 includes a non-transitory computer-readable medium according to any one of embodiments 55 to 57, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to retrieve a media segment corresponding to the media output segment identifier from a database.
[0261] Embodiment 63 includes the non-transitory computer readable medium of any one of Embodiments 55 to 62, wherein the input media stream is received over a communication channel.
[0262] Embodiment 64 comprises a non-transitory computer-readable medium according to any one of embodiments 55 to 63, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to determine that a particular media segment of the input media stream is not available for playback, and wherein the media output segment identifier corresponds to an estimate of the particular media segment.
[0263] Embodiment 65 includes a non-transitory computer-readable medium according to embodiment 64, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to cascade the estimate of the particular media segment with one or more media segments of the input media stream to generate an audio stream.
[0264] Embodiment 66 includes a non-transitory computer readable medium according to any one of Embodiments 55 to 65, wherein the output of the feature extractor includes feature data.
[0265] Embodiment 67 includes a non-transitory computer-readable medium according to any one of Embodiments 55 to 66, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to pass the at least one representation of the at least one utterance category from the first iteration of the utterance classifier as input to the utterance classifier during a subsequent iteration of the utterance classifier.
[0266] Embodiment 68 includes a non-transitory computer-readable medium according to any one of embodiments 55 to 67, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to pass the media output segment identifier from the first iteration of the segment matcher as input to the segment matcher during a subsequent iteration of the segment matcher.
[0267] Embodiment 69 comprises a non-transitory computer-readable medium as described in any one of Embodiments 55 to 68, wherein the input media stream comprises audio representing speech of at least one first person, and the media output segment identifier enables output of a corresponding speech of at least one second person.
[0268] Embodiment 70 comprises a non-transitory computer-readable medium according to any one of Embodiments 55 to 69, wherein the input media stream comprises audio representing a first speech having a first accent, and the media output segment identifier enables output of a corresponding second speech having a second accent.
[0269] Embodiment 71 comprises a non-transitory computer-readable medium according to any one of Embodiments 55 to 70, wherein the input media stream comprises audio representing speech and a first noise, and the media output segment identifier enables output of the corresponding speech without the first noise.
[0270] Embodiment 72 includes a non-transitory computer-readable medium according to any one of Embodiments 55 to 71, wherein the input media stream includes audio representing speech of at least one first person, and the media output segment identifier enables output of corresponding anonymized speech.
[0271] Embodiment 73 comprises a non-transitory computer-readable medium according to any one of embodiments 55 to 72, wherein the instructions, when executed by the one or more processors, further cause the one or more processors to cascade the media segment associated with the media output segment identifier with one or more additional media segments to generate an audio stream.
[0272] According to embodiment 74, a device includes: a component for inputting one or more segments of an input media stream into a feature extractor; a component for passing the output of the feature extractor to a discourse classifier to generate at least one representation of at least one discourse category among a plurality of discourse categories; and a component for passing the output of the feature extractor and the at least one representation to a segment matcher to generate a media output segment identifier.
[0273] Embodiment 75 comprises the apparatus of embodiment 74, wherein the media output segment identifier comprises a unit index.
[0274] Embodiment 76 includes an apparatus according to embodiment 74 or embodiment 75, wherein the media output segment identifier identifies a recorded media segment corresponding to at least one phoneme.
[0275] Embodiment 77 comprises an apparatus according to any one of embodiments 74 to 76, wherein the apparatus further comprises a component for passing the media output segment identifier to one or more memory units, wherein each of the one or more memory units comprises a set of weights representing the corresponding media segment.
[0276] Embodiment 78 includes the apparatus of embodiment 77, wherein the one or more memory units generate a speech representation based on the media output segment identifier.
[0277] Embodiment 79 includes an apparatus according to any one of Embodiments 74 to 78, further comprising means for sending data indicative of the media output segment identifier to another device.
[0278] Embodiment 80 comprises an apparatus according to Embodiment 79, wherein the input media stream includes audio representing speech, and sending the data indicating the media output segment identifier enables real-time generation of speech output at other devices.
[0279] Embodiment 81 includes an apparatus according to any one of Embodiments 74 to 76, further comprising means for retrieving a media segment corresponding to the media output segment identifier from a database.
[0280] Embodiment 82 includes the apparatus of any one of Embodiments 74 to 81, further comprising means for receiving the input media stream over a communication channel.
[0281] Embodiment 83 comprises an apparatus according to any one of embodiments 74 to 82, further comprising a component for determining that a particular media segment of the input media stream is not available for playback, and wherein the media output segment identifier corresponds to an estimate of the particular media segment.
[0282] Embodiment 84 includes the apparatus of embodiment 83, further comprising means for concatenating the estimate of the particular media segment with one or more media segments of the input media stream to generate an audio stream.
[0283] Embodiment 85 comprises an apparatus according to any one of Embodiments 74 to 84, wherein the output of the feature extractor comprises feature data.
[0284] Embodiment 86 includes an apparatus according to any one of Embodiments 74 to 85, further comprising a component for passing at least one representation of at least one utterance category from a first iteration of the utterance classifier as input to the utterance classifier during a subsequent iteration of the utterance classifier.
[0285] Embodiment 87 includes an apparatus according to any one of Embodiments 74 to 86, further comprising a component for passing the media output segment identifier from the first iteration of the segment matcher as input to the segment matcher during a subsequent iteration of the segment matcher.
[0286] Embodiment 88 includes an apparatus according to any one of Embodiments 74 to 87, wherein the input media stream includes audio representing the speech of at least one first person, and the media output segment identifier enables output of the corresponding speech of at least one second person.
[0287] Embodiment 89 comprises an apparatus according to any one of Embodiments 74 to 88, wherein the input media stream includes audio representing a first speech having a first accent, and the media output segment identifier enables output of a corresponding second speech having a second accent.
[0288] Embodiment 90 includes an apparatus according to any one of Embodiments 74 to 89, wherein the input media stream includes audio representing speech and a first noise, and the media output segment identifier enables output of the corresponding speech without the first noise.
[0289] Embodiment 91 includes an apparatus according to any one of Embodiments 74 to 90, wherein the input media stream includes audio representing speech of at least one first person, and the media output segment identifier enables output of corresponding anonymized speech.
[0290] Embodiment 92 includes an apparatus according to any one of Embodiments 74 to 91, further comprising means for concatenating the media segment associated with the media output segment identifier with one or more additional media segments to generate an audio stream.
[0291] It will also be appreciated by the skilled person that the various illustrative logic blocks, configurations, modules, circuits, and algorithmic steps described in conjunction with the specific implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of the two. Various illustrative components, blocks, configurations, modules, circuits, and steps have been generally described above in terms of their functionality. Whether such functions are implemented as hardware or processor executable instructions depends on the specific application and the design constraints imposed on the overall system. The skilled person can implement the described functions in different ways for each specific application, and such specific implementation decisions will not be interpreted as causing a departure from the scope of the present disclosure.
[0292] The steps of the method or algorithm described in conjunction with the specific implementation disclosed herein may be directly embodied in hardware, in a software module executed by a processor, or in a combination of the two. The software module may reside in a random access memory (RAM), a flash memory, a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a register, a hard disk, a removable disk, a compact disc read-only memory (CD-ROM), or any other form of non-transient storage medium known in the art. An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. In an alternative, the storage medium may be integral with the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In an alternative, the processor and the storage medium may reside in a computing device or a user terminal as discrete components.
[0293] The previous description of the disclosed aspects is provided to enable those skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the present disclosure. Thus, the present disclosure is not intended to be limited to the aspects shown herein, but should be granted the broadest scope that may be consistent with the principles and novel features as defined by the following claims.
Claims
1. A device, comprising: One or more processors configured to: Inputting one or more segments of an input media stream into a feature extractor; passing the output of the feature extractor into an utterance classifier to generate at least one representation of at least one utterance category of a plurality of utterance categories; as well as The output of the feature extractor and the at least one representation are passed to a segment matcher to determine a media output segment identifier.
2. A device according to claim 1, wherein the segment matcher is configured to obtain data representing one or more candidate frames of the media output segment based on the at least one representation, and perform a comparison of the data representing the one or more candidate frames with the output of the feature extractor, and wherein the media output segment identifier is determined based on the result of the comparison.
3. The apparatus of claim 1 , wherein the fragment matcher is configured to: obtaining data representing a plurality of candidate frames, each candidate frame of the plurality of candidate frames corresponding to a portion of a corresponding media output segment; and Determining frame match scores for the plurality of candidate frames, wherein the frame match score for a particular candidate frame indicates an estimate of similarity of the particular candidate frame to an input frame represented by the output of the feature extractor, wherein the media output segment identifier is determined based at least in part on the frame match score.
4. A device according to claim 3, wherein the fragment matcher is configured to determine the frame matching score of the specific candidate frame by passing the output of the feature extractor and the data representing the specific candidate frame into a trained machine learning model so that the trained machine learning model outputs the frame matching score of the specific candidate frame.
5. A device according to claim 3, wherein the output of the feature extractor includes one or more speech parameter values of the input frame, and wherein the fragment matcher is configured to determine the frame matching score of the specific candidate frame based on a comparison of a speech parameter value among the one or more speech parameter values of the input frame with one or more corresponding speech parameter values of the specific candidate frame.
6. The apparatus of claim 3, wherein the fragment matcher is configured to: determining one or more candidate segments based on the one or more frame matching scores; and For each of the one or more candidate segments, a segment match score is determined, wherein the segment match score for a particular candidate segment indicates an estimate of the similarity of the particular candidate segment to an input segment represented by the output of the feature extractor, wherein the media output segment identifier is determined at least in part based on the one or more segment match scores.
7. The apparatus of claim 6, wherein the media output segment identifier identifies the media output segment associated with a maximum segment matching score among the one or more candidate segments.
8. The apparatus of claim 6, wherein the fragment matcher is configured to determine the fragment matching score of the specific candidate fragment based on dynamic time warping of data representing the specific candidate fragment and data representing the input fragment.
9. A device according to claim 6, wherein the fragment matcher is configured to: for each input frame of the input fragment, determine multiple frame matching scores with respect to the input frame, wherein the input fragment includes multiple input frames, and wherein the fragment matching score of the particular candidate fragment is based on the frame matching scores of the candidate frame compared with different input frames in the multiple input frames, and also based on a memory location associated with the candidate frame.
10. The apparatus of claim 6, wherein the segment matching score for the particular candidate segment is further based on the at least one representation of at least one utterance category.
11. The apparatus of claim 1 , wherein the fragment matcher is configured to: Obtaining data representing one or more candidate frames of a media output segment; performing a comparison of the data representing the one or more candidate frames with the output of the feature extractor to identify one or more candidate segments; and A best matching media output segment among the one or more candidate segments is determined based on the at least one representation, and wherein the media output segment identifier identifies the best matching media output segment.
12. The device of claim 1, wherein the one or more processors are further configured to pass one or more constraints to the segment matcher to determine the media output segment identifier.
13. The apparatus of claim 12, wherein the one or more constraints include a speaker identifier.
14. The apparatus of claim 1, wherein the media output segment identifier identifies a recorded media segment corresponding to at least one phoneme.
15. The device of claim 1, wherein the one or more processors are further configured to pass the media output segment identifier to one or more memory units, wherein each of the one or more memory units includes a set of weights representing a corresponding media segment.
16. The device of claim 1, further comprising a modem coupled to the one or more processors, the modem configured to send data indicative of the media output segment identifier to another device.
17. The device of claim 1, further comprising one or more receivers configured to receive the input media stream over a communication channel.
18. The device of claim 1, wherein the one or more processors are further configured to determine that a particular media segment of the input media stream is not available for playback, and wherein the media output segment identifier corresponds to an estimate of the particular media segment.
19. The device of claim 18, wherein the one or more processors are further configured to concatenate the estimate of the particular media segment with one or more media segments of the input media stream to generate an audio stream.
20. The device of claim 1, wherein the input media stream includes audio representing speech of at least one first person, and the media output segment identifier enables output of a corresponding speech of at least one second person.
21. The apparatus of claim 1, wherein the input media stream includes audio representing a first speech having a first accent, and the media output segment identifier enables output of a corresponding second speech having a second accent.
22. The apparatus of claim 1, wherein the input media stream includes audio representing speech and a first noise, and the media output segment identifier enables output of the corresponding speech without the first noise.
23. The apparatus of claim 1, wherein the input media stream includes audio representing speech of at least one first person, and the media output segment identifier enables output of corresponding anonymized speech.
24. The device of claim 1, wherein the one or more processors are further configured to concatenate the media segment associated with the media output segment identifier with one or more additional media segments to generate an audio stream.
25. The device of claim 1, further comprising one or more microphones coupled to the one or more processors, the one or more microphones configured to receive audio data and generate the input media stream based on the audio data.
26. The device of claim 1, wherein the one or more processors are integrated into at least one of: a mobile phone, a tablet computer device, a wearable electronic device, a camera device, a virtual reality headset, a mixed reality headset, or an augmented reality headset.
27. The device of claim 1, further comprising one or more speakers coupled to the one or more processors, the one or more speakers being configured to output sound based on the media segment identified by the media output segment identifier.
28. A method comprising: Inputting one or more segments of an input media stream into a feature extractor; passing the output of the feature extractor into an utterance classifier to generate at least one representation of at least one utterance category of a plurality of utterance categories; and The output of the feature extractor and the at least one representation are passed into a segment matcher to generate a media output segment identifier.
29. A non-transitory computer readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to: Inputting one or more segments of an input media stream into a feature extractor; passing the output of the feature extractor into an utterance classifier to generate at least one representation of at least one utterance category of a plurality of utterance categories; and The output of the feature extractor and the at least one representation are passed into a segment matcher to generate a media output segment identifier.
30. An apparatus, comprising: means for inputting one or more segments of an input media stream into a feature extractor; means for passing the output of the feature extractor to an utterance classifier to produce at least one representation of at least one utterance class of a plurality of utterance classes; and Means for passing said output of said feature extractor and said at least one representation to a segment matcher to generate a media output segment identifier.