Method and apparatus for efficient real-time audio style transfer using granular synthesis

The method efficiently maps audio segments to a dimensional space for real-time audio style transfer, addressing slow workflows and lack of user interaction in existing methods, achieving low-latency and user-controlled audio synthesis.

WO2025158057A1PCT designated stage Publication Date: 2025-07-31ZYNAPTIQ GMBH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/EP2025/051891
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-24
Filing Date
2025-01-24
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

Existing audio synthesis methods, such as Boteas and Realtime Audio Variational autoEncoder (RAVE) models, require multiple comparisons and training for each file, leading to slow workflows and lack of intuitive user interaction for modifying audio synthesis.

Method used

A method that partitions audio files into segments, analyzes their characteristics, maps them to a dimensional space, and allows real-time modification with user input, enabling efficient and intuitive audio style transfer.

Benefits of technology

Enables low-latency, low-computational audio style transfer with user control, allowing for the creation of new audio signals that replicate original files with different acoustic properties.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure EP2025051891_31072025_PF_FP_ABST
    Figure EP2025051891_31072025_PF_FP_ABST
Patent Text Reader

Abstract

In one embodiment, a computer implemented method for audio style transfer, is disclosed. The method may include partitioning, via a processor, a first audio file into a first audio segment, wherein the first audio segment includes a portion of the first audio file; determining, via the processor, a first audio characteristic of the first audio segment based on an analysis of the first audio segment; mapping, via the processor, the first audio segment to a dimensional space based on the first audio characteristic; determining, via the processor, a second audio characteristic of a second audio segment based on an analysis of the second audio segment; identifying, via the processor, the first audio segment in the dimensional space based on the second audio characteristic; modifying, via the processor, the second audio segment based on the first audio segment; and outputting, via the processor, the second audio segment.
Need to check novelty before this filing date? Find Prior Art

Description

METHOD AND APPARATUS FOR EFFICIENT REAL-TIME AUDIO STYLE TRANSFER USING GRANULAR SYNTHESISCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This present application claims priority to U.S. Provisional Application No. 63 / 624,439 filed on January 24, 2024 entitled “Method and Apparatus For Efficient Real Time Audio Style Transfer Using Granular Synthesis,” which is hereby incorporated by reference herein in its entirety.BACKGROUND

[0002] Granular Synthesis is one of the most widely used forms of audio synthesis allowing to synthesize new sounds from small portions of existing sounds. One of these approaches is described in Boteas et al. (US Patent No. 10,606,548), which discloses an approach to generate a composite signal from a file based on comparisons of an input signal and the file. The Boteas approach, however, has the drawback that multiple comparisons need to be made to select the next portions of audio files, which can have performance implications, especially if files of arbitrary length with many portions are used. Additionally, the user cannot intuitively interact directly with this comparison to increase or decrease the likelihood of certain grains appearing in the output.

[0003] More recently, the Realtime Audio Variational autoEncoder (RAVE) models have been using techniques similar to granular synthesis. These machine learning based models use variational autoencoders to determine the properties of audio files and then synthesize audio based on training data. This process requires training for each file, making the user workflow much slower, especially if several files need to be tested for creative purposes. The whole system is a black box to the user, so the user cannot intuitively modify the synthesis, since the dimensions of the latent space will correspond to different properties for each signal.

[0004] Thus, a need exists for an improved system for the real-time synthetization of audio.BRIEF SUMMARY

[0005] Aspects and features of the present disclosure are set out in the appended claims.OVERVIEW OF DISCLOSURE

[0006] In one embodiment, a computer implemented method for audio style transfer, is disclosed. The method may include partitioning, via a processor, a first audio file into a first audio segment, wherein the first audio segment includes a portion of the first audio file; determining, via the processor, a first audio characteristic of the first audio segment based on an analysis of the first audio segment; mapping, via the processor, the first audio segment to a dimensional space based on the first audio characteristic; determining, via the processor, a second audio characteristic of a second audio segment based on an analysis of the second audio segment; identifying, via the processor, the first audio segment in the dimensional space based on the second audio characteristic; modifying, via the processor, the second audio segment based on the first audio segment; and outputting, via the processor, the second audio segment.

[0007] Optionally, in some embodiments, the second audio segment may be modified based on the first audio segment and an audio inertia value.

[0008] In some embodiments, the processing for creating the second audio segment may be done in real time.

[0009] Optionally, in some embodiments, partitioning the first audio file into the first audio segment includes applying a window function to the audio file to partition the audio file.

[0010] Optionally, in some embodiments, determining the first audio characteristic includes at least one of a peak value, root mean square value, spectral centroid, frequency domain magnitude shape, harmonic content, stochastic content and pitch, or magnitude.

[0011] Optionally, in some embodiments, mapping the first audio segment to a dimensional space based on the first audio characteristic includes: generating, via the processor, an array of audio segment metadata, wherein the array includes audio segment metadata of the first audio segment; sorting, via the processor, the array based on the first audio characteristic; generating, via the processor, a mapping function based on the sorted array; and mapping, via the processor, the audio segment metadata of the first audio segment to the dimensional space based on the mapping function.

[0012] Optionally, in some embodiments, the audio segment metadata of the first audio segment includes the first audio characteristic of the first audio segment and a reference pointer to the first audio segment.

[0013] Optionally, in some embodiments, sorting the array based on the first audio characteristic includes sorting the array based on a value of the first audio characteristic.

[0014] Optionally, in some embodiments, sorting the array based on the first audio characteristic includes further sorting the array based on an additional audio characteristic of the first audio segment.

[0015] Optionally, in some embodiments, generating a mapping function based on the sorted array includes calculating a scaled nonlinear mapping function that approximates the distribution of the sorted array.

[0016] Optionally, in some embodiments, generating a mapping function based on the sorted array includes implementing a neural network that approximates the distribution of the sorted array.

[0017] Optionally, in some embodiments, mapping the audio segment metadata of the first audio to a dimensional space based on the mapping function includes: determining one or more dimensions of analysis, wherein each dimension of analysis is based on an audio characteristic; and applying the mapping function to the audio segment metadata to distribute the audio segment metadata in each of the one or more dimensions of analysis.

[0018] Optionally, in some embodiments, determining the second audio characteristic of a second audio segment based on the analysis of the second audio segment includes: retrieving a second audio file in real time from an audio input device; partitioning the second audio file into a second audio segment; and determining at least one of a peak value, root mean square value, spectral centroid, frequency domain magnitude shape, harmonic content, stochastic content and pitch, or magnitude of the second audio segment.

[0019] Optionally, in some embodiments, identifying the first audio segment in the dimensional space based on the second audio characteristic includes: applying the mapping function to the second audio characteristic of the second audio segment to approximate the position of the second audio segment in the dimensional space; determining the dimensional distance between the second audio segment and one or more audio segments in the dimensional space; and identifying the first audio segment based on the dimensional distance, wherein the dimensional distance between the second audio segment and the first audio segment is lesser than the dimensional distance between the second audio segment and each other audio segment of the one or more audio segments in the dimensional space.

[0020] Optionally, in some embodiments, modifying the second audio segment based on the first audio segment includes: receiving a user input of the audio inertia value; modifying the audio data of the second audio segment to incorporate the first audio characteristic of the first audio segment; and modifying the audio data of the second audio segment to include audio data of the first audio segment based on the audio inertia value.

[0021] Optionally, in some embodiments, outputting the second audio segment includes outputting the modified second audio segment via an audio output device.

[0022] Optionally, in some embodiments, outputting the second audio segment includes outputting the modified second audio segment in real time.

[0023] Optionally, in some embodiments, the second audio segment is an audio segment partitioned from a second audio file, and outputting the second audio segment further includes modifying and outputting one or more additional audio segments partitioned from the second audio file, such that the output includes the entirety of the second audio file.

[0024] In one embodiment, a system for audio style transfer, includes: an audio file database; and a processor configured by instructions to perform operations including: partitioning a first audio file into a first audio segment, wherein the first audio segment includes a portion of the first audio file; determining a first audio characteristic of the first audio segment based on an analysis of the first audio segment; mapping the first audio segment to a dimensional space based on the first audio characteristic; determining a second audio characteristic of a second audio segment based on an analysis of the second audio segment; identifying the first audio segment in the dimensional space based on the second audio characteristic; modifying the second audio segment based on the first audio segment; and outputting the second audio segment.

[0025] Optionally, in some embodiments, partitioning the first audio file into the first audio segment includes applying a window function to the audio file to partition the audio file.

[0026] Optionally, in some embodiments, determining the first audio characteristic includes at least one of a peak value, root mean square value, spectral centroid, frequency domain magnitude shape, harmonic content, stochastic content and pitch, or magnitude.

[0027] Optionally, in some embodiments, mapping the first audio segment to a dimensional space based on the first audio characteristic includes: generating an array of audio segment metadata, wherein the array includes audio segment metadata of the first audio segment; sorting the array based on the first audio characteristic; generating a mapping function based on thesorted array; and mapping the audio segment metadata of the first audio segment to the dimensional space based on the mapping function.

[0028] Optionally, in some embodiments, the audio segment metadata of the first audio segment includes the first audio characteristic of the first audio segment and a reference pointer to the first audio segment.

[0029] Optionally, in some embodiments, sorting the array based on the first audio characteristic includes sorting the array based on a value of the first audio characteristic.

[0030] Optionally, in some embodiments, sorting the array based on the first audio characteristic includes further sorting the array based on an additional audio characteristic of the first audio segment.

[0031] Optionally, in some embodiments, generating a mapping function based on the sorted array includes calculating a scaled nonlinear mapping function that approximates the distribution of the sorted array.

[0032] Optionally, in some embodiments, generating a mapping function based on the sorted array includes implementing a neural network that approximates the distribution of the sorted array.

[0033] Optionally, in some embodiments, mapping the audio segment metadata of the first audio to a dimensional space based on the mapping function includes: determining one or more dimensions of analysis, wherein each dimension of analysis is based on an audio characteristic; and applying the mapping function to the audio segment metadata to distribute the audio segment metadata in each of the one or more dimensions of analysis.

[0034] Optionally, in some embodiments, determining the second audio characteristic of a second audio segment based on the analysis of the second audio segment includes: retrieving a second audio file in real time from an audio input device; partitioning the second audio file into a second audio segment; and determining at least one of a peak value, root mean square value, spectral centroid, frequency domain magnitude shape, harmonic content, stochastic content and pitch, or magnitude of the second audio segment.

[0035] Optionally, in some embodiments, identifying the first audio segment in the dimensional space based on the second audio characteristic includes: applying the mapping function to the second audio characteristic of the second audio segment to approximate the position of the second audio segment in the dimensional space; determining the dimensionaldistance between the second audio segment and one or more audio segments in the dimensional space; and identifying the first audio segment based on the dimensional distance, wherein the dimensional distance between the second audio segment and the first audio segment is lesser than the dimensional distance between the second audio segment and each other audio segment of the one or more audio segments in the dimensional space.

[0036] Optionally, in some embodiments, modifying the second audio segment based on the first audio segment includes: receiving a user input of the audio inertia value; modifying the audio data of the second audio segment to incorporate the first audio characteristic of the first audio segment; and modifying the audio data of the second audio segment to include audio data of the first audio segment based on the audio inertia value.

[0037] Optionally, in some embodiments, outputting the second audio segment includes outputting the modified second audio segment via an audio output device.

[0038] Optionally, in some embodiments, outputting the second audio segment includes outputting the modified second audio segment in real time.

[0039] Optionally, in some embodiments, the second audio segment is an audio segment partitioned from a second audio file, and outputting the second audio segment further includes modifying and outputting one or more additional audio segments partitioned from the second audio file, such that the output includes the entirety of the second audio file.

[0040] In one embodiment, a method for creating a stylistic replica of an input audio file is disclosed. The method includes segmenting by a processing element a style file into a plurality of style grains; analyzing by the processing element the plurality of style grains to identify respective style acoustic characteristics for each of the plurality of style grains; mapping the style grains to a dimensional space based on the respective style acoustic characteristics; segmenting by the processing element an input file into a plurality of input grains; analyzing the plurality of input grains to identify respective input acoustic characteristics for each of the plurality of input grains; mapping by identification in the dimensional space a match between the respective style grains and the plurality of input grains based on the input acoustic characteristics and the style acoustic characteristics; and generating an style transfer file based on the mapping, wherein the style transfer file is representative of the input file and in the acoustic style of the style file.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS

[0041] FIG. 1 illustrates an example of a computer-implemented audio style transfer system.

[0042] FIG. 2 is a flow diagram for mapping a base file to a dimensional space with the audio style transfer system of FIG. 1.

[0043] FIG. 3 is a flow diagram for resynthesizing an input file with the audio style transfer system of FIG. 1.

[0044] FIG. 4 illustrates an example user interface for displaying an audio style transfer module as described with respect to method 300.

[0045] FIG. 5 is a block diagram of an example computer system suitable for use in the audio style transfer system of FIG. 1.DETAILED DESCRIPTION

[0046] Disclosed herein is a process that assembles segments of an audio file to create a new audio signal that perceptually resembles a live audio input, in real-time, with intuitive user control, low latency and low computational resource use. The process can be considered a granular synthesis-based form of Style Transfer - it maps signal features from one audio signal on to another to create a new, hybrid sound that combines perceptual aspects of both. Stated differently, the techniques described enable creation of new audio signals (e.g., new audio files) that correlate to another audio file (e.g., live audio input or stream) to replicate an original audio file but with different acoustic properties and can be done at a low latency, and with low computational resources required, while also allowing a user full customization.

[0047] For example, the audio style transfer system may resynthesize a live input or audio file (hereinafter “input file” or “original file”), based on base characteristics of another audio file (hereinafter “base file” or “style file”). The resynthesis may include preparation operations and execution operations. The preparation operations may be performed offline before the resynthesis of the input file; the execution operations may be performed offline or in real time, as may be desired based on computational resources and timing.

[0048] Turning now to the figures, FIG. 1 illustrates an example system 100. The system 100 resynthesizes an audio input file with the audio style of an audio base file and transmits the resynthesized audio file to a user 126 via a user device 106. The system 100 includes a user device 106, an audio data store 110, an audio mixer system 112, in communication with an audio style transfer system 102 either directly or via a network 104. In some embodiments, theaudio style transfer system 102 includes a memory 116 and a processor 114. The memory 116 may include or access various types of data or instructions used by the audio style transfer system 102. Such data and instructions may include the base file data 118, input file data 120, audio mapping data 122, and audio style transfer instructions 124 in various examples. Such data and instructions may be stored on and / or executed by a computing device as described with respect to FIG. 5.

[0049] The audio style transfer system 102 and audio mixer system 112 is accessible by a user 126 through a user interface 108 provided by the user device 106, e.g., through a software application that generates a visual output on a display. In some embodiments, the audio style transfer system 102 may be in communication with one or more user devices 106, one or more audio data stores 110, and one or more audio mixer systems 112. In some embodiments, the audio style transfer system 102, audio data store 110, and audio mixer system 112 may be incorporated into the user device 106 as an application rather than as a separate system.

[0050] In some embodiments, a user 126 may engage with the system 100 through a user device 106. In some examples, the user 126 may be an end user seeking to resynthesize or otherwise transfer stylistic characteristics of an audio file with the system 100. For example the user 126 may be an end user interacting with the system 100 to modulate their voice in real time with signal features from a separate audio file. The system 100 may receive an audio input file from the user 126, transfer the audio style of the input file, and transmit the resynthesized audio file to the user 126 via the user interface 108.

[0051] In some embodiments, the user device 106 may be a device utilized by a user 126 (e.g., computer, tablet, smart phone, or the like). The user device 106 may communicate with the audio style transfer system 102 (e.g., via the network 104 or via a software installed locally on the user’s device). The user device 106 and network 104 are discussed in more detail with respect to FIG. 5. In some examples, the audio style transfer system 102 is executed on the user device 106. In such examples, communication between the audio style transfer system 102 and user device 106 may not be via network 104. In some examples, a user 126 may input a request to generate audio style transfer for an audio input file through the user interface 108. The user device 106 may communicate the request to the audio style transfer system 102. The audio style transfer system 102 may generate the audio style transfer for the audio input file and transmit the audio file to the user 126 via the user interface 108. The user interface 108 is discussed in more detail with respect to FIG. 4.

[0052] In some embodiments, the audio style transfer system 102 may be in communication with an audio data store 110. The audio data store 110 may include memory storage (e.g., in a server) for storing audio data. For example, the audio data store 110 may include data of audio base files. As used herein, audio base files represent audio files that may include base audio styles used in an audio style transfer to modify the audio style of an audio input file. The audio data store 110 may additionally include data of audio input files. As used herein audio input files represent audio files that may be resynthesized with the audio style of an audio base file. The audio data store 110 may be implemented as one storage device (e.g., physical device) or distributed across various storage devices.

[0053] In some embodiments, the audio style transfer system 102 may be in communication with an audio mixer system 112. The audio mixer system 112 may include audio mixer hardware systems, audio mixer software systems, audio modulator systems, or other such systems configured to modify audio signals. For example, the audio mixer system 112 may include an audio mixer software system configured to receive audio input signals, modify the audio input signals based on user input, and transmit or output the modified audio signals to the user 126. In some examples, the audio mixer system 112 may be hosted internally by the audio style transfer system 102, and in other examples, the audio mixer system 112 may be hosted externally (e.g., on a third-party website). In some examples, the audio style transfer system 102 may be incorporated as a module within the audio mixer system 112 and configured to modify audio signals in conjunction with the audio mixer system 112.

[0054] In some embodiments, the audio style transfer system 102 includes base file data 118 stored e.g., on the memory 116. The base file data 118 may store data related to audio base files. For example, base file data 118 may include audio data (e.g., audio signal data, frequency data, wavelength data) of an audio base file. The base file data 118 may additionally include audio data of grains of an audio base file. As used herein, a grain represents an audio segment of a sub-portion of an audio file. For example, a grain of an audio base file may be a twenty millisecond sub-portion of the base file. As can be appreciated, the time period for the grain may be varied as desired, e.g., 10ms, 30ms, 40ms, etc. The size of the audio base file may be a tradeoff based on the desired alignment with input files with processing speed. However, for files that are more similar in characteristics, larger grains may be used. The base file data 118 may additionally include slice data of an audio base file. As used herein, a slice representsmetadata of a grain, such as audio characteristics of the grain and reference pointers of the grain (e.g., spectral peaks and RMS).

[0055] The audio style transfer system 102 may receive the base file data 118 from the user device 106 and / or audio data store 110 (e.g., via the network 104). For example, the audio style transfer system 102 may receive audio data of an audio base file from the audio data store 110; the audio style transfer system 102 may receive grain configuration data (e.g., metadata representing the length of a grain) from user input via the user interface 108. The audio style transfer system 102 may store the base file data 118 in memory 116.

[0056] In some embodiments, the audio style transfer system 102 includes input file data 120 stored e.g., on the memory 116. The input file data 120 may store data related to audio input files. For example, input file data 120 may include audio data (e.g., audio signal data, frequency data, wavelength data) of an audio input file. The input file may represent the file to be recreated or replicated but with the style of the base file.

[0057] The audio style transfer system 102 may receive the input file data 120 from the user device 106 and / or audio data store 110 (e.g., via the network 104). For example, the audio style transfer system 102 may receive audio data of an input file from the audio data store 110. In another example, the audio style transfer system 102 may receive real-time audio input data from a user device 106, such as through a live audio input. The audio style transfer system 102 may store the input file data 120 in memory 116.

[0058] In some embodiments, the audio style transfer system 102 includes audio style transfer instructions 124 stored e.g., on the memory 116. The audio style transfer instructions 124 may, when executed by the processor 114, generate and / or transmit audio style transfer of an input file based on a base file (e.g., as according to method 200 and / or method 300). The audio style transfer instructions 124 may include instructions to partition a base file of the base file data 118 into an audio segment, determine an audio characteristic of the audio segment, generate a sorted array of audio slices based on the audio characteristic, map the sorted array of audio slices to a dimensional space, determine an audio characteristic of an input file, identify an audio slice in the dimensional space based on the audio characteristic of the input file, modify the audio data of the input file to transfer the audio style of the audio slice, and output the modified input file to the user 126 (e.g., via the user interface 108). The audio style transfer instructions 124 are described in further detail with respect to FIG. 2 and FIG. 3.

[0059] In some embodiments, the audio style transfer system 102 includes audio mapping data 122 stored e.g., on the memory 116. The audio mapping data 122 may store mapping functions generated by the audio style transfer system 102 according to the audio style transfer instructions 124 (e.g., as described with respect to method 200). The audio mapping data 122 may additionally store data of audio grains and slices of audio base files that are mapped to the dimensional space. The mapping acts to translate grains of the input file to grains of the base file - e.g., to replicate the input file in the style of the base file. The mapping can be based on one or more characteristics and may include thresholds and best fit metrics that are used to determine the mapping.

[0060] While the data and instructions, such as the base file data 118, input file data 120, audio mapping data 122, and audio style transfer instructions 124 are shown in FIG. 1 as being stored in the memory 116, in some examples, the data and instructions may be stored at other memory resources of the audio style transfer system 102 and / or at locations remote from the audio style transfer system 102, such as various databases or data stores (e.g., the audio data store 110). In such examples, the memory 116 of the audio style transfer system 102 may include instructions for accessing such data and instructions from remote locations, including, for example, the locations of the data and / or specific queries used to retrieve data for use by the audio style transfer system 102. For example, where the base file data 118 is stored in the audio data store 110, memory 116 may include instructions for how to retrieve or access the data from the audio data store 110.

[0061] The audio style transfer system 102 may be implemented by or at a computing device or combinations of computing resources in various embodiments. In various examples, the audio style transfer system 102 may be implemented by one or more servers, cloud computing resources, and / or other computing devices. The audio style transfer system 102 may, for example, be incorporated as a module within a mobile application, software application, or a website presented through a web browser (e.g., at a laptop or desktop computer), and the like. It should be also be noted that the output of the audio style transfer system (e.g., a modified or emulated input file that is recreated in another file) may be stored as an audio file for transfer and playback via other systems and / or may be accessed and playback via an I / O element, such as a speaker, headphones, or the like.

[0062] The components of FIG. 1 are exemplary only. In various examples, the audio style transfer system 102 may communicate with and / or include additional components and / orfunctionality not shown in FIG. 1. For example, the audio style transfer system 102 may communicate with multiple audio mixer system 112, such as different hardware and software systems.

[0063] FIG. 2 illustrates an example method 200 for mapping base file data 118 to a dimensional space with the audio style transfer system 102. The method 200 may partition, analyze, and map base file data 118 into a dimensional space based on one or more audio characteristics of a base file.

[0064] In operation 202, the audio style transfer system 102 receives base file data 118. The 102 may receive the base file data 118 from an audio data store 110 and / or user device 106 (e.g., via the network 104). In some examples, a user 126 may upload or select an audio base file via the user interface 108, and the user device 106 may communicate the base file to the audio style transfer system 102. For example, FIG. 4 portrays an example user interface 108 of the audio style transfer system 102. As portrayed with respect to FIG. 4, a user 126 may interact with the base file selector 404 to select a base file. The user 126 may select a base file based on an audio characteristic or style of the base file. For example, where the user 126 wishes to modify an input speech audio such that the speech audio has a style transfer effect of a dog barking, the user 126 may select a base file of audio of a dog barking. The user device 106 may then transmit the audio file of the dog barking to the audio style transfer system 102. In other examples, the audio style transfer system 102 may receive the base file data 118 from an audio data store 110, such as a database of audio files. The audio style transfer system 102 may store the base file data 118 in memory 116.

[0065] In operation 204, the audio style transfer system 102 partitions a base file into one or more audio segments (e.g., grains). The audio style transfer system 102 may partition the base file selected by the user 126 into one or more grains, such that audio data of the base file is segmented into the one or more grains. In partitioning the base file, the audio style transfer system 102 may apply a window function or smoothing function to the audio signal of the base file to avoid the effects of a hard cut. In some examples, the length the audio included in each grain is determined by a minimum and / or maximum length threshold. In other examples, the length of the audio included in each grain is determined by a user input received from the user device 106. For example, as portrayed with respect to FIG. 4, the user 126 may interact with a grain length selector 406 to select an audio length for grains of the base file. Where the user 126 selects a grain length of twenty milliseconds, the audio style transfer system 102 maypartition the base file into audio segments that are twenty milliseconds in length. Each grain may include all audio data of the corresponding audio segment, including audio properties, sound wave data, and the like. The audio style transfer system 102 may store the one or more grains in memory 116, e.g., in base file data 118.

[0066] In operation 206, the audio style transfer system 102 analyzes each audio segment (e.g., grain) of the one or more audio segments to determine an audio characteristic of the audio segment. Audio characteristics may include characteristics of the sound wave such as frequency, amplitude, and wavelength, and / or characteristics of the audio signal. For example, the audio style transfer system 102 may analyze each grain of the base file to determine various acoustic characteristics, such as one or more of a peak value of the grain, root mean square (RMS) value of the grain, level of the grain, spectral centroid of the grain, frequency domain magnitude shape, harmonic content, stochastic content and pitch, and / or more complex measurements, that for example describe the frequency domain magnitudes of the grain. In some examples, the analysis of the grain is configured to describe in a compressed form how humans perceive the grain psychoacoustically. The closer the analysis approximates the perception of the grain, the more exactly the resynthesis will perceptually resemble the audio input file.

[0067] In operation 208, the audio style transfer system 102 generates an audio slice based on the audio characteristic of the audio segment. The audio slice may include data representing the audio characteristic of the audio segment and a reference pointer indicating the location of the audio segment in the audio base file. For example, where an audio grain contains the first twenty milliseconds of audio of a base file, the corresponding slice of the grain may contain the RMS and peak value metrics of the audio grain and a reference to the base file indicating that the audio of the grain is located in the first twenty milliseconds of the base file.

[0068] In some examples, the slice does not include the audio data of the corresponding grain. In such examples, the absence of audio data in the slice reduces the memory required to store the slice (e.g., the audio characteristic values may be the only data stored). By lowering the memory requirements of the slice, the style transfer system 102 may decrease the memory access time and processing time needed to sort and / or map the slice in the dimensional space. Decreasing the processing time may lower the latency for generating a resynthesized audio file and allow for real-time output of the style transfer of audio. This may allow the style transfer to occur more quickly and efficiently than in other configurations.

[0069] More specifically, in embodiments where the grain is separated from the slice (e.g., the audio signal is separated from the stored characteristics), may help to take advantage of caching operations for typical computing devices. In these examples, there may be multiple levels of cache plus the working memory as a whole, and the more data needed to access for a specific operation the slower processing processes may be and data that is not in the smallest but fastest (e.g., LI -Cache) increases speed. Having to access lower caching levels or the whole working memory (RAM) in an unpredictable manner (e.g., cache miss) and the CPU has to wait for a long time (relative to conventional processing times) and the processing may be delayed (e.g., not output in real time).

[0070] In short, in many systems, separating Grain and Slice data allows faster access to the slices for mapping purposes, because there is no audio data with them, so they are smaller in working memory. In other embodiments, however, it should be noted that the grain and slice data can be stored together as well.

[0071] In operation 210, the audio style transfer system 102 generates a sorted array of one or more audio slices based on one or more audio characteristics of the one or more audio slices. For example, where a base file includes a plurality of grains and corresponding audio slices and the grains are analyzed to determine RMS values, the audio style transfer system 102 may generate an array of the audio slices, sorting the slices in ascending order based on the RMS values. In some examples, the audio style transfer system 102 may sort the slices in the array based on a plurality of audio characteristics. In such examples, the audio style transfer system 102 may generate a multi-dimensional sorted array, where each dimension of the array represents an audio characteristic. For example, if the audio style transfer system 102 analyzes grains of a base file for the RMS and peak value of the grains, the sorted array may be a two- dimensional array, where the respective slices are sorted by RMS value in one dimension, and by peak value in the second dimension. If audio characteristic includes multiple dimensions, such as a discrete cosine transform, the audio style transfer system 102 may generate a multidimensional sorted array with dimensions for each dimension of the audio characteristic.

[0072] In some examples, the audio style transfer system 102 may generate a sorted array of predetermined dimensions and sort the slices in the array based on a predetermined mode of sorting and predetermined weight values. In other examples, the audio style transfer system 102 may receive a user input indicating the dimensions of the sorted array. For example, as portrayed with respect to FIG. 4, a user 126 may interact with the spectral mode selector 412 ofthe user interface 108 to select a mode of sorting the sorted array. The user 126 may select one or more audio characteristics to sort the grains of the base file by. The user 126 may additionally interact with the spectral mode selector 412 to select a mode of sorting (e.g., ascending or descending) and assign a weight value to the selected audio characteristics. For example, after selecting a base file via the base file selector 404, the user 126 may interact with the spectral mode selector 412 to input that the grains of the base file should be sorted by RMS value in ascending order. The audio style transfer system 102 may receive the user 126 selection from the user device 106. The audio style transfer system 102 may generate a sorted array with a dimension for each audio characteristic selected by the user 126, and the audio style transfer system 102 may sort the audio slices according to the mode of sorting and weight values selected by the user. In other examples, the audio style transfer system 102 may additionally use a neural network to determine sorting characteristics of the sorted array, such as the one or more audio characteristics to sort the array by, the mode of sorting, and / or the weight values. The audio style transfer system 102 may store the sorted array in memory 116, e.g., in audio mapping data 122.

[0073] In operation 212, the audio style transfer system 102 generates a mapping function based on the sorted array, where the mapping function is configured to map the slices of the sorted array to a dimensional space based on the audio characteristics of the slices. For example, where the sorted array has N-dimensions, the mapping function may be configured to map the slices of the sorted array to an N-dimensional space, where each dimension represents an analysis dimension of the sorted array. As used herein, an analysis dimension of a sorted array represents a sorting dimension of the sorted array, including the audio characteristic, mode of sorting, and weight values.

[0074] In some examples, the audio style transfer system 102 may calculate a scaled nonlinear mapping function that approximates the distribution of the sorted values for each analysis dimension. If for example the base file is a linearly decaying sine wave, the sorted slices may indicate an approximately linear distribution of peak values. In some examples, natural base files do not have such a linear distribution but a skewed distribution, that may be approximated with a power function of the type: f(x) = xAw, with w being calculated form the median m of the distribution as: w = log(m) / log(0.5) to define the mapping. Depending on the analysis dimension the values may range from 0 to 1 or the audio style transfer system 102 may scale all values to a range from 0 to 1 before calculating the mapping.

[0075] In some examples, the audio style transfer system 102 may use a neural network to generate the mapping function (or replicating such a function) based on an audio characteristic. For example, the audio style transfer system 102 may use a neural network to approximate the distribution of slices in the dimensional space as neural networks can be interpreted as nonlinear approximators according to the universal approximation theorem. The audio style transfer system 102 may store the mapping function in memory 116, e.g., in audio mapping data 122.

[0076] In operation 212, the audio style transfer system 102 maps the sorted array of audio slices to a dimensional space based on the mapping function, where each dimension of the dimensional space represents an analysis dimension of the base file. For example, where an audio base file is partitioned into M grains, and the audio style transfer system 102 generates a sorted array of M slices corresponding to the M grains, sorting on RMS and peak value, the audio style transfer system 102 may map the M slices into a two-dimensional space, with a first dimension for RMS value and a second dimension for peak value.

[0077] FIG. 3 illustrates an example method 300 resynthesizing an audio input file with the audio style of a base file. The method 300 may modify the data of an audio input file based on an audio grain of the base file with a corresponding audio characteristic.

[0078] In operation 302, the audio style transfer system 102 receives an audio input file. The audio input file may be an audio file containing audio data for resynthesis. The audio style transfer system 102 may receive the input file from the user device 106 and / or the audio data store 110. For example, a user 126 may interact with the user interface 108 to upload an audio input file to the user device 106, record an audio input file via an i / o interface 506 (e.g., a microphone), or select an input file from an audio data store 110. The user device 106 may transmit the input file to the audio style transfer system 102. In some examples, the input file may include live audio. For example, the user 126 may record live audio via the user device 106, and the user device 106 may transmit the live audio to the audio style transfer system 102 in real time as the input file. The audio style transfer system 102 may store the input file in memory, e.g., in input file data 120.

[0079] In operation 304, the audio style transfer system 102 analyzes the input file to determine an audio characteristic of the input file. The audio style transfer system 102 may first partition the input file into one or more audio segments based on a temporal length. In some examples the audio style transfer system 102 may partition the input file into audio segmentsbased on the temporal length of the grains of the base file. For example, where the base file is partitioned into grains of twenty milliseconds, the audio style transfer system 102 may also partition the input file into audio segments of twenty milliseconds. In other examples, the temporal length of the input file audio segment may not match the temporal length of the grain of the base file.

[0080] The audio style transfer system 102 may analyze the audio segments of the input file to determine an audio characteristic of the audio segments. Audio characteristics may include characteristics of the sound wave such as frequency, amplitude, and wavelength, and / or characteristics of the audio signal. For example, the audio style transfer system 102 may analyze each audio segment of the input file to determine a peak value of the audio segment, root mean square (RMS) value of the audio segment, level of the audio segment, spectral centroid of the audio segment, frequency domain magnitude shape, harmonic content, stochastic content and pitch, and / or more complex measurements, that for example describe the frequency domain magnitudes of the audio segment. In some examples, the audio style transfer system 102 may analyze audio characteristics of the audio segment based on the analysis dimensions of the dimensional space of the base file (e.g., the dimensional space generated according to method 200). For example, where the dimensional space of the base file includes a first dimension for RMS value and a second dimension for peak value, the audio style transfer system 102 may analyze the audio segments of the input file to determine the RMS value and peak value of the input segments. In other examples, the audio style transfer system 102 may analyze the audio segments of the input file to determine one or more audio characteristics which differ from the audio characteristics of the dimensional space of the base file. The audio style transfer system 102 may store the audio segments and audio characteristics of the input file in memory 116, e.g., in input file data 120.

[0081] In some examples, the audio style transfer system 102 may further analyze the audio segments of the input file by mapping the audio segments to the dimensional space of the base file. The audio style transfer system 102 may additionally manipulate the dimensional range of the analysis of the input file to correspond with the dimensional range of the base file. The audio style transfer system 102 may learn the typical input range of each dimension to map the minimum and maximum value of the input file exactly to the minimum and maximum of the base file in the dimensional space, which will result in a larger part of the audio contained in the base file appearing in the resynthesized output. For example, where a first dimension of thedimensional space of the base file represents grains in a frequency range of 250 Hz to 500 Hz, the audio style transfer system 102 may map the audio segments of the input file to the dimensional space such that the lowest frequency audio segment is mapped to 250 Hz and the highest frequency audio segment is mapped to 500 Hz regardless of the actual frequencies of the audio segments. In other examples, the audio style transfer system 102 may further skew the analysis of the input file by using nonlinear functions controlled by the user 126, so that sections of the base file are used more frequently, or not at all.

[0082] In operation 306, the audio style transfer system 102 identifies an audio slice of the base file in the dimensional space of the base file based on the audio characteristic of the audio segment of the input file. For example, the audio style transfer system 102 may traverse the dimensional space of the base file to find the audio slice of the base file in closest proximity to the audio segment of the input file in the dimensional space. The audio style transfer system 102 may further identify the audio grain of the base file that corresponds to the identified audio slice.

[0083] In some examples, the audio style transfer system 102 may identify the audio slice by applying the inverse mapping function of the specific dimension (e.g., the inverse of the mapping function generated at operation 212 of method 200): f(x) = x^(l / w) to retrieve a linearized value from 0 to 1 for each dimension. The audio style transfer system 102 may index into the sorted slices of each dimension with a linearized value v and retrieve the index i of a target slice with the count of slices c as: i = round(c * v) to identify a matching slice. This allows efficient real-time processing, even with large base files, with intuitive user control and low computational resource use, by avoiding any search operations or comparisons between the input file and the slices of the base file.

[0084] The different mapping dimensions may be utilized in different ways. In some examples, the user 126 may interact with the user interface 108 to determine a main analysis dimension and use the other dimensions of the dimensional space only to re-identify a slice if any of those dimensions has a large difference to the identified slice. To re-identify a slice, the audio style transfer system 102 may identify one or more slices that are around or in proximity to the identified slice in the sorted main dimension.

[0085] For a multidimensional mapping that respects all dimensions equally, the audio style transfer system 102 may create a multidimensional function, that approximates the course of each analysis across all slices in the dimensional space. This function could for example be apiecewise linear function or a neural network. The audio style transfer system 102 may sort all slices, for example using a weighted sum of absolutes across all dimensions, such that the multidimensional function can be much simpler and approximation errors will be more negligible since nearby slices will provide a similarly good match across dimensions. The user 126 may select the weights of the weighted sums (e.g., by interacting with the spectral mode selector 412 as portrayed in FIG. 4) to influence the importance of the different factors, since it directly corresponds to the approximation error. To identify the slice, the audio style transfer system 102 may calculate an approximation function using the analysis results of the input file for each dimension, which then returns the best match.

[0086] In some examples, the characteristics can be weighted differently for sorting, according to user input, and / or implementation data. For example, the multidimensional mapping function can be thought of as f(xl, x2) where xl is a first characteristic and x2 is a second characteristic. The result of f(xl, x2) is a selected slice in one example and in a multidimensional case, the sorting is an additional feature that improves the consistency of the results, but can also be omitted. In various embodiments, however, the mapping may be performed oof two or more dimensions to one (2: 1) per selected slice.

[0087] In operation 308, the audio style transfer system 102 may modify the audio of the audio segment of the input file based on the identified audio slice of the base file. The audio style transfer system 102 may retrieve audio data of the grain (e.g., from memory 116) corresponding to the identified slice and modify the audio of the audio segment of to transfer the audio style of the identified slice to the audio segment. For example, the audio style transfer system 102 may overlay, mix, and / or fade in the audio corresponding with the identified slice with the audio segment of the input file. The audio style transfer system 102 may fade in the audio corresponding to the identified slice to create a clean and natural sound and may find the optimal fade points by using techniques known from advanced granular pitch shifters.

[0088] To achieve a more natural sounding result the audio style transfer system 102 may not modify the audio of every audio segments of the input file, but rather only modify the audio of the audio segment if the content and / or audio characteristics of the audio segment and the content of the base file that immediately proceeds the identified slice diverges above an inertia value (e.g., the inertia value acts as a threshold for the “fit” or match). The objective of the inertia value is to play grains of the bases file in their natural ordering as long as the resynthesis remains accurate enough, which preserves larger scale structures of the base file,which creates a more natural sounding output. For example, where the audio characteristics of an audio segment of the input file closely matches the audio characteristics of one or more grains directly proceeding the identified grain in the base file, the audio style transfer system 102 may continue to use the grains in the natural ordering of the base file rather than identify additional slices for use in modifying the input file audio. Where the inertia value is low, the proceeding grains must be more similar to the input file audio, and where the inertia value is high, the proceeding grains may be less similar to the input file audio.

[0089] In some examples, the audio style transfer system 102 may receive the inertia value from the user 126 (e.g., via the user device 106). For example, as portrayed with respect to FIG. 4, the user 126 may interact with an inertia selector 408 to input a desired inertia value. The user device 106 may communicate the user input of the inertia value to the audio style transfer system 102. In other examples, the inertia value may be a preset value stored in memory 116 or may be a value determined by a neural network.

[0090] It should be noted that the modified input file (e.g., the replicated file) may be a combination of multiple segments that are combined together to replicate the entire time lengths (or a desired subset thereof) of the input file. As such, it should be understood that while the mapping is described on a segment by segment basis (e.g., grain by grain) that the output file provided to the user may combine multiple grains together after mapping to recreate the input file, but in the style of the base file - e.g., a human speaking but with the style of a barking dog.

[0091] In operation 310, the audio style transfer system 102 may transmit and / or output the modified audio of the input file to the user 126, e.g., via the user device 106 and an acoustic output of the device. In some examples, the audio style transfer system 102 may the modified input file as data to the user 126. In some examples, the audio style transfer system 102 may playback the audio of the modified input file via an i / o interface 506 (e.g., a speaker). In some examples, the playback of the modified input file may be real-time playback, where the latency between the input of the input file and the playback of the modified input file is sufficiently low to constitute real-time or live playback. For example, the user 126 may record live speech as input to the audio style transfer system 102, and the audio style transfer system 102 may playback the modified speech with style transfer in real time. In other examples, the style transfer system may operate from prerecorded files and not in real time. For example, the system may be applied to long acoustic content, such as movie dialogue or databases of songsand the processing may be done without having to wait for the content to be loaded into the system, rather the system can access the content via memory or the like.

[0092] FIG. 4 illustrates an example user interface 108 for displaying an audio style transfer system 102. As described with respect to method 200 and method 300, the user interface 108 may include a base file selector 404, grain length selector 406, inertia selector 408, and spectral mode selector 412 in a style transfer module 402, where the style transfer module 402 is a user interface module configured to provide the user 126 with an audio style transfer experience.The style transfer module 402 may be a standalone application or may be a module within or in conjunction with one or more other applications. For example, the style transfer module 402 may be a module configured within an audio editing application, and may be displayed in conjunction with an audio mixer module 410. The style transfer module 402 may be in communication with the audio mixer module 410, and the audio mixer module 410 may modify the audio input to or output from the style transfer module 402. For example, the audio mixer module 410 may include audio editing functionality and the user 126 may edit the audio base file via the audio mixer module 410 before selecting the base with the base file selector 404 in the style transfer module 402. The components of FIG. 4 are exemplary only. In various examples, the user interface 108 may be configured to display and / or include additional components and / or functionality not shown in FIG. 4.

[0093] In some embodiments, the system and methods described herein may be configured to generate a new or emulated file (e.g., the modified input file). It should be noted that various steps and operations related to the portioning of grains from the original input file may be done for an n number of grains, depending on the length of the file. In this manner, the output file, e.g., the modified file, may be recreated by modifying the n grains based on the base file to create the modified file that has the same or substantially the same time length as the original one (e.g., by modifying n grains to match the n number of input grains). The overall new file may then be created that is the combination of the modified n grains combined together.

[0094] FIG. 5 illustrates a block diagram of an example computer system suitable for use in embodiments disclosed herein in accordance with an embodiment of the disclosure. For example, the audio style transfer system 102 may include or utilize one or several computing systems 500, and the processor 114 and memory 116 may be located at one or several computing systems 500. In various embodiments, the audio mixer system 112 is implemented by a computing system 500. In various implementations, the user device 106 and / or additionaluser devices may be implemented using any number of computing devices including, but not limited to a computer, laptop, tablet, mobile phone, smart phone, wearable device (e.g., AR / VR headset, smartwatch, smart glasses, or the like), smart speaker, vehicle (e.g., automobile), or appliance.

[0095] This disclosure contemplates any suitable number of computing systems 500. For example, the computing system 500 may be a server, a desktop computing system, a mainframe, a mesh of computing systems, a laptop or notebook computing system, a tablet computing system, an embedded computer system, a system-on-chip, a single-board computing system, or a combination of two or more of these. Where appropriate, the computing system 500 may include one or more computing systems; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which may include one or more cloud components in one or more networks. The computing system 500 may include one or more processors 502, an input / output (i / o) interface 506, one or more external devices 508, one or more memory components 510, and a network interface 512. Each of the various components may be in communication with one another through one or more buses or communication networks, such as wired or wireless networks.

[0096] In some embodiments, various components of the computing system 500 may communicate with one another through the network 104. For example, in some embodiments, the computing system 500 may be implemented as a serverless service, where computing resources for various components of the computing system 500 may be located across various computing environments (e.g., cloud platforms) and may be reallocated dynamically and / or automatically according to, for example resource usage of the computing system 500. In various implementations, the computing system 500 may be implemented using organizational processing constructs such as functions implemented by worker elements allocated with compute resources, containers, virtual machines, and the like.

[0097] The processor 114 may be any type of electronic device capable of processing, receiving, and / or transmitting instructions. For example, the processor 114 may be a central processing unit, graphics processing unit, microprocessor, processor, or microcontroller. Additionally, it should be noted that some components of the computing system 500 may be controlled by a first processor and other components may be controlled by a second processor, where the first and second processors may or may not be in communication with each other. The audio style transfer system 102, audio mixer system 112, and user device 106 may perform 1operations by executing executable instructions (e.g., software) using the processor 502. The processor 114 may be used to implement processor 114 shown in FIG. 1.

[0098] The i / o interface 506 allows a user to enter data in to computing system 500, as well as provides an input / output for the computing system 500 to communicate with other devices or services. The i / o interface 506 can include one or more input buttons, touch pads, and so on.

[0099] The external devices 508 are one or more devices that can be used to provide various inputs to the computing system 500, e.g., mouse, microphone, keyboard, trackpad, or the like. The external devices 508 may be local or remote and may vary as desired. In some examples, the external devices 508 may also include one or more additional sensors.

[0100] The memory components 510 are used by the computing system 500 to store instructions for the processor 114 and may be implemented as a data store and the like. The memory components 510 may be, for example, magneto-optical storage, read-only memory, random access memory, erasable programmable memory, flash memory, or a combination of one or more types of memory components. The memory components 510 may be used to implement the memory 116 shown in FIG. 1. The memory 116 may include various instructions for various functions of the audio style transfer system 102 which, when executed by the processor 114, perform various functions of the audio style transfer system 102. The memory 116 may further store data and / or instructions for retrieving data used by the audio style transfer system 102. Similar to the processor 114, the memory 116 utilized by the audio style transfer system 102 may be distributed across various physical computing devices. In some examples, the memory 116 may access instructions and / or data from other devices or locations, and such instructions and / or data may be read into memory 116 to implement the audio style transfer system 102.

[0101] The network interface 512 provides communication to and from the computing system 500 to other devices. The network interface 512 includes one or more communication protocols, such as, but not limited to WI-FI®, Ethernet, BLUETOOTH®, and so on. The network interface 512 may also include one or more hardwired components, such as a Universal Serial Bus (USB) cable, or the like. The configuration of the network interface 512 depends on the types of communication desired and may be modified to communicate via WIFI ®, BLUETOOTH®, and so on.

[0102] The network interface 512 may interface with the network 104. The network 104 may be implemented using one or more wired and / or wireless systems and protocols forcommunications between computing devices. In various embodiments, the network 104 or various portions of the network 104 may be implemented using the internet, a local area network, a wide area network, and / or other networks. In addition to traditional data networking protocols, in some embodiments, data may be communicated according to protocols and / or standards including near field communication, Bluetooth®, Wi-Fi, cellular connections, or the like.

[0103] The display 504 provides a visual output for the computing devices and may be varied as needed based on the device. The display 504 may be configured to provide visual feedback to the user and may include a liquid crystal display screen, light emitting diode screen, plasma screen, or the like. In some examples, the display 504 may be configured to act as an input element for the user through touch feedback or the like.

[0104] The components in FIG. 5 are exemplary only. In various examples, the computing system 500 may include additional components and / or functionality not shown in FIG. 5.

[0105] Accordingly, the audio style transfer system 102 described herein addresses particular challenges and needs presented by systems for generating audio style transfer. For example, traditional audio style transfer methods require encoder training for each base file or comparisons between signals of the base file and input file. These traditional methods require long processing times and may not allow users to intuitively interact with the system to produce a desired result. In contrast, the audio style transfer system 102 described herein maps audio segments to a dimensional space allowing for low-latency audio style transfer, and the audio style transfer system 102 includes user controls that allow users 126 to intuitively control the real-time synthesis and style transfer of audio.

[0106] The technology described herein may be implemented as logical operations and / or modules in one or more systems. The logical operations may be implemented as a sequence of processor-implemented steps directed by software programs executing in one or more computer systems and as interconnected machine or circuit modules within one or more computer systems, or as a combination of both. Likewise, the descriptions of various component modules may be provided in terms of operations executed or effected by the modules. The resulting implementation is a matter of choice, dependent on the performance requirements of the underlying system implementing the described technology. Accordingly, the logical operations making up the embodiments of the technology described herein are referred to variously as operations, steps, objects, or modules. Furthermore, it should be understood that logicaloperations may be performed in any order, unless explicitly claimed otherwise or a specific order is inherently necessitated by the claim language.

[0107] In some implementations, articles of manufacture are provided as computer program products that cause the instantiation of operations on a computer system to implement the procedural operations. One implementation of a computer program product provides a nontransitory computer program storage medium readable by a computer system and encoding a computer program. It should further be understood that the described technology may be employed in special purpose devices independent of a personal computer.

[0108] The description of certain embodiments included herein is merely exemplary in nature and is in no way intended to limit the scope of the disclosure or its applications or uses. In the included detailed description of embodiments of the present systems and methods, reference is made to the accompanying figures which form a part hereof, and which are shown by way of illustration specific to embodiments in which the described systems and methods may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice presently disclosed systems and methods, and it is to be understood that other embodiments may be utilized, and that structural and logical changes may be made without departing from the spirit and scope of the disclosure. Moreover, for the purpose of clarity, detailed descriptions of certain features will not be discussed when they would be apparent to those with skill in the art so as not to obscure the description of embodiments of the disclosure. The included detailed description therefore is not to be taken in a limiting sense, and the scope of the disclosure is defined only by the appended claims.

[0109] From the foregoing it will be appreciated that, although specific embodiments of the invention have been described herein for purposes of illustration, various modifications may be made without deviating from the spirit and scope of the invention.

[0110] Although the methods described herein (e.g., method 200 and method 300) depict a particular sequence of operations, the sequence may be altered without departing from the scope of the present disclosure. For example, some of the operations depicted may be performed in parallel or in a different sequence that does not materially affect the function of the routine. In other examples, different components of an example device or system that implements the routine may perform functions at substantially the same time or in a specific sequence.

[0111] The particulars shown herein are by way of example and for purposes of illustrative discussion of the preferred embodiments of the present disclosure and are presented in the cause of providing what is believed to be the most useful and readily understood description of the principles and conceptual aspects of various embodiments of the invention. In this regard, no attempt is made to show structural details of the invention in more detail than is necessary for the fundamental understanding of the invention, the description taken with the figures and / or examples making apparent to those skilled in the art how the several forms of the invention may be embodied in practice.

[0112] As used herein and unless otherwise indicated, the terms “a” and “an” are taken to mean “one”, “at least one” or “one or more”. Unless otherwise required by context, singular terms used herein shall include pluralities and plural terms shall include the singular.

[0113] Unless the context clearly requires otherwise, throughout the description and the claims, the words ‘comprise’, ‘comprising’, and the like are to be construed in an inclusive sense as opposed to an exclusive or exhaustive sense; that is to say, in the sense of “including, but not limited to”. Words using the singular or plural number also include the plural and singular number, respectively. Additionally, the words “herein,” “above,” and “below” and words of similar import, when used in this application, shall refer to this application as a whole and not to any particular portions of the application.

[0114] All relative, directional, and ordinal references (including top, bottom, side, front, rear, first, second, third, and so forth) are given by way of example to aid the reader’s understanding of the examples described herein. They should not be read to be requirements or limitations, particularly as to the position, orientation, or use unless specifically set forth in the claims. Connection references (e.g., attached, coupled, connected, joined, and the like) are to be construed broadly and may include intermediate members between a connection of elements and relative movement between elements. As such, connection references do not necessarily infer that two elements are directly connected and in fixed relation to each other, unless specifically set forth in the claims.

[0115] Of course, it is to be appreciated that any one of the examples, embodiments or processes described herein may be combined with one or more other examples, embodiments and / or processes or be separated and / or performed amongst separate devices or device portions in accordance with the present systems, devices and methods.

[0116] Finally, the above discussion is intended to be merely illustrative of the present system and should not be construed as limiting the appended claims to any particular embodiment or group of embodiments. Thus, while the present system has been described in particular detail with reference to exemplary embodiments, it should also be appreciated that numerous modifications and alternative embodiments may be devised by those having ordinary skill in the art without departing from the broader and Intended spirit and scope of the present system as set forth in the claims that follow. Accordingly, the specification and figures are to be regarded in an illustrative manner and are not intended to limit the scope of the appended claims.

Claims

CLAIMS1. A computer implemented method for audio style transfer, comprising: partitioning, via a processor, a first audio file into a first audio segment, wherein the first audio segment comprises a portion of the first audio file; determining, via the processor, a first audio characteristic of the first audio segment based on an analysis of the first audio segment; mapping, via the processor, the first audio segment to a dimensional space based on the first audio characteristic; determining, via the processor, a second audio characteristic of a second audio segment based on an analysis of the second audio segment; identifying, via the processor, the first audio segment in the dimensional space based on the second audio characteristic; modifying, via the processor, the second audio segment based on the first audio segment; and outputting, via the processor, the second audio segment.

2. The computer implemented method of claim 1, wherein partitioning the first audio file into the first audio segment comprises applying a window function to the audio file to partition the audio file.

3. The computer implemented method of claim 1 or 2, wherein determining the first audio characteristic comprises at least one of a peak value, root mean square value, spectral centroid, frequency domain magnitude shape, harmonic content, stochastic content and pitch, or magnitude.

4. The computer implemented method of any one of claims 1 to 3, wherein mapping the first audio segment to a dimensional space based on the first audio characteristic comprises: generating, via the processor, an array of audio segment metadata, wherein the array comprises audio segment metadata of the first audio segment; sorting, via the processor, the array based on the first audio characteristic; generating, via the processor, a mapping function based on the sorted array; andmapping, via the processor, the audio segment metadata of the first audio segment to the dimensional space based on the mapping function.

5. The computer implemented method of claim 4, wherein the audio segment metadata of the first audio segment comprises the first audio characteristic of the first audio segment and a reference pointer to the first audio segment.

6. The computer implemented method of claim 4 or 5, wherein sorting the array based on the first audio characteristic comprises sorting the array based on a value of the first audio characteristic.

7. The computer implemented method of any one of claims 4 to 6, wherein sorting the array based on the first audio characteristic comprises further sorting the array based on an additional audio characteristic of the first audio segment.

8. The computer implemented method of any one of claims 4 to 7, wherein generating a mapping function based on the sorted array comprises calculating a scaled nonlinear mapping function that approximates the distribution of the sorted array.

9. The computer implemented method of any one of claims 4 to 7, wherein generating a mapping function based on the sorted array comprises implementing a neural network that approximates the distribution of the sorted array.

10. The computer implemented method of any one of claims 4 to 9, wherein mapping the audio segment metadata of the first audio to a dimensional space based on the mapping function comprises: determining one or more dimensions of analysis, wherein each dimension of analysis is based on an audio characteristic; and applying the mapping function to the audio segment metadata to distribute the audio segment metadata in each of the one or more dimensions of analysis.

11. The computer implemented method of any one of claims 4 to 10, wherein identifying the first audio segment in the dimensional space based on the second audio characteristic comprises:applying the mapping function to the second audio characteristic of the second audio segment to approximate the position of the second audio segment in the dimensional space; determining the dimensional distance between the second audio segment and one or more audio segments in the dimensional space; and identifying the first audio segment based on the dimensional distance, wherein the dimensional distance between the second audio segment and the first audio segment is lesser than the dimensional distance between the second audio segment and each other audio segment of the one or more audio segments in the dimensional space.

12. The computer implemented method of any one of claims 1 to 11, wherein determining the second audio characteristic of a second audio segment based on the analysis of the second audio segment comprises: retrieving a second audio file in real time from an audio input device; partitioning the second audio file into a second audio segment; and determining at least one of a peak value, root mean square value, spectral centroid, frequency domain magnitude shape, harmonic content, stochastic content and pitch, or magnitude of the second audio segment.

13. The computer implemented method of any one of claims 1 to 12, wherein modifying the second audio segment based on the first audio segment comprises: receiving a user input of an audio inertia value; modifying audio data of the second audio segment to incorporate the first audio characteristic of the first audio segment; and modifying the audio data of the second audio segment to include audio data of the first audio segment based on the audio inertia value.

14. The computer implemented method of any one of claims 1 to 13, wherein outputting the second audio segment comprises outputting the modified second audio segment via an audio output device.

15. The computer implemented method of any one of claims 1 to 14, wherein outputting the second audio segment comprises outputting the modified second audio segment in real time.

16. The computer implemented method of any one of claims 1 to 15, wherein the second audio segment is an audio segment partitioned from a second audio file, and outputting the secondaudio segment further comprises modifying and outputting one or more additional audio segments partitioned from the second audio file, such that the output comprises the entirety of the second audio file.

17. A system for generating audio style transfer, comprising: an audio file database; and a processor configured by instructions to perform operations comprising: partitioning a first audio file into a first audio segment, wherein the first audio segment comprises a portion of the first audio file; determining a first audio characteristic of the first audio segment based on an analysis of the first audio segment; mapping the first audio segment to a dimensional space based on the first audio characteristic; determining a second audio characteristic of a second audio segment based on an analysis of the second audio segment; identifying the first audio segment in the dimensional space based on the second audio characteristic; modifying the second audio segment based on the first audio segment; and outputting the second audio segment.

18. The system of claim 17, wherein partitioning the first audio file into the first audio segment comprises applying a window function to the audio file to partition the audio file.

19. The system of claim 17 or 18, wherein determining the first audio characteristic comprises at least one of a peak value, root mean square value, spectral centroid, frequency domain magnitude shape, harmonic content, stochastic content and pitch, or magnitude.

20. The system of any one of claims 17 to 19, wherein mapping the first audio segment to a dimensional space based on the first audio characteristic comprises: generating an array of audio segment metadata, wherein the array comprises audio segment metadata of the first audio segment; sorting the array based on the first audio characteristic; generating a mapping function based on the sorted array; andmapping the audio segment metadata of the first audio segment to the dimensional space based on the mapping function.

21. The system of claim 20, wherein the audio segment metadata of the first audio segment comprises the first audio characteristic of the first audio segment and a reference pointer to the first audio segment.

22. The system of claim 20 or 21, wherein sorting the array based on the first audio characteristic comprises sorting the array based on a value of the first audio characteristic.

23. The system of any one of claims 20 to 22, wherein sorting the array based on the first audio characteristic comprises further sorting the array based on an additional audio characteristic of the first audio segment.

24. The system of any one of claims 20 to 23, wherein generating a mapping function based on the sorted array comprises calculating a scaled nonlinear mapping function that approximates the distribution of the sorted array.

25. The system of any one of claims 20 to 23, wherein generating a mapping function based on the sorted array comprises implementing a neural network that approximates the distribution of the sorted array.

26. The system of any one of claims 20 to 25, wherein mapping the audio segment metadata of the first audio to a dimensional space based on the mapping function comprises: determining one or more dimensions of analysis, wherein each dimension of analysis is based on an audio characteristic; and applying the mapping function to the audio segment metadata to distribute the audio segment metadata in each of the one or more dimensions of analysis.

27. The system of any one of claims 20 to 26, wherein identifying the first audio segment in the dimensional space based on the second audio characteristic comprises: applying the mapping function to the second audio characteristic of the second audio segment to approximate the position of the second audio segment in the dimensional space; determining the dimensional distance between the second audio segment and one or more audio segments in the dimensional space; andidentifying the first audio segment based on the dimensional distance, wherein the dimensional distance between the second audio segment and the first audio segment is lesser than the dimensional distance between the second audio segment and each other audio segment of the one or more audio segments in the dimensional space.

28. The system of any one of claims 17 to 27, wherein determining the second audio characteristic of a second audio segment based on the analysis of the second audio segment comprises: retrieving a second audio file in real time from an audio input device; partitioning the second audio file into a second audio segment; and determining at least one of a peak value, root mean square value, spectral centroid, frequency domain magnitude shape, harmonic content, stochastic content and pitch, or magnitude of the second audio segment.

29. The system of any one of claims 17 to 28, wherein modifying the second audio segment based on the first audio segment comprises: receiving a user input of an audio inertia value; modifying audio data of the second audio segment to incorporate the first audio characteristic of the first audio segment; and modifying the audio data of the second audio segment to include audio data of the first audio segment based on the audio inertia value.

30. The system of any one of claims 17 to 29, wherein outputting the second audio segment comprises outputting the modified second audio segment via an audio output device.

31. The system of any one of claims 17 to 30, wherein outputting the second audio segment comprises outputting the modified second audio segment in real time.

32. The system of any one of claims 17 to 31, wherein the second audio segment is an audio segment partitioned from a second audio file, and outputting the second audio segment further comprises modifying and outputting one or more additional audio segments partitioned from the second audio file, such that the output comprises the entirety of the second audio file.

33. A method for creating a stylistic replica of an input audio file comprising: segmenting by a processing element a style file into a plurality of style grains;analyzing by the processing element the plurality of style grains to identify respective style acoustic characteristics for each of the plurality of style grains; mapping the style grains to a dimensional space based on the respective style acoustic characteristics; segmenting by the processing element an input file into a plurality of input grains; analyzing the plurality of input grains to identify respective input acoustic characteristics for each of the plurality of input grains; mapping by identification in the dimensional space a match between the respective style grains and the plurality of input grains based on the input acoustic characteristics and the style acoustic characteristics; and generating an style transfer file based on the mapping, wherein the style transfer file is representative of the input file and in the acoustic style of the style file.

34. The method of claim 33, further comprising storing the style transfer file in a memory component of a computer device associated with the processing element.

35. A system for creating a stylistic replica of an input audio file, comprising a processor configured by instructions to perform the method of any of claims 33 to 34.

36. A computer program storage medium readable by a computer system and encoding a computer program which, when executed by the computer system, causes performance of the method of any of claims 1 to 16 or 33 to 34.

Citation Information

Patent Citations

  • Method of generating an audio signal

    US10606548B2

  • System and method for creating timbres

    US20210256985A1

  • US202463624439P