Machine learning media editing
A machine learning model simplifies media editing by processing natural language commands to automatically select and adjust media tools, addressing the complexity of tool interactions and improving user accessibility and consistency.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-23
- Publication Date
- 2026-04-02
AI Technical Summary
Selecting an appropriate combination of media editing tools and adjusting their settings to achieve a desired outcome in media editing is technically challenging due to the complexity of interactions between tools, requiring a deep understanding of media editing processes.
A machine learning model processes natural language commands to automatically select and adjust media editing tools and settings, allowing users to input commands in everyday language, thereby simplifying the editing process and improving consistency and repeatability.
The solution enables users to make accurate and consistent media adjustments without deep technical knowledge, enhancing accessibility and reducing the learning curve for media editing.
Smart Images

Figure US2025047572_02042026_PF_FP_ABST
Abstract
Description
MACHINE LEARNING MEDIA EDITINGCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims the benefit of priority from U.S. Provisional Application No. 63 / 698,381, filed on 24 September 2024, and from European Patent Application No. 24205196.9, filed on 08 October 2024, which are both incorporated by reference herein in their entirety.TECHNICAL FIELD
[0001] The present disclosure relates to media processing and, more particularly, to processing techniques for editing media files using machine learning models.SUMMARY
[0002] Selecting an appropriate combination of media editing tools to achieve a desired outcome can be technically challenging, even when users have a clear vision of what they want to achieve. Media editing tools, whether for video, audio, images, or any other form of media, often come with a wide array of functions and settings. Each tool may have multiple parameters that need to be precisely adjusted to affect the media to achieve the desired outcome. Furthermore, in media editing, the effects and adjustments performed by each tool may interact with those of other tools in complex ways. Thus, a deep technical understanding of these interactions may be required for the user to be able to not only select the appropriate tool or combination of tools but also to select the appropriate settings and parameters for each tool (such as, for example, the magnitude and direction of the adjustment affected by each tool).
[0003] Systems, apparatuses, methods, and techniques described in this specification provide technical solutions to these challenges (among others) by receiving natural language commands for adjusting a media file. A machine learning model maps relevant semantic information from the input natural language commands to adjustments to the media file. The adjustments may include a selection of one or more media adjustment tools and / or settings or parameters for the one or more media adjustment tools, and the media may be automatically adjusted according to the outputs of the machine learning model. Accordingly, systems, apparatuses, methods, and techniques improve technological processes associated with media editing by automating the technically challenging process of selecting the appropriate combination of media editing tools and the appropriate settings and parameters for each tool.Such automated processes additionally remove human subjectivity from the media editing process, thereby improving consistency and increasing repeatability.
[0004] In some aspects, the techniques described herein relate to a computer- implemented method including: receiving, at a media processing application, an input via a user input element of a graphical user interface; providing the input to a machine learning model to generate an output indicative of adjustments to be performed to a media file; selecting, at the media processing application, one or more tools from a set of media adjustment tools based on the output from the machine learning model; and applying, at the media processing application, the selected tools to the media file to generate an adjusted media file.
[0005] In some aspects, the techniques described herein relate to a computer- implemented method, further including: generating, at the media processing application, an adjustable element at the graphical user interface; wherein the media processing application adjusts an intensity of adjustments applied by the selected tools in response to a user interacting with the adjustable element. In some aspects, the techniques described herein relate to a computer-implemented method, wherein: the output from the machine learning model indicates a plurality of adjustments, each adjustment being associated with a different tool from the set of tools; the adjustable element includes a plurality of adjustable elements, each adjustable element corresponding to a respective tool; and the media processing application adjusts an intensity of an adjustment performed by a tool corresponding to a selected one of the plurality of adjustable elements in response to a user interacting with the selected one of the plurality of adjustable elements.
[0006] In some aspects, the techniques described herein relate to a computer- implemented method, wherein the user input element includes a text input field. In some aspects, the techniques described herein relate to a computer-implemented method, wherein the user input element includes a voice input element. In some aspects, the techniques described herein relate to a computer-implemented method, wherein: the input includes text; and the machine learning model includes a text embedding model that receives the text and generates text embeddings based on the text. In some aspects, the techniques described herein relate to a computer-implemented method, wherein: the input includes audio; the machine learning model includes a speech-to-text application that receives the audio and generates a text representation based on the audio; and the machine learning model includes a textembedding model that receives the text representation and generates text embeddings based on the text representation.
[0007] In some aspects, the techniques described herein relate to a computer- implemented method, wherein: the machine learning model includes a media embedding model that receives the media file and generates at least one of media embeddings and metadata based on the media file. In some aspects, the techniques described herein relate to a computer-implemented method, wherein: the machine learning model includes a classifier model that receives the text embeddings and maps the text embeddings to the output indicative of adjustments to be performed to the media file. In some aspects, the techniques described herein relate to a computer-implemented method, wherein: the machine learning model includes a classifier model that receives the text embeddings and the media embeddings and maps the text embeddings and the media embeddings to the output indicative of adjustments to be performed to the media file.
[0008] In some aspects, the techniques described herein relate to a computer- implemented method, further including: providing, at the media processing application, the adjusted media file as an output via the graphical user interface. In some aspects, the techniques described herein relate to a computer-implemented method, wherein the media file includes at least one of an audio file and a video file. In some aspects, the techniques described herein relate to a computer-implemented method, wherein the set of media adjustment tools includes at least one of a noise reduction tool, a de-essing tool, a stereo widening tool, a treble adjustment tool, a mids adjustment tool, and a boost tool. In some aspects, the techniques described herein relate to a non-transitory computer-readable medium including executable instructions, which when executed by an electronic processor causes the electronic processor to perform the method. In some aspects, the techniques described herein relate to a system including: memory hardware storing instructions; and processor hardware configured to execute the instructions, wherein executing the instructions causes the system to perform the method.
[0009] Other examples, embodiments, features, and aspects will become apparent by consideration of the detailed description and accompanying drawings.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] FIG. 1 is a block diagram illustrating an example computing system for processing a media file according to natural language commands.
[0011] FIG. 2 is a block diagram illustrating example data flow between a speech-to- text application and various components of a machine learning model.
[0012] FIG. 3 is a block diagram showing an example architecture of a classifier model.
[0013] FIG. 4 is a block diagram illustrating an example process for automatically adjusting a media file.
[0014] FIG. 5 is a diagram illustrating aspects of an example graphical user interface generated by a media processing application for automatically adjusting a media file.
[0015] FIG. 6 is a diagram illustrating aspects of an example graphical user interface generated by a media processing application for automatically adjusting a media file.
[0016] FIG. 7 is a diagram illustrating aspects of an example graphical user interface generated by a media processing application for automatically adjusting a media file.
[0017] FIG. 8 is a diagram illustrating aspects of an example graphical user interface generated by a media processing application for automatically adjusting a media file.
[0018] FIG. 9 is a diagram illustrating aspects of an example graphical user interface generated by a media processing application for automatically adjusting a media file.
[0019] FIG. 10 is a flowchart illustrating an example process for editing media files.
[0020] FIG. 11 is a flowchart illustrating an example process for generating a training dataset.
[0021] FIG. 12 is a schematic illustration of an example training data sample.
[0022] FIG. 13 is a diagram illustrating aspects of an example graphical user interface for collecting user descriptions.
[0023] FIG. 14 is a flowchart illustrating an example process for training a machine learning model.
[0024] In the drawings, reference numbers may be reused to identify similar and / or identical elements.DETAILED DESCRIPTION
[0025] FIG. 1 is a block diagram illustrating an example computing system 100 for processing a media file according to natural language commands. As illustrated in FIG. 1, some examples of the system 100 include a media processing and transformation platform 102, one or more user devices 104 (such as user device 104-1 and user device 104- 2), and / or a communications system 106. Although two user devices 104 are shown in the example of FIG. 1, the system 100 may include any number of user devices 104. In some examples, one or more of the user devices 104 and / or the communications system 106 are omitted from the system 100. In various implementations, the media processing and transformation platform 102 communicate with the user devices 104 via the communications system 106.
[0026] In some examples, the media processing and transformation platform 102 is a computing platform deployed across a single server or one or more servers. The media processing and transformation platform 102 may receive a media file and / or natural language commands from one or more of the user devices 104, map the natural language commands to a selection of one or more media adjustment tools and / or parameters for each selected media adjustment tool, and / or automatically adjust the media file using the selected media adjustment tools and / or parameters. In various implementations, the user devices 104 include one or more computing platforms, such as smartphones, tablet computers, laptop computers, desktop computers, computer servers, etc.
[0027] In some examples, the communications system 106 includes one or more networks, such as a General Packet Radio Service (GPRS) network, a Time-Division Multiple Access (TDMA) network, a Code-Division Multiple Access (CDMA) network, a Global System of Mobile Communications (GSM) network, an Enhanced Data Rates for GSM Evolution (EDGE) network, a High-Speed Packet Access (HSPA) network, an Evolved High- Speed Packet Access (HSPA+) network, a Long Term Evolution (LTE) network, a Worldwide Interoperability for Microwave Access (WiMAX) network, a 5th-generation mobile network (5G), an Internet Protocol (IP) network, a Wireless Application Protocol (WAP) network, or an IEEE 802.11 standards network, as well as any suitable combination of the above networks. In various implementations, the communications system 106 includes an optical network, a local area network, and / or a global communication network, such as the Internet.
[0028] As illustrated in the example of FIG. 1 , the media processing and transformation platform 102 may include system resources 108, a communications interface 110, and / or non-transitory computer-readable storage media, such as, for example, storage 112. The non- transitory computer-readable storage media may contain instructions that, when executed, cause one or more electronic processors — such as electronic processors of the system resources 108 — to perform various functions described herein. The system resources 108 may include one or more electronic processors, one or more graphics processing units, volatile computer memory, non-volatile computer memory, and / or one or more system buses interconnecting various components of the media processing and transformation platform 102. The communications interface 110 may include hardware and / or software components that communicate with other devices, platforms, and / or systems over the communications system 106. In various implementations, the communications interface 110 includes one or more transceivers for sending and / or receiving data over the communications system 106.
[0029] The storage 112 may include a speech-to-text application 114, a machine learning model 116, a media processing application 118, and / or a machine learning training application 128. FIG. 2 is a block diagram 200 illustrating example data flow between the speech-to-text application 114 and various components of the machine learning model 116. Referring collectively to FIGS. 1 and 2, the speech-to-text application 114 may receive audio inputs 202 — such as audio inputs captured from a microphone or extracted from an audio file — and output a text representation 204 of the audio inputs 202. In various implementations, the speech-to-text application 114 preprocesses the captured audio to remove noise, normalize volume, enhance clarity, etc.
[0030] The speech-to-text application 114 may transform the preprocessed audio into a compact representation that captures relevant information from the preprocessed audio, transforming the preprocessed audio into a representation suitable for processing by downstream machine learning applications. In some examples, the speech-to-text application 114 extracts features such as Mel-frequency cepstral coefficients (MFCCs) and / or spectrograms from the preprocessed audio signals. The speech-to-text application 1 14 may provide the extracted MFCCs and / or spectrograms to one or more machine learning models, which may output text representing the speech content of the audio inputs. Examples of suitable machine learning models may include acoustic models, language models, and / or deep learning algorithms that match audio features with phonetic units and / or words. The speech- to-text application 114 may perform post-processing operations on the text output from the one or more machine learning models to correct for any errors and / or format the text for further processing by downstream machine learning models.
[0031] In various implementations, a user provides natural language commands as an audio input (such as audio input 202), and the speech-to-text application 114 converts the audio input into the text representation 204 and provides the text representation 204 as inputs to the machine learning model 116. In some examples, the user provides natural language commands as text (for example, by providing the text representation 204 to the machine learning model 116). The machine learning model 1 16 receives the text representation 204 as inputs, interprets, and understands meanings present in the text, and outputs adjustments to a media file based on the meanings. In various implementations, the machine learning model 116 includes a text embedding model 120, a media embedding model 122, and / or a classifier model 124.
[0032] The text embedding model 120 receives the text representation 204 from the speech-to-text application 114 or directly from the user and processes the text representation 204 to generate text embeddings 206 that capture relevant semantic information present in the text representation 204. In some examples, the text embeddings 206 may be a dense fixed-dimensional vector representation of the text representation 204 that captures semantic meanings of the text representation 204. In various implementations, the text embedding model 120 tokenizes the input text representation 204, splitting the input text into smaller units, such as words and / or sub-words. In some examples, each token is converted into a vector representation, for example, using a lookup table from pre-trained embeddings. In various implementations, each token is provided to a machine learning model, which may, for example, generate context-aware embeddings based on the context of the words (for example, represented by tokens) surrounding each word (for example, represented by tokens). In some examples, the token embeddings are aggregated to form a single fixed-dimensional vector representing the entire input text. Examples of suitable architectures for the machine learning model include neural networks, shallow neural networks, bidirectional long shortterm memory (LSTM) networks, transformer-based models, etc.
[0033] The media embedding model 122 receives the media file 208 to be modified as an input and processes the media file to generate media embeddings 210 that capture relevant features and characteristics of the media file 208 for use by downstream machine learning models. In various implementations, the media file 208 may be any form of media file, such as an audio file or a video file. In the example where the media file 208 is an audio file, the media embedding model 122 may preprocess raw audio data from the audio file for feature extraction. Suitable audio preprocessing techniques include normalization, noise reduction, silence trimming, converting the audio signal to a mono signal having a specific sample rate,etc. The media embedding model 122 may extract features from the preprocessed audio data.For example, the media embedding model 122 extracts MFCCs that capture the power spectrum of the audio data, spectrograms that capture visual representations of frequencies in the audio data as a function of time, chroma features that represent the energy of each pitch class in the audio data, etc. The extracted features may be provided to a machine learning model, which generates embeddings based on the extracted features. Examples of suitable machine learning models include convolutional neural networks (CNNs), recurrent neural networks (RNNs), transformer-based models, etc. The outputs from the machine learning model may be converted to a fixed-size embedding vector (for example, using a pooling technique) and output from the media embedding model 122 as media embeddings 210.
[0034] In the example where the media file 208 is a video file, the media embedding model 122 may extract individual frames or sequences of frames from the video file (for example, by sampling frames at regular intervals or using key frame extraction techniques) to reduce redundancy in the video data. The media embedding model 122 may extract meaningful features from individual frames and / or sequences of frames. For example, each frame and / or sequence of frames may be provided to a machine learning model such as a CNN. The machine learning model outputs high-dimensional feature vectors representing each frame and / or sequence of frames. The media embedding model 122 may capture temporal dynamics and / or relationships between the feature vectors representing the frames and / or sequences of frames. For example, the feature vectors may be provided to a machine learning model (such as a CNN, RNN, transformer-based model, temporal segment network (TSN), SlowFast network, etc.), which outputs feature representations. The media embedding model 122 may apply a pooling technique (such as global average pooling) on the feature representations to produce media embeddings 210 by outputting a single fixed-dimensional vector representing the entirety of the video data.
[0035] In addition to generating media embeddings 210 that capture the relevant features and characteristics of the media file 208, the media embedding model 122 can also include metadata with its output. This metadata may provide additional context and information about the media file 208 and / or the media embeddings 210. For example, in the case of an audio file, the metadata may include details such as the duration of the audio, the sampling rate, the bit rate, the original file format, etc. In various implementations, the metadata may also include information about any processing or preprocessing steps performed, such as whether noise reduction or normalization was performed (and any parameters used in these processes).
[0036] In the case of a video file, the metadata may include the resolution of the video, the frame rate, the encoding format, the length of the video, etc. In various implementations, the metadata may include information about the key frames extracted, the intervals at which frames were sampled, the algorithms used for frame selection, etc. In various implementations, the metadata can include information about specific features extracted and the methods used. For example, the metadata may include details about the types of audio features extracted (e.g., MFCCs, spectrograms, chroma features, etc.) and / or the types of video features extracted (e.g., spatial features from individual frames, temporal features from frame sequences, etc.). The metadata may be used by downstream machine learning models (or other processes) that use the media embeddings 210 to gain a more comprehensive understanding of the media file 208 and the embedding process.
[0037] The classifier model 124 may receive text embeddings 206 from the text embedding model 120 and output adjustments 212 for the media file based on the text embeddings 206. In various implementations, the classifier model 124 receives text embeddings 206 from the text embedding model 120 as well as media embeddings 210 from the media embedding model 122 and outputs adjustments 212 for the media file based on the text embeddings 206 and the media embeddings 210. In some examples, the adjustments 212 include a selection of media adjustment tools (such as a selection of tools from the media adjustment tools 126 of the media processing application 118) and / or specific settings or parameters for each selected tool to achieve the desired media adjustments as indicated by the user through the natural language commands.
[0038] The classifier model 124 may receive the text embeddings 206 as inputs and process the text embeddings 206 to understand the user’s command. In various implementations, the classifier model 124 also receives the media embeddings 210 and / or metadata generated by the media embedding model 122 as inputs to understand the current state of the media file. Based on this information, the classifier model 124 identifies which tools from the media adjustment tools 126 should be used and what specific settings and parameters should be selected for each identified tool to achieve the desired media modifications. In various implementations, the classifier model 124 includes one or more machine learning models. Examples of machine learning models include multilayer perceptrons (MLPs), CNNs, RNNs, LSTM networks, and / or transformer models.
[0039] In various implementations, an MLP is used to map the text embeddings 206 and / or the media embeddings 210 to the adjustments 212. The MLP may include an inputlayer, one or more hidden layers (such as multiple dense layers) following the input layer, and an output layer (for example, with a head for each tool indicating the adjustments — if any — to be performed by each tool or with separate heads for tool selections and adjustments to be performed using each tool) following the one or more hidden layers. In some examples, a CNN is used to capture local patterns in the text embeddings 206 and / or the media embeddings 210 and map the embeddings to the adjustments 212. The CNN may include an input layer, one or more convolutional layers (for example, one or more layers with convolutional filters and / or pooling layers) following the input layer, and an output layer (for example, with a head for each tool indicating the adjustments — if any — to be performed by each tool or with separate heads for tool selections and adjustments to be performed using each tool) following the one or more convolutional layers.
[0040] In various implementations, an RNN is used to capture temporal dependencies in the text embeddings 206 and / or the media embeddings 210 and map the embeddings to the adjustments 212. Capturing temporal dependencies can be beneficial when the text embeddings 206 and / or the media embeddings 210 have sequential patterns. The RNN may include an input layer, one or more RNN layers following the input layer that capture temporal dependencies, one or more fully connected layers following the RNN layers to make predictions, and an output layer following the one or more fully connected layers (for example, with a head for each tool indicating the adjustments — if any — to be performed by each tool or with separate heads for tool selections and adjustments to be performed using each tool).
[0041] In some examples, an LSTM network is used to capture temporal dependencies in the text embeddings 206 and / or the media embeddings 210 and map the embeddings to the adjustments 212. Capturing temporal dependencies can be beneficial when the text embeddings 206 and / or the media embeddings 210 have sequential patterns. The LSTM network may include an input layer, one or more LSTM layers following the input layer that capture temporal dependencies, one or more fully connected layers following the LSTM layers to make predictions, and an output layer following the one or more fully connected layers (for example, with a head for each tool indicating the adjustments — if any — to be performed by each tool or with separate heads for tool selections and adjustments to be performed using each tool).
[0042] In various implementations, a transformer model is used to capture long-range dependencies and relationships in the text embeddings 206 and / or the media embeddings 210and map the embeddings to the adjustments 212. The transformer model may include an input layer, one or more transformer encoder layers (such as layers including self- attention and / or feed- forward networks) following the input layer, one or more fully connected layers following the transformer encoder layers to make predictions, and an output layer following the transformer encoder layers (for example, with a head for each tool indicating the adjustments — if any — to be performed by each tool or with separate heads for tool selections and adjustments to be performed using each tool).
[0043] As will be described in detail, the machine learning training application 128 may generate training data, generate graphical interfaces for labeling the training data, and train the machine learning model 116 based on the labeled training data.
[0044] FIG. 3 is a block diagram showing an example architecture of a classifier model 124. As illustrated in the example of FIG. 3, some implementations of the classifier model 124 include a multi -output classifier 302, a signal classifier 304, a zero-shot classifier 306, and one or more binary classifiers 308 (such as binary classifier 308-1 and binary classifier 308-2). The multi-output classifier 302 may be a multi-classification model, such as a decision tree or a neural network. The signal classifier 304 may be a multiclassification model, such as a decision tree or a neural network. The zero-shot classifier 306 may be a machine learning model such as a language model trained for making audio adjustments. Each binary classifier 308 may be a binary classification model, such as a logistic regression model, a neural network, etc. Although two binary classifiers 308 are shown in the example of FIG. 3, the classifier model 124 may include a binary classifier 308 corresponding to each tool of the media adjustment tools 126.
[0045] The multi-output classifier 302 receives text embeddings 206 and / or media embeddings 210, analyzes the text embeddings 206 and / or the media embeddings 210, and generates an output 310 — such as a vector — indicating high-level actions to be performed on the media file. In various implementations, the output 310 includes one or more components. For example, the output 310 may include a confidence component 312, a signal adjustment component 314, a degree of adjustment component 316, a revert component 318, and a no action component 320. The confidence component 312 may indicate a confidence of the multi-output classifier 302 in understanding the commands present in the text embeddings 206. In various implementations, the classifier model 124 may generate adjustments only in response to the confidence component 312 indicating a confidence value above a threshold. For example, the threshold may be in a range of between about 50% andabout 100% (such as about 50%, about 55%, about 60%, about 65%, about 70%, about 75%, about 80%, about 85%, about 90%, about 95%, or about 100%).
[0046] The signal adjustment component 314 may indicate whether an adjustment to the media file is necessary and / or how the media file should be adjusted. The degree of adjustment component 316 may indicate an extent of the adjustments to the media file. The reversion component 318 may indicate whether the media file should be reverted to a previous state. The no action component 320 may indicate whether the user is satisfied with the adjustments and that no further adjustments are necessary. In various implementations, the multi-output classifier 302 provides the output 310 to the signal classifier 304. In some examples, the multi-output classifier 302 provides the signal adjustment component 314 to the signal classifier 304. The signal classifier 304 analyzes the output 310 and / or the signal adjustment component 314 and generates an output 322 — such as a vector — including information related to which tools from the media adjustment tools 126 should be selected. For example, the output 322 may include a component 324 (such as component 324-1 and component 324-2) associated with each tool. Although two components 324 are illustrated in the example of FIG. 3, the output 322 may include any number of components 324 corresponding to any number of tools.
[0047] Each component 324 may be provided as an input to a corresponding binary classifier 308, and the binary classifier 308 outputs an adjustment 326 for the corresponding tool. In various implementations, a binary classifier 308 determines whether a corresponding tool is to be turned on or off and outputs an adjustment 326 indicating whether the corresponding tool is to be turned on or off. In some examples, a binary classifier 308 determines an adjustment level for a corresponding tool and outputs an adjustment 326 indicating the adjustment level for the corresponding too. In various implementations, the media adjustment tools 126 include a noise reduction tool, and the binary classifier 308 associated with the noise reduction tool outputs an adjustment 326 that indicates whether to turn the noise reduction tool on or off. In some examples, the media adjustment tools 126 include a de-essing tool, and the binary classifier 308 associated with the de-essing tool outputs an adjustment 326 that indicates whether to turn the de-essing tool on or off. In various implementations, the media adjustment tools 126 include a stereo widening tool, and the binary classifier 308 associated with the stereo widening tool outputs an adjustment 326 that indicates whether to turn the stereo widening tool on or off.
[0048] In some examples, the media adjustment tools 126 include a treble adjustment tool, and the binary classifier 308 associated with the treble adjustment tool outputs an adjustment 326 indicating whether to add treble or remove treble (for example, by a fixed decibel amount) using the treble adjustment tool. In various implementations, the media adjustment tools 126 include a mids adjustment tool, and the binary classifier 308 associated with the mids adjustment tool outputs an adjustment 326 indicating whether to add treble or remove mids (for example, by a fixed decibel amount) using the mids adjustment tool. In some examples, the media adjustment tools 126 include a bass adjustment tool, and the binary classifier 308 associated with the bass adjustment tool outputs an adjustment 326 indicating whether to add treble or remove bass (for example, by a fixed decibel amount) using the bass adjustment tool. In various implementations, the media adjustment tools 126 include a boost tool, and the binary classifier 308 associated with the boost tool outputs an adjustment 326 indicating whether to increase or decrease the overall volume of the media file (for example, by a fixed decibel amount) using the boost tool.
[0049] In some examples, the adjustments 326 are output from the classifier model 124 as adjustments 212. In various implementations, text embeddings 206 indicate that the user has indicated a degree of adjustment for the media file (for example, relative to a previous adjustment). The multi-output classifier 302 may provide the outputs 310 (for example, the degree of adjustment component 316) to the zero-shot classifier 306 and the zero-shot classifier 306 generates adjustment refinements 328 based on stored recent adjustments 330. The classifier model 124 may apply the adjusted refinements 328 to any relevant adjustments 326 to generate updated adjustments 332 and outputs any updated adjustments 332 along with any adjustments 326 not affected by the adjustment refinements 328 as adjustments 212. In some examples, the text embeddings 206 indicate that the user wishes to revert the media file to a previous state. The classifier model 124 may generate adjustment refinements 328 that revert the media file to a corresponding previous state stored in the recent adjustments 330, and the classifier model 124 outputs any updated adjustments 332 along with any adjustments 326 not affected by the adjustment refinements 328 as adjustments 212.
[0050] Returning to FIG. 1, the media processing application 118 may generate a graphical user interface for the user to upload media files and provide natural language inputs indicating how the user would like to adjust the uploaded media file. The media processing application 118 may automatically provide the user’s natural language inputs and / or the uploaded media file to the machine learning model 116 and adjust the uploaded media filebased on the adjustments output by the machine learning model 116. FIG. 4 is a block diagram 400 illustrating an example process for automatically adjusting a media file. FIGS. 5-9 are diagrams illustrating aspects of an example graphical user interface 500 generated by the media processing application 118 for automatically adjusting a media file. Referring collectively to FIGS. 4-7, in various implementations, a user device 104 accesses the graphical user interface 500 via the communications system 106. As illustrated in FIG. 5, the graphical user interface 500 may include a media input field 502. The user may access the media input field 502 and upload a media file (such as the previously referenced media file 208) to the media input field 502. The graphical user interface 500 may also include a multimodal natural language input field 504 for the user to input natural language commands.
[0051] The multimodal natural language input field 504 may include a text input field 506, a voice input button 508, and / or a submit button 510. The user may input natural language commands describing how the user would like to adjust the media file 208 as text into the text input field 506 or as a voice command by selecting the voice input button 508. After inputting the natural language commands, the user may select the submit button 510, and the media processing application 118 provides the text input to the input text field 506 (for example, the text representation 204) or the voice command (for example, the audio inputs 202) input by selecting the voice input button 508 as input commands 402 to the machine learning model 116. The machine learning model 116 analyzes the input commands 402 (for example, the text representation 204 or the audio inputs 202) and / or the media file 208 and outputs adjustments 212 (for example, accordingly to the previously described techniques).
[0052] In various implementations, the media processing application 118 automatically adjusts the media file 208 based on the adjustments 212 output from the machine learning model 116. For example, the media processing application 118 selects tools from the media adjustment tools 126 (and settings and parameters for each selected tool as appropriate) based on the adjustments 212, applies adjustments to the media file 208 using the selected tools and parameters, and generates an adjusted media file 404.
[0053] The media processing application 1 18 may generate a playback field 512 on the graphical user interface 500 and populate the playback field 512 with the original media file 208 and / or the adjusted media file 404. The user may select the original media file 208 and / or the adjusted media file 404 by selecting a corresponding element in the playback field 512 and play back and / or download the respective media file. In variousimplementations, the media processing application 118 generates a graphical user interface element allowing the user to adjust an intensity of the overall adjustment to the media file 208. For example, as illustrated in FIG. 6, the media processing application 118 may generate an adjustment slider 514. The user may provide updated input commands 406 using the adjustment slider 514. For example, in response to the user increasing the intensity of the adjustments via the adjustment slider 514, the media processing application 118 may increase the intensity of the adjustments to the media file 208 using the selected tools and generate an updated adjusted media file 404. In response to the user decreasing the intensity of the adjustments via the adjustment slider 514, the media processing application 118 may decrease the intensity of the adjustments to the media file 208 using the selected tools and generate an updated adjusted media file 532.
[0054] In some examples, the media processing application 118 generates a graphical user interface element corresponding to each tool used to perform the adjustment that allows the user to adjust an intensity of the adjustment performed by the tool. For example, as illustrated in FIG. 7, the media processing application 118 may generate adjustment slider 516 corresponding to a first selected tool, adjustment slider 518 corresponding to a second selected tool, and adjustment slider 520 corresponding to a third selected tool. Although three adjustment sliders 516-518 are shown in the example of FIG. 7, the media processing application 118 may generate any number of adjustment sliders (for example, corresponding to any number of selected tools as indicated by the adjustments 212). The user may provide updated input commands 406 using the adjustment sliders 516-518. The user may increase or decrease the intensity of the adjustment performed by each selected tool by moving the corresponding adjustment slider. The media processing application 118 may increase and / or decrease the adjustment performed by each tool (as indicated by the adjustment sliders 516— 520) and generate an updated adjusted media file 532.
[0055] In various implementations, the media processing application 1 18 generates a multimodal natural language input field 522 to allow the user to perform additional adjustments on the original media file 208 and / or the adjusted media file 404. The multimodal natural language input field 522 may include a text input field 524, a voice input button 526, and / or a submit button 530. The user may input natural language commands describing how the user would like to adjust the original media file 208 and / or the adjusted media file 404 as text into the input text field 524 or as a voice command by selecting the voice input button 526. For example, the user may indicate that they would like to perform further adjustments, revert a media file to a previous state, etc. After inputting the natural languagecommands, the user may select the submit button 530, and the media processing application 118 again provides the text input to the input text field 524 (for example, the text representation 204) or the voice command (for example, the audio inputs 202) input by selecting the voice input button 526 as updated or new input commands 406 to the machine learning model 116.
[0056] The machine learning model 116 analyzes the updated or new input commands 402 (for example, the updated or new text representation 204 or the updated or new audio inputs 202) and / or the media file 208 and outputs updated or new adjustments 212 (for example, according to the previously described techniques). The media processing application 118 may generate an updated adjusted media file 532 by applying the updated adjustments 212 output from the machine learning model 116. For example, the media processing application 118 selects tools from the media adjustment tools 126 (and settings and parameters for each selected tool as appropriate) based on the updated adjustments 212, applies adjustments to the media file 208 using the selected tools and parameters, and generates an updated adjusted media file 532. As illustrated in the examples of FIGS. 8 and 9, the media processing application 118 may populate the playback field 512 with the updated adjusted media file 532. The user may select the updated adjusted media file 532 by selecting a corresponding element in the playback field 512 and play back and / or download the updated adjusted media file 532.EXAMPLES
[0057] Systems, apparatuses, methods, and techniques implemented according to this specification provide a variety of technical benefits over conventional solutions. For example, in various implementations, the machine learning model 116 includes a large language model (LLM). The language structure encoded in the LEM may allow the machine learning model 116 to recognize seemingly different commands that require the application of the same tool. For example, the user may request to adjust a media file 208 so that it sounds “less echoey,” “closer to the microphone,” “less like in a cave,” “less wet,” “drier,” or “have less reverb.” These commands may appear to be different, but all may require application of a reverb-suppression tool (for example, from the media adjustment tools 126).
[0058] In an example where the user requests the media processing application 118 to adjust the media file 208 to “make it sound like I am in a cave,” the machine learning model 116 may map the user’ s command to an application of specific tools from the media adjustment tools 126. For example, the media adjustment tools 126 may include areverberator tool, a decorrelator tool that converts a mono audio signal to a wide stereo signal, a loudness maximizer tool, and a static equalizer tool. The machine learning model 1 16 may map the user’s command to adjustments 212 that include a selection of the reverberator tool and the decorrelator tool. The media processing application 118 may apply the reverberator tool and the decorrelator tool to the media file 208 based on the adjustments 212 and produce an adjusted media file 404 with a cave-like audio effect applied to the entire media file 208.
[0059] In another example, the user may request adjustments to the media file 208 that require knowledge of contents of the media file 208 itself. For example, the audio component of the media file 208 may include narration over background music. The media adjustment tools 126 may include a reverberator tool, a decorrelator tool that converts a mono audio signal to a wide stereo signal, a loudness maximizer tool, a static equalizer tool, and a blind source separation tool that separates speech from audio. The user may request the media processing application 118 to adjust the media file 208 to “make the voice sound like 1 am in a cave.” The machine learning model 116 may map the user’s command to adjustments 212 that include applying the blind source separation tool to select the voice object present in the media file 208, applying the reverberator tool to the selected voice object, and applying the decorrelator tool to the selected voice object. The media processing application 118 may apply the adjustments 212 to apply a cave-like audio effect to only the voice component of the media file 208 to produce the adjusted media file 404.
[0060] Thus, techniques described in herein may enable even users without a deep technical understanding of the effects and adjustments performed by each tool or the complex interactions between the tools to consistently and accurately make adjustments to the media file 208 using natural language commands.
[0061] Furthermore, implementations where the machine learning model 116 analyzes the media file 208 itself in addition to the input command 402 may achieve additional technical benefits over conventional solutions. For example, analyzing the media file 208 may allow the machine learning model 116 to identify markers of temporal locations in the media file 208 (for example, time stamps, cue points indicative of scene changes, scene cutes, etc.), spatial properties of the media file 208, objects present in the media file 208 (for example, using blind source separation processes — such as processes that extract voices from remaining audio, user guided processes — such as processes where users mark areas on a screen, etc.), and / or attributes propagated over time (for example, identifying color pallets of different media files 208, which can allow the machine learning model 116 to output adjustments 212that match the color palette of one media file 208 to the color palette of another media file 208).
[0062] Allowing users to use natural language commands to adjust media files frees users from having to learn and navigate complex graphical user interfaces including large sets of controls. However, in certain scenarios, adjusting media files using a graphical user interface element such as a slider or a knob may be more beneficial than using iterative voice commands. For example, sliders or knobs may allow for fine-tuned adjustments in a continuous range, allowing users to adjust parameters with high precision (whereas voice commands may be imprecise, and it can be challenging to communicate the exact degree of change through voice alone — particularly where the required adjustments are subtle). Thus, as previously described the machine learning model 116 may derive the dimensions and / or attributes on which adjustment is desired. The media processing application 118 then generates one or more sliders or knobs (such as the adjustment sliders 514-520 illustrated in FIGS. 6-9) that allow the user to control the amount by which each attribute is to be changed.
[0063] For example, the media file 208 may include audio narration over background music. The media adjustment tools 126 may include a reverberator tool, a decorrelator tool that converts mono audio signals to stereo audio signals, a loudness maximizer tool, a static equalizer tool, a blind source separation tool that separates speech from remaining audio, and a gain control tool for each audio object present in a media file. The user may request the media processing application 118 to adjust the media file 208 to “make the voice stand out from the music.” The machine learning model 116 may map the user’s command to adjustments 212 that include applying the source selection tool to the media file 208 to select the voice object, applying the gain control tool for the voice object to the media file 208, and generating a command to the media processing application 118 to generate a slider or knob allowing the user to control the gain of the voice object. The media processing application 118 outputs the slider or knob to a graphical user interface, allowing the user to control the gain of the voice object.
[0064] In various implementations, the media processing application 118 logs interactions and requests from the users over time. The media processing application 118 may analyze these logged interactions and identify whether users are requesting adjustments that require tools not present in the media adjustment tools 126. Such analysis may be used as a guide to further development or addition of tools to the media adjustment tools 126. In some examples, the machine learning model 116 and / or the media processing application 118 may perform asentiment analysis of the user’ s commands and / or an emotion analysis of the user’ s voice for additional feedback.
[0065] FIG. 10 is a flowchart illustrating an example process 1000 for editing media files according to the techniques described herein. In the example process 1000, the media processing application 118 receives an input according to any of the previously described techniques (at block 1002). For example, the user may provide the input commands 402 via the input field 504 of the graphical user interface 500. As previously described, the input may be in a combination of modalities (such as, for example, text and / or audio) and may describe one or more desired adjustments to be performed to the media file 208.
[0066] The user may provide the input commands 402 in natural and / or non-technical language, and / or via audio or spoken commands, which may enhance the accessibility, efficiency, and accuracy of the media editing process. For example, allowing users to input commands in natural and / or non-technical language eliminates the need for users to understand complex technical jargon or have deep expertise in the media adjustment tools 126. This enhances accessibility for a broader range of users, including those without technical skills, to perform complex media adjustments.
[0067] Furthermore, traditional media editing tools may require users to learn and master various technical controls and / or interfaces, which can be time-consuming and intimidating. Allowing for natural and / or non-technical language inputs streamlines and reduces the learning curve associated with the media editing process by allowing users to describe desired adjustments in everyday language, reducing the time and effort needed to achieve the desired media edits.
[0068] In the example process 1000, the input (e.g., the input commands 402) is provided to the machine learning model 116, which may generate an output indicative of adjustments to be performed to the media file 208 according to any of the previously described techniques (at block 1004). For example, as previously described, the classifier model 124 may map text embeddings 206 and / or media embeddings 210 generated from the input to adjustments 212, which may include a selection of tools from the media adjustment tools 126 and / or specific settings or parameters for each selected tool.
[0069] Using the machine learning model 116 to map the input commands 402 to adjustments 212 improves the consistency and accuracy of the adjustments that may be performed on the media file 208. For example, the machine learning model 116 mayconsistently interpret and execute commands based on a consistent semantic understanding of the natural and / or non-technical language inputs, rather than subjective and inconsistent human selections of tools and / or settings. This reduces variability and human error, providing more consistent results across different media editing sessions.
[0070] Furthermore, the machine learning model 116 can map natural and / or nontechnical language commands to a precise selection of tools and / or adjustments that may otherwise require manual fine-tuning. This automation allows for complex interactions between multiple tools to be handled seamlessly, improving the overall efficiency of the editing process. Additionally, the ability of the machine learning model 1 16 to interpret a wide range of phrasings for similar adjustments (e.g., “reduce echo,” “make it less echoey,” “sound less like a cave,” etc.) provides users the flexibility to communicate media edits in ways that feel natural to them, without requiring users to conform to a rigid set of commands.
[0071] In the example process 1000, the media processing application 118 selects one or more tools and / or settings for the selected tools according to any of the previously described techniques (e.g., from the media adjustment tools 126, as previously described) (at block 1006). In the example process 1000, the media processing application 118 applies the selected tools and / or settings to the media file 208 to automatically generate an adjusted media file accordingly to any of the previously described techniques (for example, adjusted media file 404 and / or updated adjusted media file 532, as previously described) (at block 1008).
[0072] Automatically generating the adjusted media file may offer a variety of technical benefits allowing for expanded scalability of the media editing process. For example, automatic tool selection and application allows the system to process numerous media files quickly without manual intervention. This automation framework also allows for the parallel processing of multiple media files, significantly reducing the overall time required for large- scale projects. Furthermore, the previously described consistency provided by the machine learning model 116 ensures that media adjustments maintain consistent quality standards and / or uniformity across batches of media files.
[0073] Furthermore, as previously described, the original media file 208 and / or adjusted media files (such as, for example, the adjusted media file 404 and / or updated adjusted media file 532) may be previewed in the playback field 512 of the graphical user interface 500. Automatically generating the adjusted media fdes and allowing users to preview the adjusted media files directly on the graphical user interface 500 provides immediate feedback on howthe automated adjustments have affected the media. This immediate insight may help users to quickly identify whether the adjustments meet their expectations or if further fine-tuning is necessary. By previewing the adjusted media, users can spot any unwanted artifacts, errors, and / or inconsistencies that may be inadvertently introduced by the automated tools. If any further adjustments are needed to correct the adjusted media, the graphical user interface 500 allows the user to immediately input the corrections (e.g., via the input field 522). This ensures that the finalized adjusted media is error free and / or meets quality standards.
[0074] Providing previews to the user also helps increase the user’ s confidence in the automated media adjustment process. When users are able to see the effects of the adjustments in real-time, users may gain confidence in the system’s capabilities and be more likely to rely on it for future tasks. Additionally, previews enable an iterative process where users can experiment with different input commands, see the results instantly, and refine their inputs or provide adjustments via different inputs to achieve the desired creative effect.TRAINING
[0075] FIG. 11 is a flowchart illustrating an example process 1100 for generating a training dataset for training the machine learning model 116. In the example process 1100, the machine learning training application 128 generates sets of audio pairs (at block 1102). Each audio pair may include an original audio sample and an adjusted audio sample. The machine learning training application 128 may generate the adjusted audio sample by applying one or more adjustments to the original audio sample using one or more of the media adjustment tools 126. The original audio sample, adjusted audio sample, and labels of the adjustments may be associated with each other to form a training data sample. Each training sample may represent a particular adjustment, and a training dataset may include multiple training data samples representing a range of different adjustments.
[0076] FIG. 12 is a schematic illustration of an example training data sample 1202. As shown in the example of FIG. 12, the training data sample 1202 may include an original audio sample 1204, an adjusted audio sample 1206, and a label of the adjustments 1208. The original audio sample 1204 may serve as a baseline or reference point for comparisons. The adjusted audio sample 1206 may be a modified version of the original audio sample 1204 with specific adjustments applied (such as, for example, any of the adjustments previously described with reference to the media adjustment tools 126). The labels of adjustments 1208 may include a structured label that describes the specific adjustments applied to the original audio sample 1204 to create the adjusted audio sample 1206. In various implementations, thelabels of adjustments 1208 include types of adjustments 1210 and / or magnitudes of adjustments 1212. The types of adjustments 1210 indicate particular modifications (such as, for example, any of the adjustments previously described with reference to the media adjustment tools 126). The magnitudes of adjustments 1212 indicate the extent or intensity of the adjustments performed.
[0077] Returning to FIG. 11, in the example process 1100, the machine learning training application 128 collects user descriptions of the adjustments performed (at block 1104). In various implementations, the machine learning training application 128 generates a graphical user interface and collects user descriptions via the graphical user interface. FIG. 13 is a diagram illustrating aspects of an example graphical user interface 1300 generated by the machine learning training application 128 for collecting user descriptions. As illustrated in FIG. 13, the graphical user interface 1300 may include a first media playback button 1302 and a second media playback button 1304. In response to the user selecting the first media playback button 1302, the machine learning training application 128 may play the original audio sample 1204 to the user. In response to the user selecting the second media playback button 1304, the machine learning training application 128 may play the adjusted audio sample 1206 to the user.
[0078] The graphical user interface 1300 may further include a first input field 1306 and a second input field 1308. After listening to the original audio sample 1204 and the adjusted audio sample 1206, the user may input a description (such as, for example, natural language descriptions) of the modifications needed to transform the original audio sample 1204 into the adjusted audio sample 1206 via the first input field 1306 (such as, for example, “increase the bass” and / or “add some echo”). The user may input a description of the modifications needed to transform the adjusted audio sample 1206 back into the original audio sample 1204 via the second input field 1308. In various implementations, the user inputs text and / or audio (e.g., voice descriptions) into the first input field 1306 and / or the second input field 1308.
[0079] Referring to FIG. 11, in the example process 1100, the machine learning training application labels the audio pairs using the collected user descriptions (at block 1106). For example, referring to FIG. 12, the machine learning training application 128 adds the user descriptions input via the first input field 1306 and / or the second input field 1308 to the training data sample 1202 as user descriptions of adjustments 1214. In various implementations, the user descriptions of adjustments 1214 includes forward adjustment descriptions 1216 and reverse adjustment descriptions 1218. The machine learning trainingapplication 128 may save the descriptions input via the first input field 1306 as forward adjustment descriptions 1216 and / or the descriptions input via the second input field 1308 as reverse adjustment descriptions 1218.
[0080] In various implementations, the machine learning training application 128 generates a training dataset including multiple training data samples. In a particular, nonlimiting example, the machine learning training application 128 generates a training dataset representing seven unique types of adjustments. The audio pairs (each corresponding to a type of adjustment) may be provided to a set of 50 listeners (such as, for example, 24 amateur listeners and 26 expert listeners), each of whom provides user descriptions of the adjustments.
[0081] The training data samples 1202 of the training dataset may be used to train the machine learning model 116. FIG. 14 is a flowchart illustrating an example process 1400 for training the machine learning model 116. In the example process 1400, the machine learning training application 128 initializes the classifier model 124 (at block 1402). For example, the classifier model 124 may initialized as a pre-trained language model, such as a transformerbased language model. In various implementations, the classifier model 124 is initialized as a pre-trained Bidirectional Encoder Representations from Transformers (BERT) model, a Robustly Optimized BERT Approach (RoBERTa) model, a DistilBERT model, A Lite BERT (ALBERT) model, an XLNet model, a Generative Pre-trained Transformer (GPT) model, a Text-to-Text Transfer Transformer (T5) model, an Enhanced Representation through Knowledge Integration (ERNIE) model, etc.
[0082] In the example process 1400, the machine learning training application 128 loads the training dataset (at block 1404). The training dataset may be generated as previously described with reference to FIGS. 11-13. In the example process 1400, the machine learning training application 128 may generate input features for the classifier model 124 based on the training dataset (at block 1406). In various implementations, the machine learning training application 128 generates text embeddings 206 from the user descriptions of adjustments 1214 and / or media embeddings 210 from the original audio samples 1204 and / or adjusted audio samples 1206. In the example process 1400, the machine learning training application 128 provides the input features to the classifier model 124 to generate predicted adjustments 212.
[0083] In the example process 1400, the machine learning training application 128 computes a loss based on the generated predicted adjustments 212 and updates parameters of the classifier model 124 (at block 1408). For example, the machine learning training application 128 compares the predicted adjustments 212 output from the classifier model 124to the labels of adjustments 1208 from the training dataset using a loss function. The loss function measures the difference between the adjustments 212 output from the classifier model 124 and the labels of adjustments 1208, which represent the ground truth, to provide a value (such as a scalar value) representing the error of the classifier model 124. In various implementations, the loss function may be a cross entropy loss function. The parameters of the classifier model 124 may be optimized according to an optimization algorithm to minimize the loss computed by the loss function. In various implementations, the optimization algorithm may be the Adaptive Moment Estimation (Adam) optimizer.
[0084] In various implementations, the operations associated with blocks 1408 and 1410 may be repeated for a number of full passes through the entire training dataset. Each pass may be referred to as an epoch. In a particular non-limiting example, the operations associated with blocks 1408 and 1410 are repeated for 30 passes (e.g., the classifier model 124 is trained for 30 epochs). In some examples, the machine learning training application 128 uses the training dataset to fine-tune the multi -output classifier 302 and each binary classifier 308 with corresponding data. In various implementations, the machine learning training application 128 splits the training dataset into 80% for training, 10% for validation, and 10% for testing.ADDITIONAL ENUMERATED EXAMPLES
[0085] Example 1. A computer-implemented method comprising: receiving, at a media processing application, an input via a user input element of a graphical user interface; providing the input to a machine learning model to generate an output indicative of adjustments to be performed to a media file; selecting, at the media processing application, one or more tools from a set of media adjustment tools based on the output from the machine learning model; and applying, at the media processing application, the selected tools to the media file to generate an adjusted media file.
[0086] Example 2. The computer-implemented method of example 1, further comprising: generating, at the media processing application, an adjustable element at the graphical user interface; wherein the media processing application adjusts an intensity of adjustments applied by the selected tools in response to a user interacting with the adjustable element.
[0087] Example 3. The computer-implemented method of any one of examples 1 or 2, wherein: the output from the machine learning model indicates a plurality of adjustments, each adjustment being associated with a different tool from the set of tools; the adjustable element includes a plurality of adjustable elements, each adjustable element corresponding toa respective tool; and the media processing application adjusts an intensity of an adjustment performed by a tool corresponding to a selected one of the plurality of adjustable elements in response to a user interacting with the selected one of the plurality of adjustable elements.
[0088] Example 4. The computer-implemented method of any one of examples 1-3, wherein the user input element includes a text input field.
[0089] Example 5. The computer- implemented method of any one of examples 1-4, wherein the user input element includes a voice input element.
[0090] Example 6. The computer-implemented method of any one of examples 1-5, wherein: the input includes text; and the machine learning model includes a text embedding model that receives the text and generates text embeddings based on the text.
[0091] Example 7. The computer-implemented method of any one of examples 1-5, wherein: the input includes audio; the machine learning model includes a speech-to-text application that receives the audio and generates a text representation based on the audio; and the machine learning model includes a text embedding model that receives the text representation and generates text embeddings based on the text representation.
[0092] Example 8. The computer-implemented method of any one of examples 6 or 7, wherein: the machine learning model includes a media embedding model that receives the media file and generates at least one of media embeddings and metadata based on the media file.
[0093] Example 9. The computer-implemented method of any one of examples 6 or 7, wherein: the machine learning model includes a classifier model that receives the text embeddings and maps the text embeddings to the output indicative of adjustments to be performed to the media file.
[0094] Example 10. The computer-implemented method of example 8, wherein: the machine learning model includes a classifier model that receives the text embeddings and the media embeddings and maps the text embeddings and the media embeddings to the output indicative of adjustments to be performed to the media file.
[0095] Example 11. The computer-implemented method of any one of examples 1-10, further comprising: providing, at the media processing application, the adjusted media file as an output via the graphical user interface.
[0096] Example 12. The computer-implemented method of any one of examples 1-11, wherein the media file includes at least one of an audio file and a video file.
[0097] Example 13. The computer-implemented method of any one of examples 1-11, wherein the set of media adjustment tools includes at least one of a noise reduction tool, a de- essing tool, a stereo widening tool, a treble adjustment tool, a mids adjustment tool, a bass adjustment tool, and a boost adjustment tool.
[0098] Example 14. A non-transitory computer-readable medium comprising executable instructions, which when executed by an electronic processor causes the electronic processor to perform the method of any one of examples 1-13.
[0099] Example 15. A system comprising: memory hardware storing instructions; and processor hardware configured to execute the instructions, wherein executing the instructions causes the system to perform the method of any one of examples 1-13.CONCLUSION
[0100] The foregoing description is merely illustrative in nature and does not limit the scope of the disclosure or its applications. The broad teachings of the disclosure may be implemented in many different ways. While the disclosure includes some particular examples, other modifications will become apparent upon a study of the drawings, the text of this specification, and the following claims. In the written description and the claims, one or more processes within any given method may be executed in a different order — or processes may be executed concurrently or in combination with each other — without altering the principles of this disclosure. Similarly, instructions stored in a non-transitory computer-readable medium may be executed in a different order — or concurrently — without altering the principles of this disclosure. Unless otherwise indicated, the numbering or other labeling of instructions or method steps is done for convenient reference and does not necessarily indicate a fixed sequencing or ordering.
[0101] It should also be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components may be utilized in various implementations. Aspects, features, and instances may include hardware, software, and electronic components or modules that, for purposes of discussion, may be illustrated and described as if the majority of the components were implemented solely in hardware. However, one of ordinary skill in the art, and based on a reading of this detailed description, would recognize that, in at least one instance, the electronic based aspects of the inventionmay be implemented in software (for example, stored on non-transitory computer-readable medium) executable by one or more processors. As a consequence, it should be noted that a plurality of hardware and software-based devices, as well as a plurality of different structural components may be utilized to implement the invention. For example, “control units” and “controllers” described in the specification can include one or more electronic processors, one or more memories including a non-transitory computer-readable medium, one or more input / output interfaces, and various connections (for example, a system bus) connecting the components.
[0102] Unless the context of their usage unambiguously indicates otherwise, the articles “a,” “an,” and “the” should not be interpreted to mean “only one.” Rather, these articles should be interpreted to mean “at least one” or “one or more.” Likewise, when the terms “the” or “said” are used to refer to a noun previously introduced by the indefinite article “a” or “an,” the terms “the” or “said” should similarly be interpreted to mean “at least one” or “one or more” unless the context of their usage unambiguously indicates otherwise.
[0103] It should also be understood that although certain drawings illustrate hardware and software located within particular devices, these depictions are for illustrative purposes only. In some embodiments, the illustrated components may be combined or divided into separate software, firmware, and / or hardware. For example, instead of being located within and performed by a single electronic processor, logic and processing may be distributed among multiple electronic processors. Regardless of how they are combined or divided, hardware and software components may be located on the same computing device or may be distributed among different computing devices connected by one or more networks or other suitable connections or links.
[0104] Thus, in the claims, if an apparatus or system is claimed, for example, as including an electronic processor or other element configured in a certain manner, for example, to make multiple determinations, the claim or claim element should be interpreted as meaning one or more electronic processors (or other element) where any one of the one or more electronic processors (or other element) is configured as claimed, for example, to make some or all of the multiple determinations collectively. To reiterate, those electronic processors and processing may be distributed.
[0105] Spatial and functional relationships between elements — such as modules — are described using terms such as (but not limited to) “connected,” “engaged,” “interfaced,” and / or “coupled.” Unless explicitly described as being “direct,” relationships betweenelements may be direct or include intervening elements. The phrase “at least one of A, B, and C” should be construed to indicate a logical relationship (A OR B OR C), where OR is a nonexclusive logical OR, and should not be construed to mean “at least one of A, at least one of B, and at least one of C.” The term “set” does not necessarily exclude the empty set. For example, the term “set” may have zero elements. The term “subset” does not necessarily require a proper subset. For example, a “subset” of set A may be coextensive with set A, or include elements of set A. Furthermore, the term “subset” does not necessarily exclude the empty set.
[0106] In the figures, the directions of arrows generally demonstrate the flow of information — such as data or instructions. The direction of an arrow does not imply that information is not being transmitted in the reverse direction. For example, when information is sent from a first element to a second element, the arrow may point from the first element to the second element. However, the second element may send requests for data to the first element, and / or acknowledgements of receipt of information to the first element. Furthermore, while the figures illustrate a number of components and / or steps, any one or more of the components and / or steps may be omitted or duplicated, as suitable for the application and setting.
[0107] The term computer-readable medium does not encompass transitory electrical or electromagnetic signals or electromagnetic signals propagating through a medium — such as on an electromagnetic carrier wave. The term “computer-readable medium” is considered tangible and non-transitory. The functional blocks, flowchart elements, and message sequence charts described above serve as software specifications that can be translated into computer programs by the routine work of a skilled technician or programmer.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented method comprising: receiving, at a media processing application, an input via a user input element of a graphical user interface; providing the input to a machine learning model to generate an output indicative of adjustments to be performed to a media file; selecting, at the media processing application, one or more tools from a set of media adjustment tools based on the output from the machine learning model; and applying, at the media processing application, the selected tools to the media file to generate an adjusted media file.
2. The computer-implemented method of claim 1 , further comprising: generating, at the media processing application, an adjustable element at the graphical user interface; wherein the media processing application adjusts an intensity of adjustments applied by the selected tools in response to a user interacting with the adjustable element.
3. The computer-implemented method of any one of claims 1 or 2, wherein: the output from the machine learning model indicates a plurality of adjustments, each adjustment being associated with a different tool from the set of tools; the adjustable element includes a plurality of adjustable elements, each adjustable element corresponding to a respective tool; and the media processing application adjusts an intensity of an adjustment performed by a tool corresponding to a selected one of the plurality of adjustable elements in response to a user interacting with the selected one of the plurality of adjustable elements.
4. The computer-implemented method of any one of claims 1-3, wherein the user input element includes a text input field.
5. The computer-implemented method of any one of claims 1-4, wherein the user input element includes a voice input element.
6. The computer-implemented method of any one of claims 1-5, wherein: the input includes text; andthe machine learning model includes a text embedding model that receives the text and generates text embeddings based on the text.
7. The computer-implemented method of any one of claims 1-5, wherein: the input includes audio; the machine learning model includes a speech-to-text application that receives the audio and generates a text representation based on the audio; and the machine learning model includes a text embedding model that receives the text representation and generates text embeddings based on the text representation.
8. The computer-implemented method of any one of claims 6 or 7, wherein: the machine learning model includes a media embedding model that receives the media file and generates at least one of media embeddings and metadata based on the media file.
9. The computer-implemented method of any one of claims 6 or 7, wherein: the machine learning model includes a classifier model that receives the text embeddings and maps the text embeddings to the output indicative of adjustments to be performed to the media file.
10. The computer-implemented method of claim 8, wherein: the machine learning model includes a classifier model that receives the text embeddings and the media embeddings and maps the text embeddings and the media embeddings to the output indicative of adjustments to be performed to the media file.
11. The computer-implemented method of any one of claims 1-10, further comprising: providing, at the media processing application, the adjusted media file as an output via the graphical user interface.
12. The computer-implemented method of any one of claims 1-11, wherein the media file includes at least one of an audio file and a video file.
13. The computer-implemented method of any one of claims 1-1 1, wherein the set of media adjustment tools includes at least one of a noise reduction tool, a de-essing tool, a stereo widening tool, a treble adjustment tool, a mids adjustment tool, a bass adjustment tool, and a boost tool.
14. A non-transitory computer-readable medium comprising executable instructions, which when executed by an electronic processor causes the electronic processor to perform the method of any one of claims 1-13.
15. A system comprising: memory hardware storing instructions; and processor hardware configured to execute the instructions, wherein executing the instructions causes the system to perform the method of any one of claims 1-13.