Sound search

By generating query subtitle embeddings and selecting subtitle embeddings in media files with the highest similarity, the difficulty of searching for specific sounds is solved, and efficient media file search is achieved.

CN120051771APending Publication Date: 2025-05-27QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073194.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-05-31
Filing Date
2023-10-04
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

As the number of media files increases, searching for media content for specific sounds becomes difficult, and the prior art lacks effective audio search solutions.

Method used

By generating query subtitle embeds, select the subtitle embeds in the media file with the highest similarity, and then identify the media file containing a specific sound.

Benefits of technology

It realizes efficient search for specific sounds in a large number of media files, improving the findability and accessibility of media content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051771A_ABST
    Figure CN120051771A_ABST
Patent Text Reader

Abstract

An apparatus includes one or more processors configured to generate one or more query caption inserts based on a query. The processor is further configured to select one or more subtitle inserts from a set of inserts associated with a set of media files of the file repository. Each subtitle embedding represents a corresponding sound subtitle, and each sound subtitle includes a natural language text description of the sound. The subtitle embedding is selected based on a similarity metric indicating a similarity between the subtitle embedding and the query subtitle embedding. The processor is further configured to generate search results identifying one or more first media files of the set of media files. Each of the first media files is associated with at least one of the caption inserts.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] I. Cross - Reference to Related Applications

[0002] This application claims priority to commonly - owned U.S. Provisional Patent Application No. 63 / 380,682, filed on October 24, 2022, and U.S. Non - Provisional Patent Application No. 18 / 326,261, filed on May 31, 2023, the entire contents of each of which are hereby expressly incorporated by reference. Technical Field

[0003] The present disclosure generally relates to searching for specific sounds in media content.

[0004] III. Description of the Related Art

[0005] Advances in technology have led to smaller and more powerful computing devices and an increase in the availability and consumption of media. For example, there are currently a variety of portable personal computing devices, including wireless telephones (such as mobile and smart phones, tablets, and laptop computers), which are small, lightweight, and easy for users to carry and enable the generation and consumption of media content almost anywhere.

[0006] Portable devices capable of capturing audio, video, or both in the form of media files have become very common. A consequence of the availability of such devices is that many people regularly capture media files and store them on their devices to preserve personal memories that they hope to be able to access later. However, as the amount of stored media (e.g., pictures, videos, audio) increases, it becomes difficult to search for desired media content. While certain modern search techniques can be used to search for pictures, there is a lack of solutions for searching the audio of media files (e.g., audio files or video files). Summary of the Invention

[0007] According to a particular aspect, a device includes one or more processors configured to generate one or more query caption embeddings based on a query. The one or more processors are further configured to select one or more caption embeddings from a set of embeddings associated with a set of media files in a file repository. Each caption embedding represents a corresponding audio caption, and each audio caption includes a natural - language text description of a sound. The one or more caption embeddings are selected based on a similarity metric indicating the similarity between the one or more caption embeddings and the one or more query caption embeddings. The one or more processors are further configured to generate search results identifying one or more first media files in the set of media files. Each first media file in the one or more first media files is associated with at least one of the one or more caption embeddings.

[0008] According to certain aspects, a method includes generating, by one or more processors, one or more query subtitle embeddings based on a query. The method further includes selecting, by the one or more processors, one or more subtitle embeddings from a set of embeddings associated with a set of media files in a file repository. Each subtitle embedding represents a corresponding audio subtitle, and each audio subtitle includes a natural language text description of the audio. The one or more subtitle embeddings are selected based on a similarity metric indicating a similarity between the one or more subtitle embeddings and the one or more query subtitle embeddings. The method further includes generating, by the one or more processors, search results identifying one or more first media files in the set of media files. Each first media file in the one or more first media files is associated with at least one of the one or more subtitle embeddings.

[0009] According to certain aspects, a non-transitory computer-readable medium stores instructions that, when executed by one or more processors, cause the one or more processors to generate one or more query subtitle embeddings based on a query. The instructions are further executable to cause the one or more processors to select one or more subtitle embeddings from a set of embeddings associated with a set of media files in a file repository. Each subtitle embedding represents a corresponding audio subtitle, and each audio subtitle includes a natural language text description of the audio. The one or more subtitle embeddings are selected based on a similarity metric indicating a similarity between the one or more subtitle embeddings and the one or more query subtitle embeddings. The instructions are further executable to cause the one or more processors to generate search results identifying one or more first media files in the set of media files. Each first media file in the one or more first media files is associated with at least one of the one or more subtitle embeddings.

[0010] According to certain aspects, a device includes means for generating, based on a query, one or more query subtitle embeddings. The device further includes means for selecting one or more subtitle embeddings from a set of embeddings associated with a set of media files in a file repository. Each subtitle embedding represents a corresponding audio subtitle, and each audio subtitle includes a natural language text description of the audio. The one or more subtitle embeddings are selected based on a similarity metric indicating a similarity between the one or more subtitle embeddings and the one or more query subtitle embeddings. The device further includes means for generating search results identifying one or more first media files in the set of media files. Each first media file in the one or more first media files is associated with at least one of the one or more subtitle embeddings.

[0011] Other aspects, advantages, and features of the present disclosure will become apparent after reviewing the entire application, including the following sections: the Brief Description of the Drawings, the Detailed Description, and the Claims. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a block diagram of specific illustrative aspects of a system operable to search for sounds in media files according to some examples of the present disclosure.

[0013] Figure 2 is according to some examples of the present disclosure Figure 1 of a view of specific aspects of the system.

[0014] Figure 3 is according to some examples of the present disclosure Figure 1 of a view of specific aspects of the system.

[0015] Figure 4 is according to some examples of the present disclosure Figure 1 of a view of specific aspects of the system.

[0016] Figure 5 is according to some examples of the present disclosure Figure 1 of a view of specific aspects of the system.

[0017] Figure 6 is of a view of specific aspects of a voice captioning engine of a system trained according to some examples of the present disclosure Figure 1 of the system.

[0018] Figure 7 is according to some examples of the present disclosure Figure 1 of a view of specific aspects of a voice captioning engine of the system.

[0019] Figure 8 Illustrates an example of an integrated circuit operable to search for sounds in media files according to some examples of the present disclosure.

[0020] Figure 9 is a view of a mobile device operable to search for sounds in media files according to some examples of the present disclosure.

[0021] Figure 10 is a view of a headset operable to search for sounds in media files according to some examples of the present disclosure.

[0022] Figure 11 is a view of a wearable electronic device operable to search for sounds in media files according to some examples of the present disclosure.

[0023] Figure 12 is a view of a mixed reality or augmented reality glasses device operable to search for sounds in media files according to some examples of the present disclosure.

[0024] Figure 13A diagram of earbuds operable to search for sounds in a media file according to some examples of the present disclosure.

[0025] Figure 14 A diagram of a voice-controlled speaker system operable to search for sounds in a media file according to some examples of the present disclosure.

[0026] Figure 15 A diagram of a camera operable to search for sounds in a media file according to some examples of the present disclosure.

[0027] Figure 16 A diagram of headphones (such as virtual reality, mixed reality, or augmented reality headphones) operable to search for sounds in a media file according to some examples of the present disclosure.

[0028] Figure 17 A diagram of a first example of a vehicle operable to search for sounds in a media file according to some examples of the present disclosure.

[0029] Figure 18 A diagram of a second example of a vehicle operable to search for sounds in a media file according to some examples of the present disclosure.

[0030] Figure 19 Is executable by Figure 1 A diagram of a specific implementation of a method for searching for sounds in a media file that can be performed by a device of.

[0031] Figure 20 A block diagram of a specific illustrative example of a device operable to search for sounds in a media file according to some examples of the present disclosure. Detailed Description

[0032] Although much information can be searched from images and videos, the auditory scenes captured by microphones include supplementary information that may not be captured by images and videos alone. Developing techniques for summarizing or understanding auditory scenes is challenging. One step towards developing audio understanding is audio tagging, which involves detecting the occurrence of any common sounds from a limited set of sounds. While audio tagging is useful in some cases, audio captioning would provide a richer set of information because captions use more natural human language to describe sounds. Natural language audio captions can also enable the description of sounds that cannot be directly tagged with predefined labels, such as any common sounds that are not easily categorized into common categories.

[0033] The search of media files is implemented using media captioning and semantic coding. For example, captions are generated to describe specific sounds detected in the media files, and semantic coding is used to generate caption embeddings representing the sound captions. In some implementations, certain sounds detected in the media files may also be processed to generate corresponding audio embeddings representing the sounds. In addition, in some implementations, certain sounds detected in the media files may be processed to generate sound tags describing the sounds, and text embeddings (e.g., tag embeddings) may be generated to represent the sound tags.

[0034] Each media file and its corresponding embedding (e.g., a subtitle embedding, an audio embedding, a tag embedding, or a combination thereof representing the sound in each media file) is stored in a file repository. Metadata associated with the media files may also be stored in the file repository. The embeddings and optionally the metadata can be used to facilitate searching the file repository for audio content of the media files.

[0035] In certain aspects, when a user provides a query, the query can be used to generate a query embedding. Search results can be generated based on similarities between the query embedding and embeddings associated with media files.

[0036] For example, if the query includes natural language text, at least a portion of the text can be used to generate a query subtitle embedding (e.g., a sentence embedding). In this example, the query subtitle embedding can be compared to the subtitle embedding associated with the media file in a subtitle embedding space to determine a similarity measure, and search results can be determined based on the similarity measure. As used herein, a "query subtitle embedding" refers to an embedding that represents multiple words that together form a semantic unit (e.g., a description of a sound).

[0037] As another example, if the query includes audio, the audio may be processed to generate one or more audio captions, which may be processed to generate a query caption embedding. Optionally, the audio may also be processed to generate one or more audio tags and corresponding tag embeddings, processed to generate one or more audio embeddings, or both. The embedding representing the query (e.g., a caption embedding, a tag embedding, an audio embedding, or a combination thereof) may be compared to an embedding associated with a media file to determine a similarity measure.

[0038] Generate search results based on similarity metrics. Determining the similarity of text-based embeddings (e.g., natural language text of captions or tags) in an embedding space provides search results representing concepts that are semantically similar to the concepts present in the query. For example, if the query states "the bell rings multiple times", the search results can list sounds that are captioned as representing "a metal object hitting a metal object". Thus, even if the query does not exactly match the caption, the search results can list sounds with semantically similar descriptors. Additionally, the sounds can include any sound that can be captured and captioned in a media file.

[0039] Certain aspects of the present disclosure are described below with reference to the accompanying drawings. In this specification, common features are designated by common reference numerals. In some of the drawings, multiple instances of a particular type of feature are used. Although these features are physically and / or logically different, the same reference numeral is used for each feature, and the different instances are distinguished by adding letters to the reference numeral. When referring in this document to a group or a type of features (e.g., when no particular feature among these features is being referred to), the reference numeral is used without the distinguishing letter. However, when a particular feature among multiple features of the same type is being referred to in this document, the reference numeral is used together with the distinguishing letter. For example, reference Figure 2 , illustrates multiple media files and these media files are associated with reference numerals 152A and 152N. When referring to a particular one of these media files (such as media file 152A), the distinguishing letter "A" is used. However, when referring to any arbitrary one of these media files or referring to these media files as a group, the reference numeral 152 is used without the distinguishing letter.

[0040] As used herein, various terms are used only for the purpose of describing particular specific implementations and are not intended to limit the specific implementations. For example, the singular forms "a", "an", and "the" are intended to also include the plural forms unless the context clearly indicates otherwise. Additionally, some of the features described herein are singular in some specific implementations and plural in other specific implementations. By way of illustration, Figure 1 depicts a device 102 that includes one or more processors ( Figure 1 the "processor" 190), which indicates that in some specific implementations, the device 102 includes a single processor 190, while in other specific implementations, the device 102 includes multiple processors 190. For ease of reference, such features are typically introduced as "one or more" features and are subsequently referred to in the singular form or the optional plural form (as indicated by "multiple" in the name of the feature) unless aspects related to multiple features of the feature are being described.

[0041] As used herein, the term "comprise" may be used interchangeably with "include". Additionally, the term "wherein" may be used interchangeably with "where". As used herein, "exemplary" indicates examples, specific implementations, and / or aspects, and should not be construed as restrictive or indicating a preference or preferred specific implementation. As used herein, ordinal terms (e.g., "first", "second", "third", etc.) used to modify elements (such as structures, components, operations, etc.) do not themselves indicate any priority or order of the element relative to another element, but merely distinguish the element from another element having the same name (but using an ordinal term). As used herein, the term "set" refers to one or more particular elements within a particular group, and the term "plurality" refers to a plurality (e.g., two or more) of particular elements.

[0042] As used herein, "coupled" may include "communicatively coupled", "electrically coupled", or "physically coupled", and may also (or alternatively) include any combination thereof. Two devices (or components) may be directly or indirectly coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) via one or more other devices, components, wires, buses, networks (e.g., wired networks, wireless networks, or combinations thereof), etc. As an illustrative, non-limiting example, two devices (or components) that are electrically coupled may be included in the same device or in different devices, and may be connected via electronics, one or more connectors, or inductive coupling. In some specific implementations, two devices (or components) that are communicatively coupled (such as electrically connected) may directly or indirectly transmit and receive signals (e.g., digital signals or analog signals) via one or more wires, buses, networks, etc. As used herein, "directly coupled" may include two devices that are coupled (e.g., communicatively coupled, electrically coupled, or physically coupled) without an intermediate component.

[0043] In the present disclosure, terms such as "determine", "calculate", "estimate", "shift", "adjust", etc. may be used to describe how to perform one or more operations. It should be noted that such terms should not be construed as restrictive, and other techniques may be utilized to perform similar operations. Additionally, as mentioned herein, "generate", "calculate", "estimate", "use", "select", "access", and "determine" may be used interchangeably. For example, "generate", "calculate", "estimate", or "determine" a parameter (or signal) may refer to actively generating, estimating, calculating, or determining the parameter (or signal), or may refer to using, selecting, or accessing (such as by another component or device) a parameter (or signal) that has already been generated.

[0044] As used herein, the term "machine learning" shall be understood to have any of its ordinary and customary meanings within the fields of computer science and data science, such meanings including, for example, a process or technique by which one or more computers can learn to perform some operation or function without being explicitly programmed to do so. As a typical example, machine learning can be used to enable one or more computers to analyze data to identify patterns in the data and generate results based on that analysis. For some types of machine learning, the generated results include data models (also referred to as "machine learning models" or simply "models"). Generally, a model is generated using a first dataset to facilitate the analysis of a second dataset. For example, a first portion of a large amount of data can be used to generate a model that can be used to analyze the remaining portion of the large amount of data. As another example, a set of historical data can be used to generate a model that can be used to analyze future data. Examples of machine learning models include, but are not limited to, perceptrons, neural networks, support vector machines, regression models, decision trees, Bayesian models, Boltzmann machines, adaptive neuro-fuzzy inference systems, and combinations, ensembles, and variants of these and other types of models. Variants of neural networks include, for example, but are not limited to, prototype networks, autoencoders, transformers, self-attention networks, convolutional neural networks, deep neural networks, deep belief networks, etc. Variants of decision trees include, for example, but are not limited to, random forests, boosted decision trees, etc.

[0045] Since a machine learning model is generated by a computer based on input data, the machine learning model can be discussed in terms of at least two different time windows (the creation / training phase and the runtime phase). During the creation / training phase, the computer creates, trains, adapts, validates, or otherwise configures the model based on input data (which is typically referred to as "training data" during the creation / training phase). Note that a trained model corresponds to software that has been generated and / or refined during the creation / training phase to perform a specific operation, such as classification, prediction, encoding, or other data analysis or data synthesis operations. During the runtime phase (or "inference" phase), the model is used to analyze input data to generate a model output. The content of the model output depends on the type of the model. For example, as a non-limiting example, a model can be trained to perform a classification task or a regression task. In some specific implementations, the model can be updated continuously, periodically, or occasionally, in which case the training time and the runtime can be interleaved, or one version of the model can be used for inference while a copy is updated, and then the updated copy can be deployed for inference.

[0046] Figure 1 is a block diagram of a specific illustrative aspect of a system 100 operable to search for sounds in media files according to some examples of the present disclosure. The system 100 includes a device 102 that includes one or more processors 190 and a memory 192. In Figure 1In [the figure], the memory 192 includes a file repository 150 that stores media files 152 and associated data, such as metadata 154 associated with each of the media files in the media files 152, audio subtitles 156 associated with the media files 152, and embeddings 158 associated with the media files 152 (e.g., subtitle embeddings 160 associated with the audio subtitles 156). Additionally, in Figure 1 [the figure], the processor 190 includes a media search engine 130 that is configured to perform a search operation to find a portion of the media files 152 that represents a specific sound. Although Figure 1 the illustrated media search engine 130 and file repository 150 are on the same device, in other embodiments, the file repository 150 is stored at a device different from the device 102 that includes the media search engine 130. In some embodiments, the file repository 150, the media search engine 130, or both are distributed across a number of devices (e.g., across a distributed computing system).

[0047] The device 102 is coupled to one or more input devices, such as a microphone 112 and a keyboard 114, via an input interface 104. The device 102 is also coupled to one or more output devices, such as a display device 116 and a speaker 118, via an output interface 106. In some embodiments, one or more of the input devices, output devices, or both are integrated with the processor 190 within the same housing. For example, the microphone 112, the keyboard 114, the display device 116, the speaker 118, or a combination thereof may be built into the device 102. In some embodiments, two or more of the input devices, output devices, or both are combined. For illustration, the display device 116 may include a touchscreen device, and the keyboard 114 may be a virtual keyboard presented via the touchscreen device. Additionally, in some embodiments, the device 102 is coupled to or includes more, fewer, or different input devices; is coupled to or includes more, fewer, or different output devices; or both of the above.

[0048] In a particular aspect, the device 102 is operable to receive a query 120 indicating a target sound via one or more input devices, perform a search operation to identify possible matches to the target sound in a media file or a portion of the media files 152, and provide search results 124 to one or more output devices based on the search operation. The query 120 may include audio or text. Examples of audio queries include the speech of a user describing the target sound. Another example of an audio query includes a non-speech sound representing the target sound. Examples of text queries include a sequence of natural language words describing the target sound.

[0049] Search result 124 may also include sound or text (and possibly other display elements such as graphical elements or hyperlinks). As an example of a search result 124 that includes sound, device 102 may transmit a portion of media file 152 that potentially matches the target sound to speaker 118. As an example of a search result 124 that includes text, device 102 may transmit text indicating media file 152 or a portion of media file 152 that potentially matches the target sound to display device 116.

[0050] In certain aspects, media search engine 130 may be operable to search for any sound in file repository 150. For example, the target sound may include any sound, including for example human voices (e.g., speech sounds) and non-speech sounds (e.g., human-generated sounds other than speech and sounds not generated by humans). By way of illustration, media file 152 may include any number and variety of sounds that may be captured using audio capture equipment (e.g., microphone 112) and stored in digital format in memory 192, and media search engine 130 may be operable to search for any one of these sounds in media file 152. Additionally, media search engine 130 may be operable to search for any type of sound in media file 152 based on a text query or based on an audio query.

[0051] In Figure 1 an example, media search engine 130 includes comparator 140 and a sound captioning engine 146 that includes one or more embedding generators 142. In some specific implementations, media search engine 130 includes additional components as further described below.

[0052] The voice captioning engine 146 is configured to generate voice captions that describe the voices detected in the input audio data. For example, when the new media file 152 is captured or stored in the file repository 150, the voice captioning engine 146 may process the new media file 152 to generate one or more voice captions that describe the voices detected in the new media file 152, and store the voice captions that describe the voices detected in the new media file 152 in the file repository 150 together with the voice captions 156. As used herein, "voice caption" refers to a natural language description of a voice. For example, a voice caption may include a sequence of words that describe the voice together (rather than individually). By way of illustration, a voice caption for the sound of a ringing bell may include text such as "a metal object hitting a metal object". In this illustrative example, the voice caption is an entire clause or sentence that describes the voice. Note that in this illustrative example, no single word of the voice caption represents the voice; rather, the voice caption as a whole is used as a semantic unit to describe the voice. In some specific implementations, as further described below, voice tags may also be used to describe voices. As used herein, "voice tag" refers to a word or word token that describes a voice. By way of illustration, a voice tag that describes the sound of a ringing bell may include text such as "bell" or "ringing". Thus, while both voice captions and voice tags are text tags that describe voices, voice captions include a more natural description of the voice (e.g., how a person presenting the voice might describe it). Thus, using voice captions may facilitate better matching with queries presented to the user. In addition, voice captions can present richer semantic information, which facilitates identifying search results from a wider range of options when used with semantic similarity-based search.

[0053] The voice captioning engine 146 includes one or more machine learning models. For example, the voice captioning engine 146 may include one or more embedding generators 142, each embedding generator corresponding to or including at least one trained machine learning model. For example, the embedding generator 142 may include an audio embedding generator that is configured to receive audio data as input and generate an audio embedding (e.g., a vector or an array) representing the audio data as output. In this example, the voice captioning engine 146 may also include a tag embedding network coupled to the audio embedding network. The tag embedding network is configured to receive the audio embedding as input and generate one or more tag embeddings as output. In this example, each tag embedding represents a word or word token of a voice tag. In addition, in Figure 1 the example, the voice captioning engine 146 may include a caption embedding generator that is configured to generate one or more voice caption embeddings based on the tag embeddings.

[0054] In some embodiments, the speech captioning engine 146 is operable to process data received via query 120 to generate one or more embeddings for use by the media search engine 130. Additionally, the speech captioning engine 146 is operable to generate embeddings 158 representing various data stored in the file repository 150. For example, when a new media file 152 is added to the file repository 150, the speech captioning engine 146 can generate caption embeddings 160 representing the speech captions 156 of the new media file 152. In some embodiments, the speech captioning engine 146 also generates one or more audio embeddings, one or more tag embeddings, or both, representing the sounds detected in the new media file 152 and stores them in the embeddings 158.

[0055] The speech captioning engine 146 can also be used to process audio data received via query 120 during a search of the file repository 150. For example, when query 120 includes audio data, the speech captioning engine 146 generates caption embeddings representing speech captions that describe the sounds detected in the query's audio data. In some embodiments, the speech captioning engine 146 can also process the audio data to generate audio embeddings of the audio data, process one or more speech tags representing the sounds detected in the audio data to generate tag embeddings, or both.

[0056] The comparator 140 is configured to determine a similarity between a query embedding based on query 120 and the embeddings 158 associated with the media file 152. For example, each embedding of a particular type can be considered a vector that specifies a point in the embedding space associated with the embedding of that type. By way of illustration, each caption embedding in the caption embeddings 160 can be considered a vector that specifies a particular location in the caption embedding space. Similarly, if the file repository 150 includes audio embeddings, each audio embedding in the audio embeddings can be considered a vector that specifies a particular location in the audio embedding space. Additionally, if the file repository 150 includes tag embeddings, each tag embedding in the tag embeddings can be considered a vector that specifies a particular location in the tag embedding space. In this example, the comparator 140 determines the similarity between the query embedding and the embeddings in the embeddings 158 based on a metric (e.g., a similarity metric) associated with the relative positions of the two embeddings in the appropriate embedding space. One benefit of such comparisons is that in a text-based embedding space, text-based embeddings with similar semantic content will tend to be closer to each other than embeddings with dissimilar content.

[0057] The generated search results 124 can be sorted (e.g., ranked) based on the values of their similarity metrics. For example, if a first subtitle embedding is closer (in the subtitle embedding space) to the query embedding than a second subtitle embedding, the search result associated with the first subtitle embedding can be ranked higher in the search results 124 than the search result associated with the second subtitle embedding.

[0058] In some specific implementations, the query 120 can be used to generate multiple types of embeddings, and these embeddings are compared with corresponding embeddings 158 (e.g., embeddings of the same type) in the file repository 150. For example, a query subtitle embedding based on the query 120 can be compared with the subtitle embedding 160, a query audio embedding based on the query 120 can be compared with the audio embedding associated with the media file 152, a query tag embedding based on the query 120 can be compared with the tag embedding associated with the media file 152, or a combination thereof. In such specific implementations, search results based on comparisons of different types of embeddings can be weighted differently to generate a ranked list of the search results 124. By way of illustration, to rank the search results 124, a first weight can be applied to the subtitle embedding similarity value, a second weight can be applied to the audio embedding similarity value, and a third weight can be applied to the tag embedding similarity value.

[0059] In some specific implementations, a particular set of embeddings 158 to be compared with the query embedding can be determined at least in part based on the metadata 154. For example, the query 120 can include information describing the target sound (e.g., target sound description) and context items. In this example, the context items can be compared with the metadata 154 to select a subset of the embeddings 158 to be compared with one or more query embeddings based on the target sound description. By way of illustration, the query 120 can include "Where is the video where the ringtone rang many times that I shot last week?" In this illustrative example, the term "video" is a context item indicating the file type of the target media file, "last week" is a context item indicating the time stamp range when the target media file was created, "I shot" is a context item indicating the source of the target media file, and "the ringtone rang many times" is a target sound description of a specific sound in the target media file. In this illustrative example, the comparator 140 compares the embedding based on the target sound description with a subset of the embeddings 158 that satisfy the filtering criteria determined from the context items and are associated with the metadata 154.

[0060] Accordingly, system 100 enables searching for a specific sound in a set of media files 152. System 100 can search for any type of sound, not just specific speech or music samples, for example. Additionally, system 100 can use intuitive search queries 120 (such as natural language text), while optionally also supporting search based on audio queries. When searching based on a text-based query 120, system 100 is able to identify search results 124 associated with sound descriptions (such as sound captions and optionally sound tags) that are semantically similar to query 120. Thus, it is not required that the user generate a query 120 that exactly matches a specific sound description in order to obtain useful search results.

[0061] Figure 2 is a diagram of a specific aspect of a system in accordance with some examples of the present disclosure. Specifically, Figure 1 illustrates an example of the operation of media search engine 130 searching file repository 150 based on text query 202. Note that in some specific implementations, text query 202 may be generated using a speech-to-text engine; however, this is not required. Media search engine 130 operates in the same manner, as described below, regardless of whether the user types text query 202 into Figure 2 device 102, or the user speaks query 120 into device 102 and the speech-to-text engine of device 102 converts the speech into text query. In either case, text query 202 includes a description of the target sound, rather than, for example, specific words to be searched for in media files 152. Figure 1 In

[0062] , media search engine 130 includes comparator 140 and sound captioning engine 146, and the sound captioning engine includes caption embedding generator 242 in Figure 2 . Caption embedding generator 242 is an example of embedding generator 142 of Figure 2 . Caption embedding generator 242 is configured to generate caption embeddings (such as query caption embedding 210) based on a set of text. For example, caption embedding generator 242 may include a sentence embedding generator, and query caption embedding 210 may include sentence embeddings based on the text of text query 202. As another example, caption embedding generator 242 may generate caption embeddings 160 based on the text of sound captions 156. In a particular specific implementation, caption embedding generator 242 includes or corresponds to one or more trained models (such as machine learning models), such as one or more neural networks. Figure 1 In a particular specific implementation, caption embedding generator 242 includes or corresponds to one or more trained models (such as machine learning models), such as one or more neural networks.

[0063] In a particular implementation, the caption embedding generator 242 passes text or one or more word tokens representing the text through one or more neural networks trained to generate query caption embeddings 210. Each query caption embedding can be regarded as a vector indicating a position in a high-dimensional text embedding space (e.g., caption embedding space 220). The one or more neural networks of the caption embedding generator 242 are trained using a large number of sounds and corresponding audio captions such that the positions in the caption embedding space 220 indicate the semantic and syntactic relationships (e.g., similarities) between the audio captions. As a result of such training, the proximity of vectors in the caption embedding space 220 indicates the similarity of the semantic content of the audio captions represented by the vectors.

[0064] In Figure 2 it, the file repository 150 includes a set of media files 252, including media file 152A, media file 152N, and possibly one or more additional media files (as Figure 2 indicated by the ellipsis in). The media files 152 include, for example, audio files, video files, virtual reality files, game files, or combinations thereof. Each media file in the media files 152 is associated with one or more audio captions 156, and each audio caption 156 is associated with a caption embedding 160. Each audio caption 156 includes text (e.g., a sequence of natural language text) describing the sound in the media file 152 associated with the audio caption 156. Since each media file 152 may include more than one sound, more than one audio caption 156 may be associated with each media file 152. In some implementations, the metadata 154 includes a timestamp for each audio caption 156, and the timestamp indicates the time index of the sound described by the audio caption 156 in the media file 152. The metadata 154 may also or alternatively include context information associated with the media file 152. For example, the metadata 154A may indicate the date when the media file 152A was created or added to the file repository 150, the location where the media file 152A was created, the user who created the media file 152A, the name of the media file 152A, etc. In a particular implementation, when a new media file is added to the file repository 150, the corresponding metadata 154, audio caption 156, and caption embedding 158 may be generated and stored together with the media file 152.

[0065] In Figure 2 it, the comparator 140 is configured to determine a similarity metric indicating how similar each caption embedding 160 is to the query caption embedding 210 representing the text query 202. Figure 2A simplified two-dimensional representation including a caption embedding space 220 is shown to illustrate the process for determining the similarity between a caption embedding 160 and a query caption embedding 210; however, in an actual implementation, the embedding space will be a high-dimensional space defined by hundreds or thousands of orthogonal axes.

[0066] In Figure 2 , an asterisk is used to indicate the position of the query caption embedding 210 in the caption embedding space 220, and a circle is used to indicate the position of each caption embedding 160 in the caption embedding space 220. During operation, the comparator 140 is configured to determine a similarity metric for each caption embedding 160 based on the distance 214 between the caption embedding 160 and the query caption embedding 210 in the caption embedding space 220. The distance 214 can be determined as, for example, a cosine distance, an Euclidean distance, or based on some other distance metric. In Figure 2 the example shown, the distance 214A between the position of the query caption embedding 210 and the position of the caption embedding 160A is less than the distance 214B between the position of the query caption embedding 210 and the position of the caption embedding 160N, which indicates that the audio caption 156A represented by the caption embedding 160A is semantically more similar to the text of the text query 202 than the audio caption 156N represented by the caption embedding 160N. In some implementations, the distance 214 is used as the similarity metric. In such cases, a smaller value of the similarity metric (indicating a smaller distance) indicates a closer match between the audio caption 156 and the text of the text query 202. In some implementations, the similarity metric is calculated based on the distance 214 such that a larger value of the similarity metric indicates a smaller distance 214 (and thus indicates a closer match between the audio caption 156 and the text of the text query 202).

[0067] The media search engine 130 generates search results 124 based on the similarity metrics associated with the audio captions 156. For example, the search results 124 can identify one or more media files 152 (or portions of the media files 152, such as a portion of the media file 152 associated with a particular audio caption 156) associated with a set of caption embeddings 160 that are most similar to the query caption embedding 210 in the set of media files 252. If the metadata 154 includes a time index associated with the particular audio identified in the search results 124, that time index can be indicated in the search results 124. In some implementations, the search results 124 include a ranked list of results. In such implementations, the search results 124 can be sorted based on their respective similarity metrics.

[0068] In some embodiments, the search results 124 list each media file (or portion of media file 152) among these media files in ranked order based on a similarity metric of the voice captions 156 of the media file 152. In some embodiments, the media search engine 130 restricts the search results 124 to include only information associated with the media file 152 that is associated with caption embeddings 160 within a threshold distance 216 of the query caption embedding 210 in the caption embedding space 220. In such embodiments, the threshold distance 216 may be preset (e.g., based on user-configurable options) or may be determined dynamically. For example, the threshold distance 216 may be determined such that a particular percentage or other ratio of the caption embeddings 160 is within the threshold distance 216. By way of illustration, the threshold distance 216 may be determined to include no more than 25%, 50%, 75%, or some other percentage of the caption embeddings 160. Although Figure 2 the threshold is illustrated in terms of distance in the caption embedding space 220, in other embodiments, the threshold may instead be applied in terms of a similarity metric. In still other embodiments, the search results 124 may be restricted to include only a specific number of the most similar results (e.g., regardless of their distance in the caption embedding space 220).

[0069] Figure 3 is a diagram of a particular aspect of a Figure 1 system in accordance with some examples of the present disclosure. Specifically, Figure 3 illustrates an example of the operation of the media search engine 130 to search the file repository 150 based on a text query 202 that includes a description of a target voice (e.g., target voice description 304) and one or more context items 306.

[0070] In Figure 3 , the media search engine 130 includes a comparator 140 and a voice captioning engine 146 (which includes a caption embedding generator 242), each of the comparator and the voice captioning engine operating as described above with reference to Figure 2 . For example, the caption embedding generator 242 is configured to generate caption embeddings (e.g., sentence embeddings) based on a set of texts of the text query 202. In the example shown in Figure 3 , the set of texts provided to the caption embedding generator 242 corresponds to the target voice description 304. In other embodiments, the target voice description 304 and the context items 306 are provided to the caption embedding generator 242. The comparator 140 is configured to determine a similarity metric that indicates how similar the caption embeddings 160 of the set of media files 252 are to the query caption embedding 210 representing the target voice description 304.

[0071] In Figure 3In [the figure], the media search engine 130 also includes a natural language processor 302 and a filter 344. The natural language processor 302 is configured to process the text query 202 to identify the part of the text query 202 corresponding to the target sound description 304 and the part of the text query 202 corresponding to the context item 306. As Figure 3 shown, at least the target sound description 304 is provided to the subtitle embedding generator 242, and the context item 306 is provided to the filter 344.

[0072] The filter 344 is configured to select, based on the context item 306, one or more media files 152 associated with metadata 154 that meet the filtering criteria from the file repository 150. For example, the context item 306 may indicate a period of interest, and the filter 344 may compare the timestamp of the metadata 154 with the period of interest to determine which media files 152 have timestamps within the specified period. As another example, the context item 306 may indicate a target file type (e.g., a video file, an audio file, or another type of file), and the filter 344 may compare the file type information of the metadata 154 with the target file type to determine which media files 152 have the target file type. In other examples, the filter 344 may apply different filtering criteria (as a supplement to or in place of the time criteria and / or the file type criteria). Non-limiting examples of such filtering criteria include the location where the media file was generated, the source of the media file, etc. In some specific implementations, the filter 344 may also receive input from other types of media search engines. By way of illustration, an image search engine may be used to label objects (e.g., faces) identified in a specific video file of the media file 152, and such object labels may be stored in the metadata 154 and used by the filter 344.

[0073] In a specific implementation, the filter 344 is configured to select a set of embeddings 346 that meet the filtering criteria from the subtitle embeddings 160 of the set of media files 252. In this implementation, the filter 344 pre-screens the subtitle embeddings 160 to reduce the number of similarity metric calculations performed by the comparator 140. For example, Figure 3 the comparator 140 in [the figure] only needs to calculate the similarity metric for the subtitle embeddings that meet the filtering criteria for the set of embeddings 346, rather than all subtitle embeddings 160 of the file repository 150. By way of illustration, in Figure 3In the caption embedding space 220, the caption embeddings 160 that satisfy the filtering criteria are indicated by circles in the caption embedding space 220 (such as caption embedding 160A), and the caption embeddings 160 that fail to pass one or more of the filtering criteria are illustrated by squares (such as caption embedding 160N). In this illustrative example, the comparator 140 determines the similarity metric for the caption embedding 160A and does not determine the similarity metric for the caption embedding 160N. Thus, the computing resources for determining the similarity metric are conserved.

[0074] In Figure 3 the example shown, the media search engine 130 generates search results 124 based on the similarity metric associated with the audio caption 156. In Figure 3 the example, the search results 124 include only results related to the media files 152 associated with the metadata 154 that satisfy the filtering criteria applied by the filter 344.

[0075] In some alternative embodiments, the filtering criteria can be used to determine the weights applied to the similarity metrics used to rank the search results 124. For example, in some such embodiments, the comparator 140 determines the similarity metric for the caption embeddings 160 associated with the metadata 154 that do not satisfy the filtering criteria; however, such caption embeddings are weighted adversely such that their ranking in the ranked search results 124 is lower than their ranking when the metadata 154 satisfies the filtering criteria. By way of illustration, in one such embodiment, the comparator 140 determines the similarity metric for the caption embedding 160N even though the metadata 154N associated with the caption embedding 160N does not satisfy the filtering criteria. In this illustrative example, the similarity metric associated with the caption embedding 160N is weighted adversely during the ranking of the search results 124. By way of illustration, in Figure 3 the caption embedding 160A and the caption embedding 160N are approximately equidistant from the query caption embedding 210, resulting in approximately equal similarity metrics; however, due to the adverse weighting, the results associated with the caption embedding 160N can be arranged lower in the search results 124 than the results associated with the caption embedding 160A.

[0076] Figure 4 is a diagram of a particular aspect of a Figure 1 system in accordance with some examples of the present disclosure. Specifically, Figure 4 illustrates an example of the operation of the media search engine 130 searching the file repository 150 based on an audio query 402 that includes an audio sample representing a target sound. For Figure 4For the purposes of the example shown, the audio query 402 includes any type of sound input (e.g., speech, non-speech sounds, ambient sounds, etc.). Additionally, even if the audio query 402 includes speech, the speech is processed like any other sound. For example, in Figure 4 speech is treated as ambient sound in the same way, and other speech is treated as non-human-generated sound.

[0077] In Figure 4 the media search engine 130 includes a sound captioning engine 146 and a comparator 140. In Figure 4 the sound captioning engine 146 includes an audio embedding generator 404, a label embedding generator 408, and a caption embedding generator 242. As described with reference to Figure 1 the sound captioning engine 146 is configured to generate sound captions that describe the sound (e.g., the sound in the media file 152, the sound in the audio query 402, or both). In Figure 4 the example shown, the sound captioning engine 146 is configured to generate one or more query caption embeddings 210 that represent sound captions that describe the sound detected in the audio query 402. For example, the audio query 402 may include audio captured by Figure 1 the microphone 112, where the audio represents ambient sounds in a soundscape, such as the rustling of leaves or the sound of a storm. In such examples, the query sound caption 210 may include "wind blowing through the leaves on the ground" or "heavy rain falling on a surface". Refer to Figure 6 for a non-limiting method of training the sound captioning engine 146.

[0078] In Figure 4 the example shown, the audio data of the audio query 402 is provided to the audio embedding generator 404. The audio embedding generator 404 generates one or more query audio embeddings 406 that represent the audio data. For example, the audio data includes information describing the waveform of the audio in the audio query 402, and the query audio embedding 406 represents the audio data.

[0079] In Figure 4 the example shown, the query audio embedding 406 is provided as an input to the label embedding generator 408 to generate one or more query label embeddings 410. Each query label embedding in the query label embeddings 410 represents a sound label that describes the sound detected in the audio data. The sound label may include a word or a word token.

[0080] In Figure 4In the example shown, the query tag embedding 410 is provided as an input to the caption embedding generator 242, which is configured to generate a query caption embedding 210 based on the query tag embedding 410. The query caption embedding 210 represents a sound caption that describes the sound detected in the audio data. Typically, the sound caption includes multiple words, such as a sequence of words that form a natural language description of the sound.

[0081] Figure 4 The comparator 140 operates as described above with reference to Figures 1 to 3 For example, the comparator 140 determines a similarity metric by comparing the query caption embedding 210 with the caption embedding 162 in the caption embedding space 220. The search result 124 is at least partially based on the similarity metric obtained by comparing the query caption embedding 210 with the caption embedding 162.

[0082] Figure 5 is an illustration of a particular aspect of a Figure 1 system according to some examples of the present disclosure. Specifically, Figure 5 illustrates an example of the operation of the media search engine 130 to search the file repository 150 based on a query 120, which may include audio or text. If the query 120 includes audio, the audio may include any type of sound input as described with reference to Figure 4 (e.g., speech, non-speech sounds, ambient sounds, etc.).

[0083] In Figure 5 the example shown, the file repository 150 includes the set of media files 252, and each media file 152 is associated with metadata 154, a sound caption 156A, and a corresponding caption embedding 162, as described above. Additionally, in Figure 5 the example, one or more of the media files 152 may optionally be associated with one or more audio embeddings 530. For example, in Figure 5 the media search engine 130 includes an audio embedding generator 404, which is configured to generate an audio embedding for a particular sound. The audio embedding generator 404 includes one or more trained machine learning models, such as one or more of the models described with reference to Figure 6 The audio embedding generator 404 is operable to process the media files 152 to generate audio embeddings 530 associated with one or more of the media files 152. Each audio embedding 530 represents the sound characteristics of an audio sample, such as waveform parameters that describe the waveform of the audio sample.

[0084] Additionally, in Figure 5In the example of, one or more media files in media file 152 may optionally be associated with one or more sound tags 512 and corresponding tag embeddings 514. For example, in Figure 5 , media search engine 130 includes a tag embedding generator 408 configured to generate sound tags that describe specific sounds. The tag embedding generator 408 includes one or more trained machine learning models, such as one or more of the models described in reference Figure 6 . The tag embedding generator 408 is operable to process media file 152 to generate sound tags 512 associated with one or more media files in media file 152. As described above, sound tags include words or word tokens that describe specific sounds.

[0085] During the operation of the media search engine 130 in Figure 5 , a preprocessor 502 of the media search engine 130 receives query 120 and determines whether the query includes audio or text. If query 120 includes text, the media search engine 130 processes the text of query 120 as described above in reference Figure 2 . For example, in Figure 5 , the preprocessor 502 optionally includes Figure 3 's natural language processor 302. In this example, the natural language processor 302 is operable to determine a target sound description 304 from the text of query 120 and provide the target sound description 304 to the tag embedding generator 408, the caption embedding generator 242, or both. In this example, the caption embedding generator 242 generates a query caption embedding 210 based on the target sound description 304. In some embodiments, the tag embedding generator 408 generates a query tag embedding 410 based on the target sound description 304, and the caption embedding generator 242 generates a query caption embedding 210 based on the query tag embedding 410. The comparator 140 determines a similarity metric by comparing the query caption embedding 210 with the caption embedding 162. In this example, the search results 124 are based on the similarity metric of comparing the query caption embedding 210 with the caption embedding 162.

[0086] In a particular implementation in which the target sound description 304 is processed by the tag embedding generator 408 to generate the query tag embedding 410, the comparator 140 may also determine a similarity metric by comparing the query tag embedding 410 with the tag embeddings 514 in the tag embedding space. For example, the comparator 140 may determine a similarity metric for each tag embedding 410 based on the distance between the query tag embedding 410 and the tag embeddings 514 in the tag embedding space 510. The distance may be determined as, for example, a cosine distance, an Euclidean distance, or based on some other distance measure. In some particular implementations, the query tag embedding 410 includes more than one tag embedding for each query caption embedding 210 (e.g., more than one sound tag associated with each detected sound). In some such particular implementations, the tag embedding 514 associated with the media file 152 may be compared with a representative query tag embedding 410 (e.g., the query tag embedding 410 closest to the centroid of the multiple query tag embeddings 410). In other such particular implementations, the tag embedding 514 associated with the media file 152 may be compared with each query tag embedding 410 of the multiple query tag embeddings 410, and a representative distance, such as the average distance between the tag embedding 514 and each query tag embedding of the multiple query tag embeddings, may be determined. In still other such particular implementations, the tag embedding 514 associated with the media file 152 may be compared with a location (such as the centroid of the locations of the multiple query tag embeddings 410) representing the locations of the multiple query tag embeddings 410 in the tag embedding space 510.

[0087] If the query 120 includes audio, the media search engine 130 processes the audio of the query 120 as described above with reference to Figure 4 . For example, in Figure 5 , the preprocessor 502 provides the audio data 520 to the sound captioning engine 146 based on the query 120. In this example, the sound captioning engine 146 determines the query caption embedding 210 to represent the sound caption describing the sound in the audio data 520. In this example, the comparator 140 determines a similarity metric by comparing the query caption embedding 210 with the caption embeddings 162 in the caption embedding space 220. In this example, the search results 124 are at least partially based on the similarity metric obtained by comparing the query caption embedding 210 with the caption embeddings 162.

[0088] Optionally, in the example of Figure 5 , the tag embedding generator 408 of the sound captioning engine 146 generates query sound tags 410 representing one or more sound tags describing the sound in the audio data 520, as described above. For example, the sound tags 512 may be generated during the process of generating the query caption embedding 210. In Figure 5In the example shown, comparator 140 can determine a similarity metric by comparing label embedding 410 in label embedding space 510 with label embedding 514, as described above. In this example, search result 124 is at least partially based on the similarity metric resulting from comparing query label embedding 410 with label embedding 514.

[0089] Optionally, in Figure 5 the example of, the audio embedding generator 404 of the voice captioning engine 146 generates a query audio embedding 406 representing the sound in the audio data 520 and provides the query audio embedding 406 to the comparator 140. For example, the query audio embedding 406 can be generated during the process of generating the query caption embedding 210.

[0090] In Figure 5 the example shown, comparator 140 can determine a similarity metric by comparing query audio embedding 406 in audio embedding space 526 with audio embedding 530. For example, comparator 140 can determine a similarity metric for each audio embedding 406 based on the distance between query audio embedding 530 and audio embedding 530 in audio embedding space 526. This distance can be determined as, for example, a cosine distance, an Euclidean distance, or based on some other distance metric. In some embodiments, the query audio embedding 406 includes more than one audio embedding for each query caption embedding 210. In some such embodiments, the audio embedding 530 associated with the media file 152 can be compared with a representative query audio embedding 406 (e.g., the query audio embedding 406 closest to the centroid of the plurality of query audio embeddings 406). In other such embodiments, the audio embedding 530 associated with the media file 152 can be compared with each query audio embedding 406 among the plurality of query audio embeddings 406, and a representative distance, such as the average distance between the audio embedding 530 and each query audio embedding among the plurality of query audio embeddings, can be determined. In still other such embodiments, the audio embedding 530 associated with the media file 152 can be compared with a location (such as the centroid of the locations of the plurality of query audio embeddings 406) in the audio embedding space 526 representing the locations of the plurality of query audio embeddings 406.

[0091] In Figure 5In the example shown, query 120 can be used to generate multiple types of embeddings, which are compared with corresponding embeddings in file repository 150. For example, query caption embedding 210 can be compared with caption embedding 162, query audio embedding 406 can be compared with audio embedding 530 associated with media file 152, query tag embedding 410 can be compared with tag embedding 514 associated with media file 152, or a combination thereof. In some embodiments, when comparing multiple types of embeddings with corresponding embeddings in file repository 150, each comparison can generate one or more search results. In such embodiments, search results based on comparisons of different types of embeddings can be weighted differently to generate a ranked list of search results 124. For illustration, to sort search results 124 based on ranking, a first weight can be applied to the similarity value based on the comparison of caption embedding 162 with query caption embedding 210, a second weight can be applied to the similarity value based on the comparison of audio embedding 530 with query audio embedding 406, and a third weight can be applied to the similarity value based on the comparison of tag embedding 514 with query tag embedding 410.

[0092] Figure 6 is a training according to some examples of the present disclosure Figure 1 of a sound captioning engine (such as sound captioning engine 670) of the system shown. In Figure 6 it, sound captioning engine 670 includes multiple machine learning models, which include audio embedding generator 680, tag embedding generator 684, and caption embedding generator 688. In a particular embodiment, audio embedding generator 680, tag embedding generator 684, and caption embedding generator 688 respectively represent examples of audio embedding generator 404, tag embedding generator 408, and caption embedding generator 242 during a training process (e.g., before the machine learning parameters (such as link weights) of audio embedding generator 404, tag embedding generator 408, and caption embedding generator 242 are fixed).

[0093] In Figure 6 the training process shown, loss calculator 640 determines a loss metric based on one or more differential calculations ( Figure 6 "differential calculations" in) among multiple differential calculations. The training process is iterative, and during each iteration, the change in the loss metric is used to adjust the machine learning parameters of one or more of the machine learning models in sound captioning engine 670. For example, machine learning optimizer 650 can use one or more backpropagation operations (or another machine learning optimization process) to adjust the machine learning parameters of the machine learning models to reduce the loss metric.

[0094] The training process uses a set of captioned training data 602. The captioned training data 602 includes a large number of audio data samples and corresponding labels. Each audio data sample includes a representation of a particular sound, and each label associated with the audio data sample includes a description of the sound. The labels can include, for example, sound labels, sound captions, or both that are considered to be correct. For example, each label assigned to a sound can be based on a description generated by a person after listening to the sound.

[0095] During an iteration of the training process, audio data 604 representing a sound is provided as input to an audio embedding generator 680. The audio embedding generator 680 generates one or more audio embeddings 682 representing the audio data 604, and the audio embeddings 682 are provided as input to a label embedding generator 684. As an example, the audio embedding generator 680 includes a neural network that is configured to take as input a spectrogram of the audio data. In this example, the audio embedding generator 680 can include one or more convolutional layers that are configured to and optionally pre-trained to process the audio data 604 to generate the audio embeddings 682 (e.g., the audio embedding generator 680 can be a convolutional neural network (CNN)).

[0096] The predicted token embedding 610 is determined based on the state or output of one or more layers of the label embedding generator 684. In some embodiments, the predicted token embedding 610 is output by the last layer of the label embedding generator 684. In other embodiments, the predicted token embedding 610 is determined based on the state or output of one or more hidden layers of the label embedding generator 684. For example, the output layer of the label embedding generator 684 can be configured to generate a one-hot vector identifying a single label for the input audio embedding 682. In this example, the predicted token embedding 610 can include a vector of floating-point values for generating the one-hot vector. In some such embodiments, the predicted token identifier 620 ( Figure 6 "predicted token ID" in

[0097] In some specific implementations, the audio embedding generator 680, the tag embedding generator 684, or both are pre-trained machine learning models. Examples of machine learning models that can be used as or included in the audio embedding generator 680 and the tag embedding generator 684 include PANN; YAMNet; VGGish; and modifications to AlexNet, Inception V3, or ResNet (PANN refers to the neural network described by Kong et al. in the paper "Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition". YAMNet refers to a pre-trained audio event classifier purchased from TensorFlow Hub. VGGish is a pre-trained convolutional neural network purchased from Google. Modifications to AlexNet, Inception V3, and ResNet refer to the neural networks described by Hershey et al. in the paper "CNN ARCHITECTURES FOR LARGE-SCALE AUDIO CLASSIFICATION". In a specific implementation where the audio embedding generator 680 and the tag embedding generator 684 are pre-trained or trained independently of the caption embedding generator 688, the machine learning optimizer 650 may not modify the machine learning parameters (e.g., weights) of the audio embedding generator 680 during the training process, may not modify the machine learning parameters (e.g., weights) of the tag embedding generator 684 during the training process, or both of the above. Alternatively, the pre-trained audio embedding generator 680 and tag embedding generator 684 can be used as a starting point for further training, in which case the machine learning optimizer 650 can Figure 6 further optimize (e.g., modify) the machine learning parameters (e.g., weights) of the audio embedding generator 680, the tag embedding generator 684, or both during the training process shown.

[0098] In Figure 6 the example shown, the predicted token embedding 610 is provided as input to the caption embedding generator 688. The caption embedding generator 688 generates the predicted caption embedding 630 based on a set of one or more predicted token embeddings 610. In a specific implementation, the caption embedding generator 688 is a multi-head self-attention network, such as a transformer network (such as Sentence-Bert). Figure 7 An example of using a transformer-based embedding generator as part of the caption embedding generator 630 is shown.

[0099] The loss calculator determines a loss metric based on difference calculation 634. In some specific implementations, the loss metric is further based on either or both of difference calculation 614 and difference calculation 624. Difference calculation 614 is based on a comparison of the predicted token embedding 610 and the ground truth token embedding 612 for the same sound. In the context of training, "ground truth" indicates that a value or parameter (e.g., a label) is assigned by a human or otherwise sufficiently verified to be considered reliable. In Figure 6 the example shown, the ground truth token embedding 612 for each sound can be indicated in the captioned training data 602. In a particular specific implementation, the loss metric can be determined at least in part based on a similarity metric indicating the similarity between the ground truth token embedding 612 for a sound and the predicted token embedding 610 for the sound. By way of illustration, the similarity metric can be determined as the cosine distance between the ground truth token embedding 612 for a sound and the predicted token embedding 610.

[0100] Difference calculation 624 is based on a comparison of the predicted token embedding 620 and the ground truth token identifier 622 for a particular sound. In Figure 6 the example shown, the ground truth token identifier 622 for each sound can be indicated in the captioned training data 602. In a particular specific implementation, the loss metric can be determined at least in part based on a similarity metric indicating the similarity between the ground truth token identifier 622 for a sound and the predicted token identifier 620 for the sound. By way of illustration, the similarity metric can be determined as the cross-entropy loss between the ground truth token identifier 622 for a sound and the predicted token identifier 620.

[0101] Difference calculation 634 is based on a comparison of the predicted caption embedding 630 and the ground truth caption embedding 632 for a particular sound. In Figure 6 the example shown, the ground truth caption embedding 632 for each sound can be determined by the caption embedding generator 688 based on the ground truth token embedding 612 for the particular sound. In other specific implementations, the ground truth caption embedding 632 for each sound can be indicated in the captioned training data 602. In a particular specific implementation, the loss metric can be determined at least in part based on a similarity metric indicating the similarity between the ground truth caption embedding 632 for a sound and the predicted caption embedding 630 for the sound. By way of illustration, the similarity metric can be determined as the cosine distance between the ground truth caption embedding 632 for a sound and the predicted caption embedding 630.

[0102] The machine learning optimizer 650 is operable to modify machine learning parameters (e.g., weights) of the audio embedding generator 680, the label embedding generator 684, the caption embedding generator 688, or a combination thereof to reduce a loss metric. In some embodiments, the audio embedding generator 680 is pre-trained and static, and the machine learning optimizer 650 is operable to modify machine learning parameters (e.g., weights) of the label embedding generator 684, the caption embedding generator 688, or both, to reduce a loss metric. In some embodiments, as referenced Figure 7 The audio embedding generator 680, the label embedding generator 684, and the caption embedding generator 688 are selected to be differentiable such that backpropagation can be used to modify (e.g., train or fine-tune) machine learning parameters (e.g., weights) of each based on the loss metric, as described.

[0103] As a specific non-limiting example, the PANN machine learning model can be used as the audio embedding generator 680, and a stacked arrangement of two transformer decoder layers with four heads and gelu activation can be used as the label embedding generator 684. In this example, the label embedding generator 684 can be trained to generate word / token embeddings (e.g., the predicted token embeddings 610) that are provided to the caption embedding generator 688. The word / token embeddings are further projected into a space whose dimension is equal to the size of the vocabulary such that the prediction can be expressed as a one-hot encoded vector (e.g., as the predicted token identifier 620 corresponding to each predicted token embedding 610). For example, the predicted token identifier 620 can be based on 128-dimensional word2vec.

[0104] The loss calculator 640 attempts to reduce (e.g., minimize) the cross-entropy loss between the one-hot encoded vector of the ground truth token identifier 622 and the corresponding predicted token identifier 620. Training to make the word / token embeddings (e.g., the predicted token embeddings 610) more accurate can also be improved by configuring the loss calculator 640 to determine the loss metric based in part on the cosine distance between the word / token embeddings (e.g., the ground truth token embedding 612 and the corresponding predicted token embedding 610).

[0105] Because sentence embeddings can represent the gist of multiple labels (e.g., semantic and syntactic content), the loss calculator 640 can also be configured to determine the loss metric based at least in part on the cosine similarity between the ground truth caption embedding 632 and the corresponding predicted caption embedding 630.

[0106] In addition, attaching the subtitle embedding generator 688 to the label embedding generator 684 allows the machine learning optimizer 650 to directly update the machine learning parameters of the label embedding generator 684, the subtitle embedding generator 688, or both via backpropagation. For example, when training the label embedding generator 684, the predicted token embeddings 610 generated by the label embedding generator 684 are directly fed into the subtitle embedding generator 688, and the differential calculation 634 is used to update the weights of the label embedding generator 684. Thus, the weights of the label embedding generator 684 can be directly optimized to reduce (e.g., minimize) the distance between the subtitle embeddings 630, 632, and thus make the generated subtitles closer in meaning to the reference subtitles.

[0107] In some specific embodiments, Sentence-BERT is used as the subtitle embedding generator 688 and is configured or trained to distinguish whether two sentences entail each other, contradict each other, or are neutral to each other. In this example, the label embedding generator 684 is also a BERT network, such that the predicted token embeddings 610 generated by the label embedding generator 684 can be directly input into the subtitle embedding generator 688 (e.g., Sentence-BERT) to achieve end-to-end backpropagation. In other specific embodiments, other machine learning models are used to replace or supplement Sentence-BERT. For example, word2vec or FastText can be used as the subtitle embedding generator 688.

[0108] Figure 7 is according to some examples of the present disclosure Figure 1 is a diagram of a specific aspect of the voice captioning engine of the system of. Specifically, Figure 7 illustrates an example of a subtitle embedding generator 788. During training, the subtitle embedding generator 788 corresponds to Figure 6 an example of the subtitle embedding generator 688 of. During use (e.g., after training and during inference), the subtitle embedding generator 788 corresponds to Figures 2 to 5 an example of the subtitle embedding generator 242 of any of.

[0109] In Figure 7 the example shown, the subtitle embedding generator 788 uses a neural network-based embedding generator 716 (e.g., a BERT network, another multi-head attention-based network, a word2vec network, a FastText network, etc.) and one or more pooling layers 718. However, compared with conventional language processing models, the neural network-based embedding generator 716 is configured to receive based on the label embedding generated by a label embedding generator (e.g., Figure 4 or Figure 5 the label embedding generator 408 of, or Figure 6The input of the token embeddings 706 output by the label embedding generator 684 (e.g., the generator input 714), so all operations of the subtitle embedding generator 788 are differentiable, which enables backpropagation training for each of the machine learning models of the voice captioning engine 670. Figure 6 For the machine learning models of the voice captioning engine 670.

[0110] For example, conventional language processing models include non-differentiable operations 702 for preparing the input for a neural network (e.g., a BERT model). As Figure 7 shown, examples of such non-differentiable operations 702 include tokenization of the input and / or a lookup operation to determine token embeddings based on the input tokens. Omitting the non-differentiable operations 702 enables simultaneous backpropagation training of the weights of all machine learning models of the voice captioning engine (e.g., Figures 1 to 5 the voice captioning engine 146 or Figure 6 the voice captioning engine 670) or any subset thereof. For example, referring to Figure 6 , the machine learning optimizer 650 can use backpropagation to train or fine-tune the weights of the audio embedding generator 680, the label embedding generator 684, and / or the subtitle embedding generator 688 to minimize the loss function based on the differential calculation 634. This arrangement has the additional benefit of using a loss function that directly represents the metric of interest (e.g., how closely the captions generated by the voice captioning engine match the captions that would be assigned by a human). Additionally, training using a loss function based on subtitle embeddings (e.g., the differential calculation 634) enables optimization of the semantic similarity of the captions rather than the same caption language. For example, instead of using subtitle embeddings, the loss function for training can use the predicted captions (or the predicted labels represented by the differential calculation 624). In this case, the goal would be to minimize the cross-entropy loss of the one-hot vectors, which promotes the generation of the predicted captions or labels that are the same as the ground truth captions or labels (e.g., represented by the ground truth label ID 622 or the one-hot encoding based on the ground truth subtitle embedding 632). However, loss functions based on label embeddings (e.g., the differential calculation 614) and / or subtitle embeddings (e.g., the differential calculation 634) can be used to minimize the distance in the embedding space, which effectively translates to increased semantic similarity between the predicted captions and the ground truth captions, between the predicted labels and the ground truth labels, or both.

[0111] In Figure 7 the example shown, the non-differentiable operations 702 are omitted and only the differentiable operations 704 are used. For example, the generator input 714 is based on the output of the label embedding generator (e.g., Figure 4 or Figure 5 the label embedding generator 408 or Figure 6The token embeddings 706 generated by the label embedding generator 684), thus avoiding tokenization and embedding lookup operations. In Figure 7 In the example shown, the generator input 714 of the neural network-based embedding generator 716 may also include a segment embedding 710 and a position embedding 712.

[0112] In Figure 7 In the example shown, the neural network-based embedding generator 716 generates a caption representation 720 based on the generator input 714. The caption representation 720 is aggregated by the pooling layer 718 to generate a caption embedding 730. In some specific implementations, the caption representation 720 corresponds to the output of the last layer of the neural network-based embedding generator 716. In other specific implementations, the caption representation 720 corresponds to the output or state of one or more hidden layers of the neural network-based embedding generator 716. The caption embedding 730 corresponds to or includes Figures 1 to 5 the caption embedding 160 of Figures 2 to 5 the query caption embedding 210 of Figure 6 the ground truth caption embedding 632 of Figure 6 or an example of any of the predicted caption embeddings 630 of

[0113] Figure 8 depicts a specific implementation 800 of the device 102 as an integrated circuit 802 including one or more processors 190. The integrated circuit 802 includes a signal input 804 (such as one or more bus interfaces) to enable the reception of a query 120 for processing. The integrated circuit 802 also includes a signal output 806 (such as a bus interface) to enable the transmission of an output signal, such as an output 124 representing search results. In Figure 8 In the example shown, the processor 190 includes a comparator 140 and a speech captioning engine 146. The integrated circuit 802 enables the operation of searching for speech in media content. The integrated circuit 802 may be integrated within one or more other devices, such as a mobile phone or tablet computer as depicted in Figure 9 as depicted in Figure 10 headphones as depicted in Figure 11 a wearable electronic device as depicted in Figure 12 a mixed reality or augmented reality glasses device as depicted in Figure 13 earbuds as depicted in Figure 14 a voice-controlled speaker system as depicted in Figure 15 a camera as depicted in Figure 16 a virtual reality, mixed reality or augmented reality headset, or as depicted in Figure 17 or Figure 18 a vehicle as depicted in

[0114] As an illustrative, non-limiting example, Figure 9 Illustrative embodiment 900 is depicted in which device 102 includes a mobile device 902, such as a phone or tablet computer. Mobile device 902 includes a microphone 112, a camera 906, and a display screen 904. The components of processor 190, including media search engine 130, are integrated within mobile device 902 and are illustrated using dashed lines to indicate internal components that are generally not visible to the user of mobile device 902. In a particular example, media search engine 130 is operable to search user-generated media content, downloaded media content, and / or other media content stored in the on-board memory of mobile device 902 or stored at a remote memory, such as a server or cloud-based file repository. For example, a user may provide a query (e.g., Figure 1 query 120) by entering text via the touch screen of mobile device 902, by providing speech via microphone 112 that describes a target sound, or by capturing a target sound via microphone 112. In this example, media search engine 130 performs one or more of the search operations described in reference Figures 1 to 5 to generate search results based on the query. The search results may be provided to the user on display screen 904 or as an audio output to speaker 118.

[0115] Figure 10 Illustrative embodiment 1000 is depicted in which device 102 includes a headset device 1002. Headset device 1002 includes a microphone 112. The components of processor 190, including media search engine 130, are integrated within headset device 1002. In a particular example, media search engine 130 is operable to search user-generated media content, downloaded media content, and / or other media content stored in the on-board memory of headset device 1002 or stored at a remote memory (e.g., at a mobile device, gaming system, computer, server, or cloud-based file repository accessible to headset device 1002). For example, a user may provide a query (e.g., Figure 1 query 120) by providing speech via microphone 112 that describes a target sound, or by capturing a target sound via microphone 112. In this example, media search engine 130 performs one or more of the search operations described in reference Figures 1 to 5 to generate search results based on the query. The search results may be provided to the user on a display of a computing device or as an audio output to speaker 118.

[0116] Figure 11Illustrates a specific implementation 1100 in which device 102 includes a wearable electronic device 1102 (illustrated as a "smartwatch"). The wearable electronic device 1102 includes a processor 190 and a display screen 1104. Components of the processor 190, including the media search engine 130, are integrated in the wearable electronic device 1102. In a specific example, the media search engine 130 is operable to search for user-generated media content, downloaded media content, and / or other media content stored in the on-board memory of the wearable electronic device 1102 or stored at a remote memory (e.g., a mobile device, a gaming system, a computer, a server, or a cloud-based file repository accessible by the wearable electronic device 1102). For example, a user can provide a query (e.g., Figure 1 query 120) by entering text via the display screen 1104, by providing speech describing a target sound via the microphone 112, or by capturing the target sound via the microphone 112. In this example, the media search engine 130 performs one or more of the search operations described with reference to Figures 1 to 5 to generate search results based on the query. The search results can be provided to the user on the display screen 1104 or as a sound output to the speaker 118.

[0117] Figure 12 Illustrates a specific implementation 1200 in which device 102 includes a portable electronic device corresponding to augmented reality or mixed reality glasses 1202. The glasses 1202 include a holographic projection unit 1204 configured to project visual data onto the surface of the lens 1206 or reflect the visual data from the surface of the lens 1206 and reflect it onto the retina of the wearer. Components of the processor 190, including the media search engine 130, are integrated in the glasses 1202. In a specific example, the media search engine 130 is operable to search for user-generated media content, downloaded media content, and / or other media content stored in the on-board memory of the glasses 1202 or stored at a remote memory (e.g., a mobile device, a gaming system, a computer, a server, or a cloud-based file repository accessible by the glasses 1202). For example, a user can provide a query (e.g., Figure 1 query 120) by providing speech describing a target sound via the microphone 112 or by capturing the target sound via the microphone 112. In this example, the media search engine 130 performs one or more of the search operations described with reference to Figures 1 to 5 to generate search results based on the query. The search results can be provided to the user via projection onto the lens 1206 or as a sound output to the speaker.

[0118] Figure 13Illustrates a specific implementation 1300 in which the device 102 includes a portable electronic device corresponding to a pair of earbuds 1306, the pair of earbuds including a first earbud 1302 and a second earbud 1304. Although earbuds are described, it should be understood that the disclosed technology can be applied to other in-ear or over-ear playback devices.

[0119] The first earbud 1302 includes a microphone 112, which Figure 13 may include a high signal-to-noise ratio microphone positioned to capture the voice of the wearer of the first earbud 1302; an array of one or more other microphones configured to detect ambient sounds and spatially distributed to support beamforming, illustrated as microphones 1322A, 1322B, and 1322C; an "inner" microphone 1324 near the wearer's ear canal (e.g., to assist with active noise cancellation); and a bone conduction microphone 1326, such as configured to convert sound vibrations of the wearer's earbone or skull into an audio signal. The second earbud 1304 may be configured in a substantially similar manner to the first earbud 1302.

[0120] In Figure 13 which, components of the processor 190 (including the media search engine 130) are integrated into one or both of the earbuds 1306. In a particular example, the media search engine 130 is operable to search for user-generated media content, downloaded media content, and / or other media content stored in the on-board memory of the earbuds 1306 or stored at a remote memory (e.g., a mobile device, a gaming system, a computer, a server, or a cloud-based file repository accessible by the earbuds 1306). For example, a user may provide a query (e.g., Figure 1 query 120) by providing speech describing a target sound via the microphone 112, or by capturing the target sound via the microphone 112 or via one or more of the microphones 1322. In this example, the media search engine 130 performs one or more of the search operations described in reference Figures 1 to 5 to generate search results based on the query. The search results may be provided to the user as a sound output to the speaker 118.

[0121] Figure 14 Is a specific implementation 1400 in which the device 102 includes a wireless speaker and a voice-activated device 1402. The wireless speaker and the voice-activated device 1402 may have a wireless network connection and be configured to perform auxiliary operations. Figure 14The wireless speaker and voice activation device 1402 includes a processor 190, which includes a media search engine 130. Additionally, the wireless speaker and voice activation device 1402 includes a microphone 112 and a speaker 118. During operation, in response to receiving a query (e.g., Figure 1 query 120), the media search engine 130 searches for user-generated media content, downloaded media content, and / or other media content stored in the on-vehicle memory of the wireless speaker and voice activation device 1402 or stored at a remote memory (such as a mobile device, a gaming system, a computer, a server, or a cloud-based file repository accessible by the wireless speaker and voice activation device 1402). For example, a user may provide a query (e.g., Figure 1 query 120) by providing speech describing a target sound via the microphone 112 or by capturing the target sound via the microphone 112. In this example, the media search engine 130 performs one or more of the search operations described in reference Figures 1 to 5 to generate search results based on the query. The search results may be provided to the user as an audio output to the speaker 118.

[0122] Figure 15 FIG. depicts a specific implementation 1500 in which the device 102 is integrated into or includes a portable electronic device corresponding to the camera 1502. In Figure 15 , the camera 1502 includes a microphone 112. Additionally, components of the processor 190, including the media search engine 130, may be integrated into the camera 1502. In a specific example, the media search engine 130 is operable to search for user-generated media content, downloaded media content, and / or other media content stored in the on-vehicle memory of the camera 1502 or stored at a remote memory (e.g., at a mobile device, a gaming system, a computer, a server, or a cloud-based file repository accessible by the camera 1502). For example, a user may provide a query (e.g., Figure 1 query 120) by providing speech describing a target sound via the microphone 112 or by capturing the target sound via the microphone 112. In this example, the media search engine 130 performs one or more of the search operations described in reference Figures 1 to 5 to generate search results based on the query. The search results may be provided to the user via an observation screen (e.g., arranged on the back of the camera 1502).

[0123] Figure 16Depicts a specific implementation 1600 in which device 102 includes a portable electronic device corresponding to an extended reality headset 1602 (e.g., a virtual reality headset, a mixed reality headset, an augmented reality headset, or a combination thereof). The extended reality headset 1602 includes a microphone 112 and a speaker 118. In certain aspects, the visual interface device 1604 is positioned in front of the user's eyes to enable the display of augmented reality, mixed reality, or virtual reality images or scenes to the user when wearing the extended reality headset 1602. In a particular example, the visual interface device 1604 is configured to display a notification that indicates a user voice detected in the audio signal from the microphone 112. In a particular implementation, components of the processor 190, including the media search engine 130, are integrated in the extended reality headset 1602. In a particular example, the media search engine 130 is operable to search user-generated media content, downloaded media content, and / or other media content stored in the on-board memory of the extended reality headset 1602 or stored at a remote memory (e.g., a mobile device, a gaming system, a computer, a server, or a cloud-based file repository accessible by the extended reality headset 1602). For example, a user may provide a query (e.g., Figure 1 query 120) by providing speech via the microphone 112 that describes a target sound or by capturing the target sound via the microphone 112. In this example, the media search engine 130 performs one or more of the search operations described in reference Figures 1 to 5 to generate search results based on the query. The search results may be provided to the user via the visual interface device 1604 or as a sound output to the speaker 118.

[0124] Figure 17 Depicts a specific implementation 1700 in which device 102 corresponds to a vehicle 1702 (illustrated as a manned or unmanned aerial device (e.g., a package delivery drone)) or is integrated within the vehicle. The microphone 112 and the speaker 118 are integrated into the vehicle 1702. The vehicle 1702 may also include one or more cameras 1704. In a particular implementation, components of the processor 190, such as the media search engine 130, are also integrated in the vehicle 1702. During operation, the microphone 112 may generate an input media stream that may be stored as a media file in a file repository (e.g., Figure 1in a media file 152 of the file repository 150) or as a media search query (e.g., query 120). For example, the audio captured by the microphone 112 can be provided as the target sound of the query to determine whether the audio matches a previously recorded sound (such as a gunshot sound, a siren sound, a car collision sound, etc.). In this example, the media search engine 130 performs one or more of the search operations described in reference Figures 1 to 5 to generate search results based on the query. The search results can be used to determine subsequent actions performed by the vehicle 1702, such as reporting the detection of a specific sound to a monitoring station.

[0125] Figure 18 depicts another specific implementation 1800 in which the device 102 corresponds to a vehicle 1802 (illustrated as a car) or is integrated within the vehicle. The vehicle 1802 includes a processor 190 that includes a media search engine 130. The vehicle 1802 also includes a microphone 112, a speaker 118, and a display device 116. The microphone 112 is positioned to capture the words of the operator of the vehicle 1802 or the passengers of the vehicle 1802. During operation, the user can provide a query via the display device 116 of the microphone 112 to initiate a search for media content. In response to receiving the query (e.g., Figure 1 query 120), the media search engine 130 searches for user-generated media content, downloaded media content, and / or other media content stored in the on-vehicle memory of the vehicle 1802 or in another memory accessible by the vehicle 1802 (such as the memory of a mobile device in the vehicle 1802 or the memory at a server or cloud-based file repository accessible by the vehicle 1802). In this example, the media search engine 130 performs one or more of the search operations described in reference Figures 1 to 5 to generate search results based on the query. The search results can be provided to the user via the display device 116 or as a sound output to the speaker 118.

[0126] Reference Figure 19 , shows a specific implementation of a method 1900 for searching for sounds in media files. In a particular aspect, one or more operations of the method 1900 are performed by at least one of the media search engine 130, the processor 190, the device 102, Figure 1 system 100 of, or a combination thereof.

[0127] The method 1900 enables searching media files for a specific sound (e.g., Figures 1 to 5 media file 152). The method 1900 can be performed by an executing media search engine (e.g., Figures 1 to 5one or more processors (e.g., of the media search engine 130) initiate, execute, or control. Media files can include user-generated media files, downloaded media files, or media files accessed in some other way. Additionally, media files can include any file that stores audio content, such as audio files, video files, virtual reality files, or combinations thereof. Figure 1 The processor 190) of.

[0128] Method 1900 includes generating one or more query subtitle embeddings at block 1902 based on a query. For example, Figures 1 to 5 the voice captioning engine 146) of is operable to generate a query subtitle embedding 210 based on query 120. Query 120 can include text (e.g., as in Figure 2 and Figure 3 the text query 202) or audio (e.g., as in Figure 4 the audio query 402). Query 120 can include, for example, a sequence of natural language words that describe a non-speech sound, or an audio sample that represents a sound.

[0129] In some particular implementations, query 120 includes a first set of words that describe a target sound (e.g., Figure 3 or Figure 5 the target sound description 304). In some such particular implementations, the query can also include a second set of words that describe the context (e.g., Figure 3 the context item 306). In such particular implementations, one or more query subtitle embeddings can be determined based on the first set of words (e.g., the target sound description 304), and filtering criteria for selecting a set of embeddings to search can be determined based on the second set of words (e.g., the context item 306). For example, each media file (or at least one subset of the media files) in the media files can be associated with file metadata that indicates the context associated with the media file and an embedding that represents the sound of the media file. In this example, a set of embeddings associated with the media files to be compared with the query subtitle embedding 210 of query 120 can be selected based on the file metadata and the filtering criteria based on the second set of words (e.g., the context item 306). Examples of context information that can be stored as part of the file metadata include, but are not limited to, a timestamp associated with the media file, a location associated with the media file, a file type associated with the media file, a source of the media file, non-audio content of the media file (e.g., a person present in an image of the media file), or combinations thereof.

[0130] In some specific implementations, the query may include audio data representing the sound to be searched (such as different from the description of the sound). For example, the user may capture (using a microphone) the audio data representing the sound, and the audio data representing the sound can be used as an audio query. In such specific implementations, method 1900 may include generating one or more query sound captions based on the query audio data. In such specific implementations, one or more query captions of the query are embedded based on one or more query sound captions. For example, Figure 4 the audio embedding generator 404 can generate a query audio embedding 406 based on the audio query 402, and the tag embedding generator 408 can generate a query tag embedding 410 based on the query audio embedding 406. In this example, the query caption embedding 210 is based on the query tag embedding 410.

[0131] Method 1900 further includes selecting, at block 1904, one or more caption embeddings from the set of embeddings associated with the set of media files in the file repository. Each caption embedding represents a corresponding sound caption, and each sound caption includes a natural language text description of the sound.

[0132] In a particular aspect, the one or more caption embeddings are selected based on a similarity metric indicating the similarity between the one or more caption embeddings and the one or more query caption embeddings. For example, method 1900 may include determining the value of the similarity metric based on the distance between the caption embedding and the query caption embedding in the embedding space.

[0133] Method 1900 further includes generating, at block 1906, search results identifying one or more first media files in a set of media files, wherein each first media file in the one or more first media files is associated with at least one of the one or more caption embeddings. In some specific implementations, the search results indicate media files including the sound corresponding to the query and the playback time of the sound in the media file. For example, the caption embedding may describe a specific sound associated with a specific media file, and the caption embedding may be associated with a time index indicating the approximate playback time of the specific media file where the specific sound appears. In this example, the search results may include information identifying the media file, the specific sound (such as the sound caption or sound tag describing the specific sound), and the time index associated with the specific sound.

[0134] Although in Figure 19The method 1900 illustrated in the flowchart describes searching for subtitle embeddings of media files based on query subtitle embeddings representing a query, but method 1900 may also include searching for other types of embeddings to generate search results. For example, in some embodiments, method 1900 may also include searching for tag embeddings of media files based on tag embeddings representing a query. In this example, one or more media files in the set of media files are associated with one or more tag embeddings (e.g., Figure 5 tag embedding 514 of Figure 5 sound tag 512) representing one or more sound tags associated with the media file.

[0135] In some such embodiments, method 1900 includes generating one or more query tag embeddings based on a query. For example, the tag embedding generator 408 may generate a query tag embedding 410 based on the query 120. Additionally, in such embodiments, method 1900 may also include selecting one or more tag embeddings from the set of embeddings, where the one or more tag embeddings are selected based on a similarity metric indicating the similarity between the one or more tag embeddings and the one or more query tag embeddings. In such embodiments, the search results further identify one or more media files associated with at least one of the one or more tag embeddings.

[0136] Additionally or alternatively, in some embodiments, one or more media files in the set of media files are associated with one or more audio embeddings of one or more sounds in the media file. In such embodiments, method 1900 may also include generating one or more query audio embeddings based on the audio data of the query and comparing the query audio embeddings with the audio embeddings associated with the media file. Further, in such embodiments, a query subtitle embedding representing the query may be generated based on the audio data of the query. For example, Figure 4 or Figure 5 the audio embedding generator 404 may generate a query audio embedding 406 based on the audio data of the audio query 402. In this example, the tag embedding generator 408 generates a query tag embedding 410 based on the query audio embedding 406, and the subtitle embedding generator 242 generates a query subtitle embedding 210 based on the query tag embedding 410.

[0137] In some such specific implementations, query tag embeddings, query audio embeddings, or both can also be used to search for specific sounds represented in an audio query in a media file. For example, method 1900 can include selecting one or more audio embeddings from the set of embeddings, where the one or more audio embeddings are selected based on a similarity metric indicating the similarity between the one or more audio embeddings and the query audio embedding. In this example, the search results further identify one or more media files associated with at least one of the one or more tag embeddings.

[0138] In some specific implementations, method 1900 also includes sorting the search results based on a rank associated with each search result. For example, the rank of each search result can be based on the value of the similarity metric (e.g., a search result that is more similar to the query can be assigned a higher rank value in the search results). When method 1900 includes searching based on multiple types of embeddings (caption / sentence embeddings, tag embeddings, and / or audio embeddings), the similarity metrics associated with the different types of embeddings can be weighted to assign a rank for sorting the search results. By way of illustration, a first set of media files for the search results can be identified based on comparing a query caption embedding based on the query with the caption embeddings of the media files, a second set of media files for the search results can be identified based on comparing a query audio embedding based on the query with the audio embeddings of the media files, and a third set of media files for the search results can be identified based on comparing a query tag embedding based on the query with the tag embeddings of the media files. In this illustrative example, the weighting of the similarity metric associated with the first set of media files is different from the weighting of the similarity metric associated with the second set of media files, different from the weighting of the similarity metric associated with the third set of media files, or both of the above.

[0139] In some specific implementations, method 1900 can also include the operation of adding one or more new media files to a file repository. For example, in such specific implementations, method 1900 includes obtaining an additional media file for storage at the file repository and processing the additional media file to detect one or more sounds represented in the additional media file. In this example, method 1900 also includes generating one or more embeddings (e.g., audio embeddings, tag embeddings, caption embeddings, or a combination thereof) associated with the one or more sounds detected in the additional media file and storing the additional media file and the one or more embeddings in the file repository. In this example, in response to receiving a subsequent query, method 1900 includes searching for the one or more embeddings associated with the additional media file.

[0140] Figure 19The method 1900 can be implemented by a field programmable gate array (FPGA) device, an application specific integrated circuit (ASIC), a processing unit (such as a central processing unit (CPU)), a DSP, a controller, another hardware device, a firmware device, or any combination thereof. For example, Figure 19 the method 1900 can be executed by a processor that executes instructions, such as those described with reference to Figure 20 .

[0141] With reference to Figure 20 , a block diagram depicting a particular illustrative embodiment of a device is shown and generally designated as 2000. In various embodiments, the device 2000 may have more or fewer components than Figure 20 shown. In an illustrative embodiment, the device 2000 may correspond to the device 102. In an illustrative embodiment, the device 2000 may perform one or more operations described with reference to Figures 1 to 19 .

[0142] In a particular embodiment, the device 2000 includes a processor 2006 (e.g., a central processing unit (CPU)). The device 2000 may include one or more additional processors 2010 (e.g., one or more DSPs). In a particular aspect, Figure 1 the processor 190 corresponds to the processor 2006, the processor 2010, or a combination thereof. The processor 2010 may include a voice and music codec 2008, which includes a voice codec (“vocoder”) encoder 2036, a vocoder decoder 2038, a media search engine 130, or a combination thereof.

[0143] The device 2000 may include a memory 192 and a codec 2034. The memory 192 may include instructions 2056 that can be executed by one or more additional processors 2010 (or the processor 2006) to implement the functionality described with reference to the media search engine 130. In the Figure 20 example shown, the memory 192 also includes a file repository 150.

[0144] In Figure 20 , the device 2000 includes a modem 2070 coupled to an antenna 2052 via a transceiver 2050. The modem 2070, the transceiver 2050, and the antenna 2052 are operable to receive an input media stream, transmit an output media stream, or both. For example, the device 2000 may receive a media file to be stored in the file repository 150 via the modem 2070, may receive a query via the modem 2070, or may transmit search results to another device via the modem 2070.

[0145] Device 2000 may include a display device 116 coupled to a display controller 2026. A speaker 118 and a microphone 112 may be coupled to a CODEC 2034. The CODEC 2034 may include a digital-to-analog converter (DAC) 2002, an analog-to-digital converter (ADC) 2004, or both. In a particular implementation, the CODEC 2034 may receive an analog signal from the microphone 112, use the analog-to-digital converter 2004 to convert the analog signal into a digital signal, and provide the digital signal to a voice and music codec 2008. The voice and music codec 2008 may process the digital signal, and the digital signal may be further processed by a media search engine 130. In a particular implementation, the voice and music codec 2008 may provide the digital signal to the CODEC 2034. The CODEC 2034 may use the digital-to-analog converter 2002 to convert the digital signal into an analog signal, and may provide the analog signal to the speaker 118.

[0146] In a particular implementation, device 2000 may be included in a system-in-package or system-on-chip device 2022. In a particular implementation, a memory 192, a processor 2006, a processor 2010, a display controller 2026, a CODEC 2034, and a modem 2070 are included in the system-in-package or system-on-chip device 2022. In a particular implementation, an input device 2030 and a power supply 2044 are coupled to the system-in-package or system-on-chip device 2022. Additionally, in a particular implementation, as Figure 20 shown, the display device 116, the input device 2030, the speaker 118, the microphone 112, the antenna 2052, and the power supply 2044 are external to the system-in-package or system-on-chip device 2022. In a particular implementation, each of the display device 116, the input device 2030, the speaker 118, the microphone 112, the antenna 2052, and the power supply 2044 may be coupled to a component (such as an interface or a controller) of the system-in-package or system-on-chip device 2022.

[0147] Device 2000 may include a smart speaker, a speaker bar, a mobile communication device, a smartphone, a cellular phone, a laptop computer, a computer, a tablet computer, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, headphones, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and voice-activated device, a portable electronic device, an automobile, a computing device, a communication device, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, a mobile device, or any combination thereof.

[0148] In connection with the specific implementations described, an apparatus includes components for generating one or more query caption embeddings based on a query. For example, the components for generating one or more query caption embeddings may correspond to media search engine 130, voice captioning engine 146, embedding generator 142, caption embedding generator 242, caption embedding generator 688, caption embedding generator 788, processor 190, processor 2006, processor 2010, one or more other circuits or components configured to generate query caption embeddings, or any combination thereof.

[0149] In connection with the specific implementations described, the apparatus also includes components for selecting one or more caption embeddings from a set of embeddings associated with a set of media files in a file repository, where each caption embedding represents a corresponding voice caption and each voice caption includes a natural language text description of the voice, and where the one or more caption embeddings are selected based on a similarity metric indicating the similarity between the one or more caption embeddings and the one or more query caption embeddings. For example, the components for selecting one or more caption embeddings may correspond to media search engine 130, comparator 140, processor 190, processor 2006, processor 2010, one or more other circuits or components configured to select caption embeddings, or any combination thereof.

[0150] In connection with the specific implementations described, the apparatus also includes components for generating search results identifying one or more first media files in the set of media files, where each first media file in the one or more first media files is associated with at least one of the one or more caption embeddings. For example, the components for generating search results may correspond to media search engine 130, comparator 140, processor 190, processor 2006, processor 2010, one or more other circuits or components configured to generate search results, or any combination thereof.

[0151] In some specific implementations, a non-transitory computer-readable medium (e.g., a computer-readable storage device such as memory 192) includes instructions (e.g., instructions 2056) that, when executed by one or more processors (e.g., one or more processors 190, one or more processors 2010, or processor 2006), cause the one or more processors to generate one or more query subtitle embeddings based on a query. The instructions are further executable by the one or more processors to select one or more subtitle embeddings from a set of embeddings associated with a set of media files in a file repository, where each subtitle embedding represents a corresponding audio subtitle and each audio subtitle includes a natural language text description of the audio, and where the one or more subtitle embeddings are selected based on a similarity metric indicating the similarity between the one or more subtitle embeddings and the one or more query subtitle embeddings. The instructions are also executable by the one or more processors to generate search results identifying one or more first media files in the set of media files, where each first media file in the one or more first media files is associated with at least one of the one or more subtitle embeddings.

[0152] Specific aspects of the present disclosure are described below in groups of related embodiments:

[0153] According to Embodiment 1, a device includes one or more processors configured to: generate one or more query subtitle embeddings based on a query; select one or more subtitle embeddings from a set of embeddings associated with a set of media files in a file repository, where each subtitle embedding represents a corresponding audio subtitle and each audio subtitle includes a natural language text description of the audio, where the one or more subtitle embeddings are selected based on a first similarity metric indicating the similarity between the one or more subtitle embeddings and the one or more query subtitle embeddings; and generate search results identifying one or more first media files in the set of media files, where each first media file in the one or more first media files is associated with at least one of the one or more subtitle embeddings.

[0154] Embodiment 2 includes the device according to Embodiment 1, where the query includes a sequence of natural language words describing a non-speech sound.

[0155] Embodiment 3 includes the device according to Embodiment 1 or Embodiment 2, where the query includes a first set of words describing a target sound and a second set of words describing a context, and where the one or more processors are configured to determine the one or more query subtitle embeddings based on the first set of words.

[0156] Example 4 includes the apparatus according to Example 3, wherein each media file of at least one subset of the set of media files is associated with file metadata indicating the context associated with the media file, and wherein the one or more processors are configured to select the set of embeddings from which to select the one or more subtitle embeddings based on the second set of words of the query and the file metadata.

[0157] Example 5 includes the apparatus according to Example 4, wherein the file metadata of a particular media file indicates a timestamp associated with the media file, a location associated with the media file, or both.

[0158] Example 6 includes the apparatus according to any one of Examples 1 to 5, wherein a particular subtitle embedding describes a particular sound associated with a particular media file, and wherein the particular subtitle embedding is associated with a time index indicating an approximate playback time of the particular media file at which the particular sound occurs.

[0159] Example 7 includes the apparatus according to any one of Examples 1 to 6, wherein the search results further indicate, for a particular media file, a time index associated with a particular sound.

[0160] Example 8 includes the apparatus according to any one of Examples 1 to 7, wherein the set of media files includes one or more audio files, one or more video files, one or more virtual reality files, or a combination thereof.

[0161] Example 9 includes the apparatus according to any one of Examples 1 to 8, wherein the query includes query audio data, and wherein the one or more query subtitle embeddings are based on the query audio data.

[0162] Example 10 includes the apparatus according to any one of Examples 1 to 9, wherein a particular media file in the set of media files is further associated with one or more audio embeddings of one or more sounds in the particular media file, and wherein the set of embeddings associated with the set of media files includes the one or more audio embeddings.

[0163] Example 11 includes the apparatus according to Example 10, wherein the one or more query captions are embedded in query audio data based on the query, and the one or more processors are further configured to: generate a query audio embedding based on the query audio data; and select one or more audio embeddings from the set of embeddings, wherein the one or more audio embeddings are selected based on a second similarity metric indicative of a similarity between the one or more audio embeddings and the query audio embedding, wherein the search results further identify one or more second media files in the set of media files, each of the one or more second media files being associated with at least one of the one or more audio embeddings.

[0164] Example 12 includes the apparatus according to Example 11, wherein the one or more processors are further configured to rank the search results based on a similarity value, and wherein the value of the first similarity metric associated with the one or more first media files and the value of the second similarity metric associated with the one or more second media files are weighted differently to rank the search results.

[0165] Example 13 includes the apparatus according to any one of Examples 1 to 12, wherein a particular media file in the set of media files is further associated with one or more tag embeddings representing one or more sound tags associated with the particular media file, and wherein the set of embeddings associated with the set of media files includes the one or more tag embeddings.

[0166] Example 14 includes the apparatus according to Example 13, wherein the one or more processors are further configured to: generate one or more query tag embeddings based on the query; and select one or more tag embeddings from the set of embeddings, wherein the one or more tag embeddings are selected based on a third similarity metric indicative of a similarity between the one or more tag embeddings and the one or more query tag embeddings, wherein the search results further identify one or more third media files in the set of media files, each of the one or more third media files being associated with at least one of the one or more tag embeddings.

[0167] Example 15 includes the apparatus according to Example 14, wherein the one or more processors are further configured to rank the search results based on a similarity value, and wherein the value of the first similarity metric associated with the one or more first media files and the value of the third similarity metric associated with the one or more third media files are weighted differently to rank the search results.

[0168] Example 16 includes the apparatus according to any one of Examples 1 to 15, wherein the one or more processors are further configured to: obtain additional media files for storage at the file repository; process the additional media files to detect one or more sounds represented in the additional media files; generate one or more embeddings associated with the one or more sounds detected in the additional media files; store the additional media files and the one or more embeddings in the file repository; and in response to receiving a subsequent query, search for the one or more embeddings associated with the additional media files.

[0169] Example 17 includes the apparatus according to Example 16, wherein generating the one or more embeddings includes generating caption embeddings associated with the one or more sounds.

[0170] Example 18 includes the apparatus according to Example 16, wherein in order to generate the one or more embeddings associated with the one or more sounds represented in the additional media files, the one or more processors are configured to generate audio embeddings representing specific sounds detected in the additional media files.

[0171] Example 19 includes the apparatus according to any one of Examples 1 to 18, wherein the one or more processors are further configured to determine the similarity metric based on the distance between the one or more caption embeddings and the one or more query caption embeddings in the embedding space.

[0172] According to Example 20, a method includes: generating, by one or more processors, one or more query caption embeddings based on a query; selecting, by the one or more processors, one or more caption embeddings from a set of embeddings associated with a set of media files of a file repository, wherein each caption embedding represents a corresponding sound caption and each sound caption includes a natural language text description of a sound, and wherein the one or more caption embeddings are selected based on a first similarity metric indicating a similarity between the one or more caption embeddings and the one or more query caption embeddings; and generating, by the one or more processors, search results identifying one or more first media files of the set of media files, wherein each first media file of the one or more first media files is associated with at least one of the one or more caption embeddings.

[0173] Example 21 includes the method according to Example 20, wherein the query includes a sequence of natural language words describing a non-speech sound.

[0174] Example 22 includes the method according to Example 20 or Example 21, wherein the query includes a first set of words describing a target sound and a second set of words describing a context, and the method further includes determining the one or more query subtitle embeddings based on the first set of words.

[0175] Example 23 includes the method according to Example 22, wherein each media file of at least one subset of the set of media files is associated with file metadata indicating the context associated with the media file, and the method further includes selecting the set of embeddings from which to select the one or more subtitle embeddings based on the second set of words of the query and the file metadata.

[0176] Example 24 includes the method according to Example 23, wherein the file metadata of a particular media file indicates a timestamp associated with the media file, a location associated with the media file, or both.

[0177] Example 25 includes the method according to any one of Examples 20 to 24, wherein a particular subtitle embedding describes a particular sound associated with a particular media file, and wherein the particular subtitle embedding is associated with a time index indicating the approximate playback time of the particular media file at which the particular sound occurs.

[0178] Example 26 includes the method according to any one of Examples 20 to 25, wherein the search results further indicate, for a particular media file, a time index associated with a particular sound.

[0179] Example 27 includes the method according to any one of Examples 20 to 26, wherein the set of media files includes one or more audio files, one or more video files, one or more virtual reality files, or a combination thereof.

[0180] Example 28 includes the method according to any one of Examples 20 to 27, wherein the query includes query audio data, and wherein the one or more query subtitle embeddings are based on the query audio data.

[0181] Example 29 includes the method according to any one of Examples 20 to 28, wherein a particular media file in the set of media files is further associated with one or more audio embeddings of one or more sounds in the particular media file, and wherein the set of embeddings associated with the set of media files includes the one or more audio embeddings.

[0182] Example 30 includes the method according to Example 29, wherein the one or more query captions are embedded based on query audio data of the query, and the method further includes: generating a query audio embedding based on the query audio data; and selecting one or more audio embeddings from the set of embeddings, wherein the one or more audio embeddings are selected based on a second similarity metric indicating a similarity between the one or more audio embeddings and the query audio embedding, wherein the search results further identify one or more second media files in the set of media files, and each of the one or more second media files is associated with at least one of the one or more audio embeddings.

[0183] Example 31 includes the method according to Example 30, further including ranking the search results based on a similarity value, and wherein the value of the first similarity metric associated with the one or more first media files and the value of the second similarity metric of the one or more second media files are differently weighted to rank the search results.

[0184] Example 32 includes the method according to any one of Examples 20 to 31, wherein a particular media file in the set of media files is further associated with one or more tag embeddings representing one or more sound tags associated with the particular media file, and wherein the set of embeddings associated with the set of media files includes the one or more tag embeddings.

[0185] Example 33 includes the method according to Example 32, further including: generating one or more query tag embeddings based on the query; and selecting one or more tag embeddings from the set of embeddings, wherein the one or more tag embeddings are selected based on a third similarity metric indicating a similarity between the one or more tag embeddings and the one or more query tag embeddings, wherein the search results further identify one or more third media files in the set of media files, and each of the one or more third media files is associated with at least one of the one or more tag embeddings.

[0186] Example 34 includes the method according to Example 33, further including ranking the search results based on a similarity value, and wherein the value of the first similarity metric associated with the one or more first media files and the value of the third similarity metric of the one or more third media files are differently weighted to rank the search results.

[0187] Example 35 includes the method according to any one of Examples 20 to 34, further comprising: obtaining additional media files for storage at the file repository; processing the additional media files to detect one or more sounds represented in the additional media files; generating one or more embeddings associated with the one or more sounds detected in the additional media files; storing the additional media files and the one or more embeddings in the file repository; and in response to receiving a subsequent query, searching for the one or more embeddings associated with the additional media files.

[0188] Example 36 includes the method according to Example 35, wherein generating the one or more embeddings includes generating caption embeddings associated with the one or more sounds.

[0189] Example 37 includes the method according to Example 35, wherein generating the one or more embeddings associated with the one or more sounds represented in the additional media files includes generating audio embeddings representing specific sounds detected in the additional media files.

[0190] Example 38 includes the method according to any one of Examples 20 to 37, further comprising determining the similarity metric based on the distance between the one or more caption embeddings and the one or more query caption embeddings in the embedding space.

[0191] According to Example 39, a non-transitory computer-readable storage device stores instructions that can be executed by one or more processors to cause the one or more processors to: generate one or more query caption embeddings based on a query; select one or more caption embeddings from a set of embeddings associated with a set of media files in a file repository, wherein each caption embedding represents a corresponding sound caption and each sound caption includes a natural language text description of a sound, and wherein the one or more caption embeddings are selected based on a first similarity metric indicating the similarity between the one or more caption embeddings and the one or more query caption embeddings; and generate search results identifying one or more first media files in the set of media files, wherein each first media file in the one or more first media files is associated with at least one of the one or more caption embeddings.

[0192] Example 40 includes the non-transitory computer-readable storage device according to Example 39, wherein the query includes a sequence of natural language words describing a non-speech sound.

[0193] Example 41 includes the non-transitory computer-readable storage device according to Example 39 or Example 40, wherein the query includes a first set of words describing a target sound and a second set of words describing a context, and wherein the instructions are further executable to cause one or more processors to determine the one or more query subtitle embeddings based on the first set of words.

[0194] Example 42 includes the non-transitory computer-readable storage device according to Example 41, wherein each media file of at least one subset of the set of media files is associated with file metadata indicating the context associated with the media file, and wherein the instructions are further executable to cause one or more processors to select the set of embeddings from which to select the one or more subtitle embeddings based on the second set of words of the query and the file metadata.

[0195] Example 43 includes the non-transitory computer-readable storage device according to Example 42, wherein the file metadata of a particular media file indicates a timestamp associated with the media file, a location associated with the media file, or both.

[0196] Example 44 includes the non-transitory computer-readable storage device according to any one of Examples 39 to 43, wherein a particular subtitle embedding describes a particular sound associated with a particular media file, and wherein the particular subtitle embedding is associated with a time index indicating an approximate playback time of the particular media file at which the particular sound occurs.

[0197] Example 45 includes the non-transitory computer-readable storage device according to any one of Examples 39 to 44, wherein the search results further indicate, for a particular media file, a time index associated with a particular sound.

[0198] Example 46 includes the non-transitory computer-readable storage device according to any one of Examples 39 to 45, wherein the set of media files includes one or more audio files, one or more video files, one or more virtual reality files, or a combination thereof.

[0199] Example 47 includes the non-transitory computer-readable storage device according to any one of Examples 39 to 46, wherein the query includes query audio data, and wherein the one or more query subtitle embeddings are based on the query audio data.

[0200] Example 48 includes the non-transitory computer-readable storage device according to any one of Examples 39 to 47, wherein a particular media file in the set of media files is further associated with one or more audio embeddings of one or more sounds in the particular media file, and wherein the set of embeddings associated with the set of media files includes the one or more audio embeddings.

[0201] Example 49 includes the non-transitory computer-readable storage device according to Example 48, wherein the one or more query caption embeddings are based on query audio data of the query, and the instructions are further executable to cause one or more processors to: generate a query audio embedding based on the query audio data; and select one or more audio embeddings from the set of embeddings, wherein the one or more audio embeddings are selected based on a second similarity metric indicating a similarity between the one or more audio embeddings and the query audio embedding, wherein the search results further identify one or more second media files in the set of media files, each of the one or more second media files being associated with at least one of the one or more audio embeddings.

[0202] Example 50 includes the non-transitory computer-readable storage device according to Example 49, wherein the instructions are further executable to cause one or more processors to rank the search results based on a similarity value, and wherein the value of the first similarity metric associated with the one or more first media files and the value of the second similarity metric of the one or more second media files are differently weighted to rank the search results.

[0203] Example 51 includes the non-transitory computer-readable storage device according to any one of Examples 39 to 50, wherein a particular media file in the set of media files is further associated with one or more tag embeddings representing one or more sound tags associated with the particular media file, and wherein the set of embeddings associated with the set of media files includes the one or more tag embeddings.

[0204] Example 52 includes the non-transitory computer-readable storage device according to Example 51, wherein the instructions are further executable to cause one or more processors to: generate one or more query tag embeddings based on the query; and select one or more tag embeddings from the set of embeddings, wherein the one or more tag embeddings are selected based on a third similarity measure indicative of a similarity between the one or more tag embeddings and the one or more query tag embeddings, wherein the search results further identify one or more third media files in the set of media files, and each third media file of the one or more third media files is associated with at least one of the one or more tag embeddings.

[0205] Example 53 includes the non-transitory computer-readable storage device according to Example 52, wherein the instructions are further executable to cause one or more processors to rank the search results based on a similarity value, and wherein the value of the first similarity measure associated with the one or more first media files and the value of the third similarity measure of the one or more third media files are differently weighted to rank the search results.

[0206] Example 54 includes the non-transitory computer-readable storage device according to any one of Examples 39 to 53, wherein the instructions are further executable to cause one or more processors to: obtain additional media files for storage at the file repository; process the additional media files to detect one or more sounds represented in the additional media files; generate one or more embeddings associated with the one or more sounds detected in the additional media files; store the additional media files and the one or more embeddings in the file repository; and in response to receiving a subsequent query, search for the one or more embeddings associated with the additional media files.

[0207] Example 55 includes the non-transitory computer-readable storage device according to Example 54, wherein generating the one or more embeddings includes generating caption embeddings associated with the one or more sounds.

[0208] Example 56 includes the non-transitory computer-readable storage device according to Example 54, wherein to generate the one or more embeddings associated with the one or more sounds represented in the additional media files, the instructions are executable to cause one or more processors to generate audio embeddings representing a particular sound detected in the additional media files.

[0209] Example 57 includes the non-transitory computer-readable storage device according to any one of Examples 39 to 56, wherein the instructions are further executable to cause one or more processors to determine the similarity metric based on a distance between the one or more caption embeddings and the one or more query caption embeddings in an embedding space.

[0210] According to Example 58, a device includes: means for generating one or more query caption embeddings based on a query; means for selecting one or more caption embeddings from a set of embeddings associated with a set of media files in a file repository, wherein each caption embedding represents a corresponding audio caption and each audio caption includes a natural language text description of the audio, wherein the one or more caption embeddings are selected based on a first similarity metric indicating a similarity between the one or more caption embeddings and the one or more query caption embeddings; and means for generating a search result identifying one or more first media files in the set of media files, each first media file in the one or more first media files being associated with at least one of the one or more caption embeddings.

[0211] Example 59 includes the device according to Example 58, wherein the query includes a sequence of natural language words describing a non-speech sound.

[0212] Example 60 includes the device according to Example 58 or Example 59, wherein the query includes a first set of words describing a target sound and a second set of words describing a context, and the device further includes means for determining the one or more query caption embeddings based on the first set of words.

[0213] Example 61 includes the device according to Example 60, wherein each media file in at least one subset of the set of media files is associated with file metadata indicating a context associated with the media file, and the device further includes means for selecting the set of embeddings from which the one or more caption embeddings are selected based on the second set of words of the query and the file metadata.

[0214] Example 62 includes the device according to Example 61, wherein the file metadata of a particular media file indicates a timestamp associated with the media file, a location associated with the media file, or both.

[0215] Example 63 includes the device according to any one of Examples 58 to 62, wherein a particular caption embedding describes a particular sound associated with a particular media file, and wherein the particular caption embedding is associated with a time index indicating an approximate playback time of the particular media file at which the particular sound occurs.

[0216] Embodiment 64 includes the apparatus according to any one of Embodiments 58 to 63, wherein the search results further indicate a time index associated with a specific sound for a specific media file.

[0217] Embodiment 65 includes the apparatus according to any one of Embodiments 58 to 64, wherein the set of media files includes one or more audio files, one or more video files, one or more virtual reality files, or a combination thereof.

[0218] Embodiment 66 includes the apparatus according to any one of Embodiments 58 to 65, wherein the query includes querying audio data, and wherein the one or more query captions are embedded based on the query audio data.

[0219] Embodiment 67 includes the apparatus according to any one of Embodiments 58 to 66, wherein a specific media file in the set of media files is further associated with one or more audio embeddings of one or more sounds in the specific media file, and wherein the set of embeddings associated with the set of media files includes the one or more audio embeddings.

[0220] Embodiment 68 includes the apparatus according to Embodiment 67, wherein the one or more query captions are embedded based on the query audio data of the query, and the apparatus further includes: components for generating query audio embeddings based on the query audio data; and components for selecting one or more audio embeddings from the set of embeddings, wherein the one or more audio embeddings are selected based on a second similarity metric indicating a similarity between the one or more audio embeddings and the query audio embeddings, wherein the search results further identify one or more second media files in the set of media files, and each of the one or more second media files is associated with at least one of the one or more audio embeddings.

[0221] Embodiment 69 includes the apparatus according to Embodiment 68, further including components for ranking the search results based on similarity values, and wherein the values of the first similarity metric associated with the one or more first media files and the values of the second similarity metric of the one or more second media files are weighted differently to rank the search results.

[0222] Embodiment 70 includes the apparatus according to any one of Embodiments 58 to 69, wherein a specific media file in the set of media files is further associated with one or more tag embeddings representing one or more sound tags associated with the specific media file, and wherein the set of embeddings associated with the set of media files includes the one or more tag embeddings.

[0223] Embodiment 71 includes the apparatus according to Embodiment 70, and further includes: a component for generating one or more query tag embeddings based on the query; and a component for selecting one or more tag embeddings from the set of embeddings, wherein the one or more tag embeddings are selected based on a third similarity metric indicating the similarity between the one or more tag embeddings and the one or more query tag embeddings, wherein the search results further identify one or more third media files in the set of media files, and each of the one or more third media files is associated with at least one of the one or more tag embeddings.

[0224] Embodiment 72 includes the apparatus according to Embodiment 71, and further includes a component for ranking the search results based on a similarity value, and wherein the value of the first similarity metric associated with the one or more first media files and the value of the third similarity metric of the one or more third media files are weighted differently to rank the search results.

[0225] Embodiment 73 includes the apparatus according to any one of Embodiments 58 to 72, and further includes: a component for obtaining additional media files to be stored at the file repository; a component for processing the additional media files to detect one or more sounds represented in the additional media files; a component for generating one or more embeddings associated with the one or more sounds detected in the additional media files; a component for storing the additional media files and the one or more embeddings in the file repository; and a component for searching for the one or more embeddings associated with the additional media files in response to receiving a subsequent query.

[0226] Embodiment 74 includes the apparatus according to Embodiment 73, wherein generating the one or more embeddings includes generating caption embeddings associated with the one or more sounds.

[0227] Embodiment 75 includes the apparatus according to Embodiment 73, wherein generating the one or more embeddings associated with the one or more sounds represented in the additional media files includes generating audio embeddings representing specific sounds detected in the additional media files.

[0228] Embodiment 76 includes the apparatus according to any one of Embodiments 58 to 75, and further includes determining the similarity metric based on the distance between the one or more caption embeddings and the one or more query caption embeddings in the embedding space.

[0229] Those skilled in the art will also appreciate that the various illustrative logical blocks, configurations, modules, circuits, and algorithm steps described in connection with the specific implementations disclosed herein can be implemented as electronic hardware, computer software executed by a processor, or a combination of both. The various illustrative components, blocks, configurations, modules, circuits, and steps have been described generally above in terms of their functionality. Whether such functionality is implemented as hardware or processor-executable instructions depends upon the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in varying ways for each particular application, and such implementation decisions will not be interpreted as causing a departure from the scope of the present disclosure.

[0230] The steps of a method or algorithm described in connection with the specific implementations disclosed herein can be embodied directly in hardware, in a software module executed by a processor, or in a combination of both. A software module may reside in random access memory (RAM), flash memory, read only memory (ROM), programmable read only memory (PROM), erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), registers, a hard disk, a removable disk, a compact disc read only memory (CD-ROM), or any other form of non-transitory storage medium known in the art. An exemplary storage medium is coupled to the processor such that the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an application specific integrated circuit (ASIC). The ASIC may reside in a computing device or a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a computing device or a user terminal.

[0231] The foregoing description of the disclosed aspects is provided to enable a person skilled in the art to make or use the disclosed aspects. Various modifications to these aspects will be readily apparent to those skilled in the art, and the principles defined herein may be applied to other aspects without departing from the scope of the disclosure. Thus, the disclosure is not intended to be limited to the aspects shown herein but is to be accorded the widest scope consistent with the principles and novel features defined by the following claims.

Claims

1. A device, the device comprises: one or more processors configured to: generate one or more query subtitle embeddings based on a query; select one or more subtitle embeddings from a set of embeddings associated with a set of media files in a file repository, where each subtitle embedding represents a corresponding audio subtitle and each audio subtitle includes a natural language text description of the audio, and where the one or more subtitle embeddings are selected based on a first similarity metric indicating the similarity between the one or more subtitle embeddings and the one or more query subtitle embeddings; and generate search results identifying one or more first media files in the set of media files, where each first media file in the one or more first media files is associated with at least one of the one or more subtitle embeddings.

2. The device according to claim 1, wherein the query comprises a sequence of natural language words describing a non-speech audio.

3. The device according to claim 1, wherein the query comprises a first set of words describing a target audio and a second set of words describing a context, and wherein the one or more processors are configured to determine the one or more query subtitle embeddings based on the first set of words.

4. The device according to claim 3, wherein each media file in at least one subset of the set of media files is associated with file metadata indicating the context associated with the media file, and wherein the one or more processors are configured to select the set of embeddings from which to select the one or more subtitle embeddings based on the second set of words of the query and the file metadata.

5. The device according to claim 1, wherein a specific subtitle embedding describes a specific audio associated with a specific media file, and wherein the specific subtitle embedding is associated with a time index indicating the approximate playback time of the specific media file at which the specific audio appears.

6. The device according to claim 1, wherein the set of media files comprises one or more audio files, one or more video files, one or more virtual reality files, or a combination thereof.

7. The device according to claim 1, wherein the query comprises query audio data, and wherein the one or more query subtitle embeddings are based on the query audio data.

8. The device according to claim 1, wherein a specific media file in the set of media files is further associated with one or more audio embeddings of one or more audios in the specific media file, and wherein the set of embeddings associated with the set of media files comprises the one or more audio embeddings.

9. The device according to claim 8, wherein the one or more query subtitle embeddings are based on the query audio data of the query, and the one or more processors are further configured to: generate a query audio embedding based on the query audio data; and Select one or more audio embeddings from the set of embeddings, wherein the one or more audio embeddings are selected based on a second similarity metric indicative of a similarity between the one or more audio embeddings and the query audio embedding, wherein the search results further identify one or more second media files in the set of media files, and each second media file in the one or more second media files is associated with at least one of the one or more audio embeddings.

10. The apparatus according to claim 9, wherein the one or more processors are further configured to rank the search results based on similarity values, and wherein the values of the first similarity metric associated with the one or more first media files and the values of the second similarity metric of the one or more second media files are differently weighted to rank the search results.

11. The apparatus according to claim 1, wherein a particular media file in the set of media files is further associated with one or more tag embeddings representing one or more sound tags associated with the particular media file, and wherein the set of embeddings associated with the set of media files includes the one or more tag embeddings.

12. The apparatus according to claim 11, wherein the one or more processors are further configured to: generate one or more query tag embeddings based on the query; and select one or more tag embeddings from the set of embeddings, wherein the one or more tag embeddings are selected based on a third similarity metric indicative of a similarity between the one or more tag embeddings and the one or more query tag embeddings, wherein the search results further identify one or more third media files in the set of media files, and each third media file in the one or more third media files is associated with at least one of the one or more tag embeddings.

13. The apparatus according to claim 12, wherein the one or more processors are further configured to rank the search results based on similarity values, and wherein the values of the first similarity metric associated with the one or more first media files and the values of the third similarity metric of the one or more third media files are differently weighted to rank the search results.

14. The apparatus according to claim 1, wherein the one or more processors are further configured to: obtain additional media files for storage at the file repository; process the additional media files to detect one or more sounds represented in the additional media files; generate one or more embeddings associated with the one or more sounds detected in the additional media files; store the additional media files and the one or more embeddings in the file repository; and in response to receiving a subsequent query, search for the one or more embeddings associated with the additional media files.

15. The apparatus according to claim 1, wherein the one or more processors are further configured to determine the first similarity measure based on a distance between the one or more caption embeddings and the one or more query caption embeddings in the embedding space.

16. A method, the method comprising: generating, by one or more processors, one or more query caption embeddings based on a query; selecting, by the one or more processors, one or more caption embeddings from a set of embeddings associated with a set of media files in a file repository, wherein each caption embedding represents a corresponding audio caption and each audio caption includes a natural language text description of the audio, and wherein the one or more caption embeddings are selected based on a first similarity measure indicative of a similarity between the one or more caption embeddings and the one or more query caption embeddings; and generating, by the one or more processors, search results identifying one or more first media files in the set of media files, wherein each first media file in the one or more first media files is associated with at least one of the one or more caption embeddings.

17. The method according to claim 16, wherein the query comprises query audio data or a sequence of natural language words describing non-speech sounds.

18. The method according to claim 16, wherein the query comprises a first set of words describing a target sound and a second set of words describing a context, and the method further comprises determining the one or more query caption embeddings based on the first set of words.

19. The method according to claim 18, wherein each media file in at least one subset of the set of media files is associated with file metadata indicative of a context associated with the media file, and the method further comprises selecting the set of embeddings from which the one or more caption embeddings are selected based on the second set of words of the query and the file metadata.

20. The method according to claim 16, wherein a particular caption embedding describes a particular sound associated with a particular media file, and wherein the particular caption embedding is associated with a time index indicative of a approximate playback time of the particular media file at which the particular sound occurs.

21. The method according to claim 20, wherein the one or more query caption embeddings are based on query audio data of the query, and the method further comprising: generating a query audio embedding based on the query audio data; and selecting one or more audio embeddings from the set of embeddings, wherein the one or more audio embeddings are selected based on a second similarity measure indicative of a similarity between the one or more audio embeddings and the query audio embedding, and wherein the search results further identify one or more second media files in the set of media files, wherein each second media file in the one or more second media files is associated with at least one of the one or more audio embeddings.

22. The method according to claim 16, the method further comprising: generating one or more query tag embeddings based on the query; and selecting one or more label embeddings from the set of embeddings, wherein the one or more label embeddings are selected based on a third similarity measure indicative of a similarity between the one or more label embeddings and the one or more query label embeddings, wherein the search result further identifies one or more third media files in the set of media files, each of the one or more third media files being associated with at least one of the one or more label embeddings.

23. The method according to claim 16, the method further comprising: obtaining additional media files for storage at the file repository; processing the additional media files to detect one or more sounds represented in the additional media files; generating one or more embeddings associated with the one or more sounds detected in the additional media files; storing the additional media files and the one or more embeddings in the file repository; and in response to receiving a subsequent query, searching for the one or more embeddings associated with the additional media files.

24. A non-transitory computer-readable storage device storing instructions that, when executed by one or more processors, cause the one or more processors to: generate one or more query caption embeddings based on a query; select one or more caption embeddings from a set of embeddings associated with a set of media files in a file repository, wherein each caption embedding represents a corresponding audio caption and each audio caption includes a natural language text description of a sound, wherein the one or more caption embeddings are selected based on a first similarity measure indicative of a similarity between the one or more caption embeddings and the one or more query caption embeddings; and generate a search result identifying one or more first media files in the set of media files, each of the one or more first media files being associated with at least one of the one or more caption embeddings.

25. The non-transitory computer-readable storage device according to claim 24, wherein the query includes a first set of words describing a target sound and a second set of words describing a context, and wherein the instructions are further executable to cause one or more processors to determine the one or more query caption embeddings based on the first set of words.

26. The non-transitory computer-readable storage device according to claim 24, wherein the instructions are further executable to cause the one or more processors to: obtain additional media files for storage at the file repository; process the additional media files to detect one or more sounds represented in the additional media files; generate one or more embeddings associated with the one or more sounds detected in the additional media files; store the additional media files and the one or more embeddings in the file repository; and in response to receiving a subsequent query, search for the one or more embeddings associated with the additional media files.

27. The non-transitory computer-readable storage device according to claim 26, wherein generating the one or more embeddings includes generating caption embeddings associated with the one or more sounds.

28. The non-transitory computer-readable storage device according to claim 24, wherein the instructions are further executable to cause one or more processors to determine the first similarity metric based on a distance between the one or more caption embeddings and the one or more query caption embeddings in an embedding space.

29. An apparatus, the apparatus comprising: means for generating one or more query caption embeddings based on a query; means for selecting one or more caption embeddings from a set of embeddings associated with a set of media files in a file repository, wherein each caption embedding represents a corresponding audio caption and each audio caption includes a natural language text description of a sound, and wherein the one or more caption embeddings are selected based on a first similarity metric indicating a similarity between the one or more caption embeddings and the one or more query caption embeddings; and means for generating a search result identifying one or more first media files in the set of media files, each first media file in the one or more first media files being associated with at least one of the one or more caption embeddings.

30. The apparatus according to claim 29, wherein the means for generating the one or more query caption embeddings, the means for selecting one or more caption embeddings, and the means for generating the search result are integrated in a speaker bar, a mobile communication device, a smart phone, a cellular phone, a laptop computer, a computer, a tablet computer, a personal digital assistant, a display device, a television, a game console, a music player, a radio, a digital video player, a digital video disc (DVD) player, a tuner, a camera, a navigation device, a vehicle, a headset, an augmented reality headset, a mixed reality headset, a virtual reality headset, an aircraft, a home automation system, a voice-activated device, a wireless speaker and voice-activated device, a portable electronic device, an automobile, an Internet of Things (IoT) device, a virtual reality (VR) device, a base station, or any combination thereof.