Systems and methods for searching and presenting media content
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2026-03-05
AI Technical Summary
Searching for specific audio information within vast digital media is challenging due to the sheer volume of available data, and existing systems lack efficient organization and classification methods.
An AI-driven search engine that converts audio files to text data, generates indexed vectors with timestamps and metadata, and uses a large language model to classify segments, enabling deep linking and commerce options based on user queries.
Facilitates precise audio content retrieval with deep linking to relevant segments and provides commerce options, enhancing user interaction and content discovery.
Smart Images

Figure US2025035016_05032026_PF_FP_ABST
Abstract
Description
SYSTEMS AND METHODS FOR SEARCHING AND PRESENTING MEDIACONTENTCROSS-REFERENCE
[0001] This application claims priority to U.S. Provisional Patent Application No. 63 / 664,305, filed on June 26, 2024, which is entirely incorporated herein by reference for all purposes.BACKGROUND
[0002] Media refers to communication tools used to store and deliver information or data, and includes but is not limited to mass media, multimedia, news media, published media, hypermedia, broadcast media, advertising media, social media (i.e., media disseminated through social interaction) and / or other types of media. Media can be delivered over electronic communication networks. Media communications can include digital media (i.e., electronic media used to store, transmit, and receive digitized information). Media includes privately communicated information or data, publicly communicated information or data, or a combination thereof. Media can be Social media. Social media includes web-based and mobile technologies which may or may not be associated with social networks. For example, social media such as blogs may not be associated with a social network. Social media includes privately communicated information or data, publicly communicated information or data, or a combination thereof. Currently there are multiple social networks that enable users to post content.
[0003] Given the vast amount of digital information readily available, searching for specific information particularly in the form of audio can be a challenge.SUMMARY
[0004] This present disclosure provides systems and methods for searching media content. In particular, systems and methods of the present disclosure provide improved search engine and media platform for ingesting and organizing audio content into multiple orthogonal dimensions then classifying audio content into multiple perspectives / lenses within a dimension leveraging artificial intelligence (Al) or machine learning technology. Systems and methods provide an audio content search engine along with capability to deep link an exact moment / spot of a segment of the audio content that is relevant to the query. Additionally, one or more commerce options may be provided to a user based on the query and the media content. Such media content may be in the form of audio data such as podcasts, speeches, recorded presentations, recorded lectures, phone calls, broadcast, meetings, and the like.
[0005] In an aspect of the present disclosure, a method for searching audio content is provided. The method comprises converting a plurality of audio files into a text data, generating indexed data based on the text data; receiving a search query on an electronic device of a user to search across the plurality of audio files and identifying one or more relevant audio files and one or more relevant segments within each of the one or more relevant audio files in response to the search query; adding a deep link to each of the one or more relevant segments for the user to choose to play a selected relevant segment via GUI; and displaying one or more commerce options within the GUI, wherein the one or more commerce options are identified based at least in part on the search query and the one or more relevant segments within each of the one or more relevant audio files.
[0006] In an aspect, a computer-implemented method for querying and presenting audio content is provided. The method comprises: ingesting a plurality of audio files from a plurality of sources; converting the ingested plurality of audio files into indexed data, wherein the indexed data comprises a vector representation of a transcription of a segment within an audio file, timestamps associated with the segment, and extracted metadata; augmenting the indexed data by assigning a category to the respective segment with aid of a large language model (LLM); and displaying, on a graphical user interface (GUI), a first filter dimension corresponding to filtering across the plurality of sources and a second filter dimension corresponding to filtering across a plurality of categories.
[0007] In a related yet separate aspect, a system for querying and presenting audio content is provided. The system comprises one or more computer processors are individually or collectively programmed to: ingest a plurality of audio files from a plurality of sources; convert the ingested plurality of audio files into indexed data, wherein the indexed data comprises a vector representation of a transcription of a segment within an audio file, timestamps associated with the segment, and extracted metadata; augment the indexed data by assigning a category to the respective segment with aid of a large language model (LLM); and display, on a graphical user interface (GUI), a first filter dimension corresponding to filtering across the plurality of sources and a second filter dimension corresponding to filtering across a plurality of categories.
[0008] In some embodiments, the plurality of audio files comprise a plurality of podcasts. In some embodiments, ingesting the plurality of audio files comprises using a domain-aware semantic crawler to detect media content in audio-format. In some cases, ingesting the plurality of audio files comprises assessing a credibility of a source from the plurality of sources using a classification model.
[0009] In some embodiments, the plurality of sources comprise public websites, manual uploads and content stream sources. In some embodiments, converting the ingested plurality of audio files into indexed data comprises chunking an audio file into one or more segments and associating each segment with a start timestamp and an end timestamp. In some embodiments, the extracted metadata comprises at least one of language, duration, source of the audio file.
[0010] In some embodiments, the LLM is fine-tuned using training dataset comprising a label of a category and a rationale for assigning the category. In some cases, the method further comprises evaluating the LLM’s performance for generating a rationale for assigning a category.
[0011] In some embodiments, the method further comprises receiving a search query via the GUI and identifying one or more relevant segments in response to the search query. Upon receiving the search query, the method comprises sorting the one or more relevant segments based at least in part on a category assigned to each of the one or more relevant segments and one or more metadata fields associated with the one or more relevant segments.
[0012] In some cases, the search query is received by receiving different gestures corresponding to the first filter dimension and the second filter dimension. In some cases, the one or more relevant segments are identified by converting the search query into a vector format and performing a semantic search by processing the search query in the vector format and the indexed data. In some instances, the method comprises displaying, via the GUI, the one or more relevant segments with the associated category and an option to share one or more segments selected from the one or more relevant segments. In some cases, the method comprises embedding a deep link to each of the one or more relevant segments for the user to choose to play a selected relevant segment via the GUI and further displaying a translated transcription of the one or more metadata fields, a translated transcription in a user preferred language. In some cases, the method comprises displaying one or more commerce options within the GUI, wherein the one or more commerce options are identified based at least in part on the search query and the one or more relevant segments.
[0013] In another aspect, a computer-implemented method for querying and presenting audio content is provided, the computer-implemented comprising: providing a graphical user interface (GUI) for searching audio content, wherein the GUI comprises a first filter dimension corresponding to filtering across the plurality of sources and a second filter dimension corresponding to filtering across a plurality of categories associated with the audio content gathered from the plurality of sources; receiving a search query from the GUI for filtering the audio content; identifying one or more relevant segments from the audio content in response to the search query, wherein a segment is converted to an indexed data comprising a vectorrepresentation of a transcription of the segment within an audio file, and is associated with timestamps, extracted metadata and an assigned category; and displaying the one or more relevant segments as search result within the GUI, wherein each relevant segment is displayed with the assigned category, a playback option or an option to share the relevant segment.
[0014] In some embodiments, the assigned category is generated by a large language model (LLM). In some cases, the LLM is fine-tuned using training dataset comprising a label of a category and a rationale for assigning the category. In some cases, the search query is received by receiving different gestures corresponding to the first filter dimension and the second filter dimension.
[0015] Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.INCORPORATION BY REFERENCE
[0016] All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and / or take precedence over any such contradictory material.BRIEF DESCRIPTION OF THE DRAWINGS
[0017] The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also “Figure” and “FIG.” herein), of which:
[0018] FIG. 1 shows an example of a search engine and various features provided by the search engine;
[0019] FIG. 2 shows a non-limiting example of a computing device; in this case, a device with one or more processors, memory, storage, and a network interface, in accordance with some embodiments;
[0020] FIG. 3 shows a non-limiting example of a web / mobile application provision system; in this case, a system providing browser-based and / or native mobile user interfaces, in accordance with some embodiments; and
[0021] FIG. 4 shows a non-limiting example of a cloud-based web / mobile application provision system; in this case, a system comprising an elastically load balanced, auto-scaling web server and application server resources as well synchronously replicated databases, in accordance with some embodiments.
[0022] FIG. 5 shows various components of the search engine system, in accordance with some embodiments of the present disclosure.
[0023] FIG. 6 shows examples of perspectives provided by the search engine system herein.
[0024] FIGs. 7-8 show examples of GUI allowing users to navigate the audio content via different gestures (e.g., swipe left / right, swipe up / down).
[0025] FIG. 9 shows an example of switching from filtering along the lens dimension to the source dimension.
[0026] FIG. 10 shows an example of audio media content search result presented in the form of timestamped segment.
[0027] FIG. 11 shows an example of multilingual content presentation feature of the GUI.DETAILED DESCRIPTION
[0028] While various embodiments of the invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. Numerous variations, changes, and substitutions may occur to those skilled in the art without departing from the invention. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention.
[0029] The term “media source,” as used herein, generally refers to a service that stores or otherwise provides media content. A media provider may be a “social media source,” which generally refers to a service that stores or otherwise provides social media content. A media provider can turn communication into an interactive dialogue. Some media providers, such as social media providers, can use web-based and mobile technologies to provide such interactive dialogues.
[0030] The term “media content,” as used herein, generally refers to textual, graphical, audio and / or video information (or content) on a computer system or a plurality of computer systems. In some cases, media content can include uniform resource locators (URL’s). Media content may be information provided on the Internet or one or more intranets. In some cases, media content may comprise social media content, which can include textual, graphical, audio and / or video content on one or more social media networks, web sites, blog sites or other web-based pages. Social media content may be related to a social entity or social contributor. In some cases, media content may be generated by any third-party software such as conferencing, chatting software, software for audio / video conference calls and various others.
[0031] In particular, systems and methods of the present disclosure provide improved search engine and media platform for ingesting and organizing audio content into multiple orthogonal dimensions then classifying audio content into multiple perspectives / lenses within a dimension leveraging artificial intelligence (Al) or machine learning technology. In some embodiments, a media player (e.g., audio media player, video player) with improved search engine may be provided. Systems and methods provide an audio content search engine along with capability to deep link an exact moment / spot of a segment of the audio content that is relevant to the query. Additionally, one or more commerce options may be provided to a user based on the query and the media content. Such media content may be in the form of audio data such as podcasts, speeches, recorded presentations, recorded lectures, phone calls, broadcast, meetings, and the like.
[0032] The terms “dimensions” may refer to multi-dimensional views of media content streams, such as across a source dimension, a contributor dimension, or topic dimension. Such dimensions (e.g., sources / entities, contributors / speaker / author and topic / tags / key words are also referred to as “dimensions” or “social dimensions” herein. In some embodiments, a social dimension can refer to a particular type of dimension associated with social media content, while a dimension can refer to a dimension associated with any type of media content. Unlike conventional organization of content using a tree- structure (e.g., category branched or divided into subcategories), the dimensions can be orthogonal dimensions — that is, dimensions that are mutually exclusive such that filtering along a selected dimension can get deeper - drilling down (e.g., applying further new topic filters along the topic dimension) as many detail levels as a user prefers. The orthogonality of the multiple dimensions also permits drilling down a media content from any drill level along any dimension i.e., different dimensions can be available at any given detail level to further drill down the content at the given level without reverting to previous level of drill. As an alternative, some of the dimensions can be inclusive of other dimensions. Forexample, the social entity dimension can overlap at least to some extent with the social contributor dimension.
[0033] The terms “perspective,” “lens” as utilized herein may refer to classifications such as topics / industries / key words for classifying or organizing the medica content within the topic dimension. As described later herein, the systems and methods herein may provide a unique audio content processing algorithm leveraging Al with human defined perspectives as guardrails for providing a balanced and interpretative perspective. Perspectives may also be referred to as classifiers enriched by human curated interpretative framing.
[0034] Whenever the term “at least,” “greater than,” or “greater than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “at least,” “greater than” or “greater than or equal to” applies to each of the numerical values in that series of numerical values. For example, greater than or equal to 1, 2, or 3 is equivalent to greater than or equal to 1, greater than or equal to 2, or greater than or equal to 3.
[0035] Whenever the term “no more than,” “less than,” or “less than or equal to” precedes the first numerical value in a series of two or more numerical values, the term “no more than,” “less than,” or “less than or equal to” applies to each of the numerical values in that series of numerical values. For example, less than or equal to 3, 2, or 1 is equivalent to less than or equal to 3, less than or equal to 2, or less than or equal to 1.
[0036] Large Language Models
[0037] A large language model (LLM) is a machine learning model that is trained to predict the next word in a sequence of words, given the context of the previous words. These models are typically trained on a very large dataset of text, such as a corpus of books, articles and / or the Internet, and are able to generate human-like text that is coherent and flows naturally. The size of a large language model refers to the number of parameters it has, which determines its capacity to learn and generate text. Models with more parameters are able to learn more complex patterns in the data, but also require more computational resources to train and use.Systems and methods for querying and presenting audio data
[0038] In some embodiments, an Al-based search engine for audio data is provided. In some embodiments, a media player (e.g., audio media player, video player) with improved search engine may be provided. Though the search engine is described in the context of audio content or video content, it should be noted that various algorithms, and methods herein can be applied to any media content that may not be audio data. In some embodiments, a media player (e.g., an audio media player or video player) may be equipped with an enhanced search engine thatprovides improved content discovery, retrieval, and user interaction capabilities. FIG. 1 shows an example of the search engine integrated into a media player and various features provided by the search engine. FIG. 1 shows an example of the search engine and various features provided by the search engine.
[0039] FIG. 1 shows an example of a graphical user interface (GUI) provided by the search engine, in accordance with some embodiments. The search engine herein may receive input data comprising audio data to be searched later by any user. The audio data may be uploaded by a user, collected from the Internet, or recorded in real time. The audio data may include, for example, podcasts, speeches, videos, audio lectures, phone calls, broadcast, meetings, audio data embedded in any websites, and the like. The systems and methods herein may provide a unique content discovery engine configured to identify, discover and ingest the audio data. Details about the content discovery engine are described with respect to FIG. 5.
[0040] The search engine may convert the audio data into text content. In some cases, deep learning technologies may be employed to transcribe the audio to obtain text data in a text format, and the text data may be sliced to obtain a set of data segment for processing by a large language model (LLM). In some cases, the search engine may employ optimized speech-to- text models that are fine-tuned on the target audio domains to increase transcription accuracy.
[0041] The text data may then be indexed such that the text content transcribed from a plurality of audio files (e.g., podcasts) may be indexed, archived and searchable by the search engine. In some cases, the transcribed data may be converted to vectorial representation. In some cases, the search engine may use one or more chunking techniques to split text from one or more audio files into semantically similar chunks that can then be embedded and use in search lookups. In some cases, chunking is performed based at least in part on a fixed chunk size with some chunk overlap for text that is too long. In some cases, indexes are generated for the text data. In some cases, the indexes may comprise sparse index each may index or correspond to an element such as a token, a chunk, a word, a passage or other data structures for efficient query. It should be noted that any other suitable data structures may be used for the search of the elements. In some cases, the indexes may also comprise a dense index each may correspond to or index elements such as a representation of textual items (e.g., word embeddings of a chunk of texts). For example, a word embedding (a real-valued vector that encodes the meaning of the word such that the words that are closer in the vector space are expected to be similar in meaning) may be indexed and a dense index may be created. In some cases, sparse and dense indexing is utilized for all text data. The search engine may transform the transcribed text into parsed text which maythen be converted into chunks. The chunks may be embedded and stored. A chunk may, for example, be a sentence, a section, a passage, a paragraph, that forms a query result.
[0042] The above preprocessing and indexing operations may be performed by an indexing engine which may comprise distributed components or modules configured to listen for incoming data from the content discovery engine (i.e., ingestion phase) and perform parallelized preprocessing operations to prepare the content for efficient retrieval and downstream analysis. Details about the indexing engine are described later herein.
[0043] The processed text content may then be searchable or queried by a user. For example, a shown in the GUI, a user may search across a plurality of audio files (e.g., podcasts) by inputting a query 101. Upon receiving the query, the search engine herein may provide a list of one or more relevant podcasts 103. For example, the search result may comprise a list of podcasts 105. Each podcast may be displayed along with the transcription of relevant content 107 and an option for the user to play the audio 109 of the relevant content.
[0044] The relevant content may be identified using a generative Large Language Model (LLM). The LLM can be any suitable LLM. For instance, the LLM may be the standard nextword prediction model. The relevant content (e.g., sentences, passages, etc.) may be identified by a generative LLM in response to a user’s question or query. The LLM may enrich the ingested data and classify the data into a plurality of perspectives. The LLM may be classifiers enriched by human curated interpretative framing. The LLM may be trained utilizing a unique method such that the perspectives are balanced (i.e., relatively equally represented). In some cases, the LLM may be fine tuned to learn the human interpretation rules. Details about the LLM training, and fine-tuning for augmentation or data enrichment into unique perspectives are described later herein.
[0045] The search operations may comprise routing queries initiated (e.g., from the GUI of the search engine application) through a semantic search layer configured for speed, relevance, and contextual understanding. The search engine may generate deep link for each identified relevant content (e.g., sentences, passages, etc.), for example, one or more timestamps may be displayed within a podcast that is identified as relevant to the users’ query and the one or more timestamps are associated with one or more deep links to direct a user to the exact spot in the podcast.
[0046] In some cases, the search engine may be capable of ranking the one or more relevant podcasts and one or more sections within a podcast by degree of relevance based on the user input query and the content indexes built as described above.
[0047] In some embodiments, each entry of the result (e.g., podcast and its relevant content) may further comprise one or more commerce options 111. The commerce option 111 may be provided to a user based at least in part on a product (i.e., commerce option), a user query input and a content of the podcast. For instance, LLM may be employed to understand a user’s intent (e.g., desired information) without being intrusive. The LLM may take as input the input query, identified relevant information and available product based on intent match, phrase match, semantic matching, broad match or others. In some cases, the commerce option 111 may be a local service that is identified based at least in part on a geolocation of the user.
[0048] The commerce options may be retrieved from a managed products inventory. Alternatively, the search engine may include a data mining module for searching the Internet and / or an intranet to collect product or commerce options of interest to a user.
[0049] FIG. 5 shows various components of the search engine system 500, in accordance with some embodiments of the present disclosure. The search engine system 500 may comprise a content discovery engine 510 configured to identify, discover and ingest content data from a variety of sources, globally and automatically.
[0050] In some embodiments, the media content data may comprise audio data to be searched later by any user. The audio data may be uploaded by a user, collected from the Internet, or recorded in real time. The audio data may include, for example, podcasts, speeches, videos, audio lectures, phone calls, broadcast, meetings, audio data embedded in any websites, and the like. The content discovery engine 510 may be a scalable infrastructure configured to identify and collect the audio content data from a variety of sources.
[0051] In some embodiments, the scalable infrastructure may comprise a public audio discovery engine, a suite of application programming interfaces (APIs), features for receiving and processing manual uploads of audio content, streaming hooks and various other components.
[0052] In some embodiments, the public audio discovery engine may be a unique tool designed to programmatically identify and catalog publicly available audio content across institutional, corporate, governmental, academic websites, and any other publicly available sources. The public audio discovery engine may comprise a domain-aware semantic crawler. In some cases, the domain-aware semantic crawler may be executed to target known public institutions such as universities, government agencies, organizations, companies using seed URLs. The domain-aware semantic crawler may be a specialized web crawler designed to understand the context and structure of different domains (e.g., academic, government, corporate), interpret content semantically, rather than just scraping links blindly, and prioritize discovery of audio files or pages with embedded audio.
[0053] The seed URLs may be curated and / or dynamically generated. The domain-aware semantic crawler may be capable of managing seed URLs. For instance, the domain-aware semantic crawler may use a list of seed URLs as starting points (e.g., harvard.edu, nasa.gov, who.int) and / or use APIs or search queries to find new institutional or organizational domains to update the list of seeds (i.e., dynamically generated seed URLs). This beneficially provides high- quality entry points for domain-specific exploration such as starting crawling from reliable public domains. The domain-aware semantic crawler may then use domain classifiers to tag and cluster domains (e.g., .edu —> academic, .gov —> government).
[0054] In some cases, the domain-aware semantic crawler may extract metadata by detecting audio-relevant formats (e.g., MP3, WAV, OGG, AAC etc.) and associated metadata schemas embedded within web pages, including but not limited to HTML5<audio> tags, downloadable links, and RSS enclosures. Examples of metadata extracted may include, for example, title, duration, speaker, topic, source domain and the like. The crawler may be domain-aware and may adjust crawling behavior based on domain type. The crawler may comprise a semantic parser to understand the intent and content structure of each page and focus on detecting and prioritizing publicly available audio content from web domains. As an example, the crawler may scan HTML pages and files for links to . mp3, .wav, .ogg, .aac, ,m4a, .webm, .flac and inspect <audio> and <source> tags, detect <enclosure> tags in RSS / XML for RSS feed parsing, and perform metadata extraction from HTML by extracting <title>, <meta name="description">, <meta property="og:audio">, extracting <title>, <description>, <itunes:duration>, <pubDate>, <itunes:author> from RSS feed.
[0055] In some cases, the domain-aware semantic crawler may apply content classification models to determine topical relevance and source credibility, allowing prioritization of high- value datasets. The domain-aware semantic crawler may comprise a topical relevance classifier such as a fine-tuned transformer model (e.g., distilBERT, MiniLM) to classify transcripts or page text into topics (e.g., healthcare, policy, education, corporate communications). The input to the topical relevance classifier may include, for example, page text, transcript (if available), RSS item descriptions and the output may include a confidence score per topic. In some cases, the domain-aware semantic crawler may utilize pre-trained credibility scoring models or heuristic rules to perform credibility estimation. For instance, domain-level heuristics may include assigning higher credibility to domains such as .edu, .gov, presence of known institutional names or based on authorship metadata.
[0056] In some cases, the domain-aware semantic crawler outputs a structured index of discovered sources, metadata, and access pathways for ingestion into the search engine system.In the structured index of the discovered audio content data and metadata may be designed for search integration. For example, the ingested audio content data may be stored in index document format including title, source URL, domain, format, duration, published data, topics, credibility score, access type, and the like. The ingested structured data may be stored in a database such as Elasticsearch (for full-text search & filtering) or PostgreSQL with vector similarity (e.g., pgvector).
[0057] In some embodiments, the content discovery engine may comprise APIs integration features allowing for integrations with public and private APIs to fetch structured metadata and media files efficiently. These APIs (e.g., YouTube Data API, Spotify API, and RSS feed readers) may support periodic or event-driven polling for new audio content, ensuring timely updates from major platforms. For instance, the API integration feature may comprise scheduler or event listener using CRON jobs, serverless triggers (AWS Lambda, Cloud Scheduler) or message queues (e.g., Kafka) to periodically check for updates. In some cases, the API integration features may comprise API request module to adapt to various rate limits or authentication requirement (e.g., OAuth2 authentication handling for Spotify and YouTube). The API integration feature may Parser or Normalizer configured to normalize API-specific data into the unique ingestion schema as described above.
[0058] In some embodiments, the content discovery engine may comprise features permitting data ingestion from manual uploaded media content. For instance, the manual upload features may comprise a secure portal or backend mechanism for manual contribution of content, used primarily for exclusive or proprietary material. The system herein may provide secure file uploads such as uploading to AWS S3, GCP Cloud Storage, or on-prem storage via presigned URLs. The search engine system may provide UI allowing user to upload the media content. The manual upload features may automatically extract metadata fields (e.g., Title, Author, Language, Topic, Rights, Description, Tags, etc) from the uploads.
[0059] In some embodiments, the content discovery engine may comprise streaming hooks or webhooks allowing for real-time ingestion of media stream. The real-time ingestion mechanisms may subscribe to content streams and automatically ingest media upon publication. For instance, the real-time ingestion mechanism may ingest media content from event sources such as webhook-enabled platforms (e.g., SoundCloud, proprietary CMS, GitHub for podcast repos) or RTMP / RTSP streams for live events. The real-time ingestion mechanism may comprise a webhook receiver (Expresses, FastAPI, Flask), signature validation to verify authenticity and event parser for metadata extraction.
[0060] In some embodiments, the search engine system 500 may comprise an indexing engine 520 configured for preprocessing of the ingested data and performing indexing operations. Instead of storing the ingested data directly, the indexing engine may improve the memory efficiency and downstream search functions by providing unique indexing algorithms. In some cases, the ingested data may be stored in cloud-native databases, local file systems or a combination of both. The databases may send the ingested data and files to the indexing engine 520 via API calls (e.g., a REST or gRPC API).
[0061] In some embodiments, the indexing engine may comprise distributed components or modules configured to listen for incoming data from the content discovery engine (i.e., ingestion phase) and perform parallelized preprocessing operations to prepare the content for efficient retrieval and downstream analysis. In some cases, the plurality of distributed components may comprise at least a transcription module 521, a metadata extractor 523 and a segmentation and stamp component 535.
[0062] The architecture of the indexing engine ensures the system remains horizontally scalable and fault-tolerant, allowing new content to be processed and indexed in near real-time.
[0063] In some embodiments, the transcription module 521 may be configured to transcribe audio / video content using automated speech recognition (ASR) models, producing time-stamped transcripts. In some cases, the transcribed data may be converted to vectorial representation. In some cases, indexes are generated for the text data. In some cases, the indexes may comprise sparse index each may index or correspond to an element such as a token, a chunk, a word, a passage or other data structures for efficient query. In some cases, the indexes may also comprise a dense index each may correspond to or index elements such as a representation of textual items (e.g., word embeddings of a chunk of texts). For example, a word embedding (a real-valued vector that encodes the meaning of the word such that the words that are closer in the vector space are expected to be similar in meaning) may be indexed and a dense index may be created. In some cases, sparse and dense indexing is utilized for all text data. The search engine may transform the transcribed text into parsed text which may then be converted into chunks. The chunks may be embedded and stored. A chunk may, for example, be a sentence, a section, a passage, a paragraph, that forms a query result.
[0064] In some cases, the transcription module 521 may comprise an ASR engine that may comprise preprocessing of the input files (e.g., audio file, video file) such as by normalizing audio, removing silence, converting video to audio file. The transcription module outputs time- aligned text segments.
[0065] In some embodiments, the metadata extractor 523 may extract metadata such as speaker identity (if available), language, duration, and source platform. The metadata extractor may be configured for identifying and extracting structured, machine-readable information from audio and / or video content during or after ingestion. For instance, the metadata extractor 523 may extract speaker names from the associated metadata (e.g., podcast RSS feeds, YouTube descriptions, embedded XML) in the ingested data to identify the speaker identity, use speaker diarization models to detect who spoke when, assigning speaker IDs. In some cases, the metadata extractor 523 may identify the primary spoken language in the audio content to aid in transcript accuracy, indexing, and search. The language may be identified by a built-in language detection feature of the ASR model or a separate language ID model. In some cases, the metadata extractor 523 may measure the total length of the audio or video file such as by using tools like ffprobe (part of ffmpeg) to extract file duration directly from the media. In some cases, the metadata extractor 523 may identify the source platform or domain from which the media originated (e.g., YouTube, Spotify, university website) such as by parsing the source URL to infer the platform, using regex or rules to extract host / platform metadata, or using API response fields (e.g., provider, uploader, channel, feedTitle) when the content data is ingested via API.
[0066] In some embodiments, the segmentation and stamp component 535 may be configured to segment content into coherent topical chunks, aligning transcript text with timestamps for improved navigability. In some cases, the segmentation and stamp component 535 may use one or more chunking techniques to split text from one or more audio files into semantically similar chunks that can then be embedded and use in search lookups. In some cases, chunking is performed based at least in part on a fixed chunk size with some chunk overlap for text that is too long. In some cases, the segment length may be configurable by a user. For example, a user may be permitted to set up a preferred segment lengths via a GUI (e.g., 30 seconds, 2 minutes) and the system may use the user-set segment length as target length for generating the segments. For instance, two or more preliminary segments may be further assembled into a segment within a target range set by the user-defined segment length.
[0067] The segmentation and stamp component 535 may divide transcripts into coherent topical chunks that preserve semantic meaning. In some cases, the segmentation and stamp component 535 may receive structured transcript (from ASR module) and perform chunking. The chunking may comprise semantic chunking using language models to detect natural topic shifts or coherence boundaries and / or a fixed-size chunking with overlap. The semantic chunking may utilize natural language processing techniques such as lexical cohesion to group related sentences. In some cases, semantic chunking may employ transformer-based embeddings such as sentence embeddings (e.g., SBERT) to calculate cosine similarity between adjacent blocks andmark segment boundaries based on similarity threshold. The fixed size chunking with overlap methods may comprise breaking the transcript into chunks of N sentences or AT tokens / words with each chunk sharing a portion with the previous chunk to preserve context.
[0068] The segmentation and stamp component 535 may be configured to align each chunk with corresponding audio timestamps for easy navigation, preview, sharing, and search result anchoring. For instance, each chunk may be associated with an accurate start and end timestamps for time-synced playback, transcript alignment, search, sharing and various other downstream analysis. In some cases, the segmentation and stamp component 535 may aggregate timestamps of words / sentences within each chunk and use the first sentence's start time and last sentence's end time.
[0069] In some cases, the indexing engine 520 may store the processed output in a vectorized format using embeddings (e.g., via transformer models) and index them in a high-performance vector database for fast semantic search). In some cases, the transcribed text data or chunks may then be converted into embeddings (vectors) such as by the normalization and augmentation module 530 for semantic processing. Alternatively, the topical chunks with timestamps may be converted to vectorized representations by the transcription module 521 or the segmentation and stamp component 535. Each finalized chunk may be passed to an embedding model for semantic vectorization. In some cases, the input to the embedding model may comprise the timestamped segments which can be the output of the ASR module, the segmentation and stamping component and / or the metadata extractor. The input features may comprise, for example, media ID< segment ID, start time stamp, end time stamp, text, metadata such as speaker, language, source, etc.
[0070] The indexing engine may use pre-trained or fine-tuned models to generate vector embeddings of the timestamped segment text. The transfer model may be Sentence Transformers (SBERT) for semantic similarity, OpenAI Embeddings, or any other transformer model such as custom BERT variants.
[0071] The indexed data or embeddings may be stored in high-performance vector database that can be cloud-native databases, local file systems or a combination of both. Unlike conventional SQL databases that aren't optimized for high-dimensional similarity search (e.g., cosine or dot product), a vector database beneficially allows for approximate nearest neighbor (ANN) search, real-time ranking of relevant results by semantic proximity and various other functions. The content in the vector format may be stored in the database by lining to metadata (e.g., media ID< timestamps, topic, etc).
[0072] The vector format of the segments can allow for fast semantic search. For instance, a user query may be converted to vector using the same embedding model by the search module 540 as described later herein.
[0073] The search engine system 500 may comprise a normalization and augmentation module 530 to generate enriched data. The normalization and augmentation module 530 may receive the segmented data with metadata in the vector format and enrich the data by Al generated classifications or perspectives.
[0074] In some cases, perspectives may be classifiers enriched by human curated interpretative framing. FIG. 6 shows examples of perspectives 610. Each of the perspective (e.g., politics, government, technology, social life, etc.) may be associated with human interpretation guide 613. For instance, the interpretation guide for politics perspective may be “discussion of ideologies, political actors (e.g., parties, campaigns), elections, activism, or debates on laws and values. Focused on power struggles or value-driven conflict.”
[0075] The perspectives may be generated by training a model using training datasets. The training dataset may comprise annotated content data. For example, each episode may be listened by annotators and then manually categorized. In some cases, the training data may comprise rationale for the categorization. FIG. 6 shows examples of the training dataset 620. As shown in the example, a training data may comprise title, excerpt, labels (i.e., ground truth of the perspective), and rationale. The training data may be created to be balanced such that each perspective is equally represented in the training dataset. Next, the system may finetune LLM so that it learns human interpretation rules. The finetuned LLM may enrich the already-segmented, vectorized transcript data by assigning interpretive perspectives (e.g., politics, technology, social issues) and generating rationale explanations for those labels.
[0076] In the input to the LLM classifiers may comprise the structured and enriched audio data in the vector format. The output of the LLM classifiers may include perspectives or high- level interpretive lenses (e.g., Politics, Social Life, Technology, Science, etc.) and each perspective is associated with curated framing guides written by subject matter experts. Such as framing guides may be used for model training to provide rationale-focused fine-tuning and post- hoc interpretability.
[0077] The training dataset may be balanced that is equal representation of all perspectives, multi-labeled that is one chunk can have multiple labels and include rationale for supervised learning and explanation generation.
[0078] In some cases, the LLM base model may be a transformer LLM. The above- mentioned dataset may be used for finetuning the base model. In some cases, finetuning may comprise employing the following strategies: multi-label classification head, optional sequence- to-sequence generation head (for rationale). In some cases, the training process may use Binary Cross Entropy (BCE) loss for the perspective classification, and use cross-entropy loss for rationale generation.
[0079] In an exemplary training process, the model (e.g., a fine-tuned Transformer like BERT, RoBERTa, or T5) is trained to classify transcript chunks into one or more perspective labels (e.g., Politics, Technology), the process involves forward pass that is the input transcript chunk fed through the transformer model, producing a set of logits or probability scores for each possible label, loss computation that is a multi-label loss function (like Binary Cross Entropy) comparing the model's predicted label probabilities against the ground-truth labels, backward pass (B ackpropagation) that is the gradient of the loss with respect to all model parameters is computed and propagated from the output layer back through the entire network, and parameter update where an optimizer (e.g., Adam) may be employed to use these gradients to update the model weights to minimize future prediction errors. In some cases, a sequence-to-sequence architecture (e.g., T5, BART, GPT-style) of the model may be employed to generate rationale text to explain why a chunk was labeled with a perspective. During training, loss gradients are backpropagated through the model layers (encoder and decoder) to improve future generation accuracy. This helps the model imitate human reasoning structures when fine-tuned on annotated examples.
[0080] The finetuned model may be evaluated through automatic metrics such as Accuracy, Fl-score, Precision, Recall (multi-label setup), and BLEU or ROUGE for rationale explanations. For instance, the evaluation methods may use accuracy, precision / recall / Fl (micro, macro, weighted) for evaluating the classification or predicting the perspectives. The methods may utilize BLEU / ROUGE to assess overlap with human-written rationales for evaluating the rationale generation.
[0081] The finetuned model may be further evaluated in the quality and interpretability of the perspective labels and coherence and relevance of the rationale. For instance, quality and relevance of the assigned perspectives and the generated rationale may be evaluated by human expert.
[0082] Referring back to FIG. 5, the search engine system 500 may comprise a search module 540 configured to receive queries initiated from a user interface (UI). The queries may be routed through a semantic search layer designed for speed, relevance, and contextualunderstanding. In some cases, the search module 540 may perform query parsing such as parsing natural language queries and generate embeddings of the queries.
[0083] In some cases, the search module 540 may be configured to perform similarity Search. For instance, the embedded query may be matched against the vector index of the ingested content using approximate nearest neighbor (ANN) techniques to retrieve the most semantically relevant content chunks. The search module 540 may be configured to perform semantic retrieval of content by comparing vectorized user queries to the vector embeddings of previously ingested and indexed content chunks. The same transformer embedding model used during content ingestion may be used to generate the query embeddings to ensure vector space compatibility. Compared to conventional similarity search algorithm that finds exact nearest neighbor which can be too slow for large-scale data, approximate nearest neighbor (ANN) finds the top-K most semantically similar vector entries from the indexed content which beneficially improves speed. As described above, during ingestion, the content chunks may be converted into embeddings and stored in an ANN index along with the timestamps and metadata (e.g., assigned perspectives, source, speaker, etc.).
[0084] In some cases, the search module 540 may be configured to perform metadata filtering. For instance, optional filters (e.g., date range, speaker, platform, topic) are applied to narrow the search results based on the structured metadata stored with each vector segment. Upon executing the ANN algorithm, the top-K candidate chunks or vector segments may be retrieved based on cosine or L2 distance, the associated metadata may also be fetched and used to further sort or filter the results (e.g., top-K candidate vector segments) by the metadata fields such as by source, timestamp, perspective, speaker, or to support hybrid ranking (e.g., combining vector similarity with traditional keyword match or popularity scores).
[0085] In some cases, the search module 540 may be configured to perform ranking. For instance, the retrieved results are scored and ranked using a hybrid model that combines semantic similarity, recency, popularity, and user preferences. The hybrid scoring model aggregates multiple factors into a final ranking score for each retrieved result. These factors may include semantic similarity (e.g., output from the ANN search algorithm such as cosine or dot-product similarity between query and content chunk embeddings), recency (e.g., preference for newer content (based on publication or recording date) by applying a decay function on age of the publication), popularity (e.g., aggregate usage metrics (e.g., listens, shares, likes, citation count)), and user preferences (e.g., personalization parameters such as language, topics of interest, or interaction history). The ranking score may be a weighted linear combination of the abovefactors. Alternatively, the ranking score may be generated by a dynamic model such as a learn- to-rank (LTR) model.
[0086] In some cases, the search module 540 may be configured to perform aggregation. The retrieved results are finally aggregated by perspectives and presented to users accordingly. In some cases, a search result to a query may comprise a list of results each result may comprise one or more timestamped segments within an episode. In some cases, the aggregation rule may be configurable. For instance, a user may be permitted to set up how the results are aggregated and presented (e.g., aggregated per episode, per source, per speaker, arranged by timestamps, arranged by speaker, topics, etc). Rather than simply listing chunks, the system groups related results based on contextual metadata or user-configurable rules — such as episode, speaker, topic, or perspective — improving both navigability and insight discovery. In some cases, the input to the aggregation logic may comprise the results from semantic retrieval and hybrid ranking as described above which may include a list of timestamped result chunks, each linked to a source (e.g., episode, file, speaker, topic). The aggregation logic may group segments that share the same interpretive perspective (e.g., Politics, Technology) and within each perspective, the results can be further grouped by episode and sorted by timestamp.
[0087] The GUI 550 may be provided to display the filter results received from the search module in a configurable manner. For instance, matching segments are returned with context windows and timecodes, enabling users to jump directly to relevant moments in the source content and optionally playback them in the media player application. The GUI may be configured to render the filtered and aggregated search results to the user which allows direct interaction with time-aligned content segments, enabling preview, playback, filtering, and navigation through semantically indexed media. For example, each search result may comprise the transcript of the segment, graphical element for clickable to play the segment, start / end timecodes (from original audio / video source), source metadata (e.g., episode title, speaker, date, platform), highlighted match terms (if keyword or hybrid search is used), perspective or topic tags. In some cases, each segment may comprise a context window — a configurable number of seconds or sentences before and after the matched chunk to improve comprehension. The media playback feature of the GUI may stream audio / video inline at the precise start timestamp associated with the segment, highlight or caption the matched portion during playback, and various other features.
[0088] The highly granular structure of the indexed content enables precise and flexible sharing options. For instance, timestamp-based sharing features may allow users to share exact moments from the source content, down to the sentence or phrase level, using timestamped linksor auto-generated video / audio snippets. For example, each segment or selected snippet can be encoded into a shareable link (deep link) that directly references its position in the source. The shareable link may allow any users to listen to only the associated segments from the start timestamp to the end timestamp. In optional cases, the system may generate temporary or cached video / audio snippets containing only the selected segment. In some cases, the GUI may allow users to select one or more segments, select sentences within a segment (highlight a sentence or speaker turn in the transcript)), assemble sentences or segments, and click a "share" icon. The system may generate a deep link or shareable link for the content selected by the user. In some cases, the GUI may display options for users to copy a timestamped URL, generate a mini-clip, share via social / email, or export transcript + link. The system’s fine-grained indexing architecture — which includes semantic chunking, time-stamped transcripts, and metadata — makes it possible for users to share specific, meaningful moments from large media files (e.g., podcasts, interviews, lectures) with high precision.
[0089] In some cases, the GUI may display concise, human-readable content — such as summaries, highlights, and quotes — from the selected segments, and / or from the longer audio / video material. The system may provide shareable Al-generated summaries, topic highlights, or key quotes that can be exported as text, cards, or social media-ready graphics. The summaries, highlights or quotes may be generated by fine-tuned LLMs.
[0090] In some cases, the GUI may provide embed features. For instance, users may embed clips, summaries, and transcripts into third-party platforms such as blogs, newsletters, or websites with rich media players and context-aware previews. In some cases, the system 500 may support various embed types such as video / audio clip embeds for small player that starts at a specific timestamp, summary widgets for collapsible blocks with title, summary, key highlights, quote embeds for tweet-style quote box with media link or transcript previews for inline text browser with play-on-click functionality. In some cases, the GUI may display options for embed such as one-click “Generate Shareable Card”, auto-translated summaries that automatically translate the summary in a user preferred language, or title suggestions generated by LLM that takes into account of the context of the third-party platform.
[0091] Gesture-based navigational access to audio media content
[0092] The GUIs described herein can provide a hierarchical, multi-dimensional category view of media content (e.g., audio content) to the user in response to various navigational gestures. Systems and methods of the disclosure can provide a gesture-based navigational access to timestamped audio / video media contents aggregated by Al-generated perspectives (e.g., lifestyle, technology, sports, news, etc), tags or media sources (e.g., CNN, Mashable-Tech, DukeGlobal Health Institute) as described above. In some embodiments, navigation may be accomplished using touch-based gestures, such as swipes and taps. Such a method may permit a user to drill down, drill up within the same dimension or across different dimensions, perform search or filtering action on a touch-screen device using a thumb of the user while the user is holding the device. The systems may allow a user to perform a search / filter across dimensions such as words and phrases (i.e., tags or social tags), perspectives (i.e., industries) or media resources. A user is allowed to switch across such dimensions efficiently using user gestures.
[0093] FIGs. 7-8 show examples of GUI allowing users to search, filter or navigate the audio content via different gestures (e.g., swipe left / right, swipe up / down). The GUIs may be provided via a media player. These GUIs may be implemented within a media player application and are configured to support search, filtering, and navigation of media content using various touch or pointer gestures. For example, users may swipe left or right to skip forward or backward to apply filters to search media content, and may be presented with search results in timestamped segments across audio tracks, chapters, or multiple timestamped segments within a track (e.g., in a podcast or audiobook). In some cases, the different gestures may correspond to different dimensions for applying filters. For example, as shown in FIG. 7, a left / right swipe may correspond to filtering along sources dimension within a selected perspective. As shown in the example, swipe right to left may allow users to explore different sources in ‘technology’ perspective / lens. A swipe to left may switch to different sources (e.g., OCDevel, August Bradley, Steve Gibson). A user may begin with selecting a topic (e.g., Bitcoin) within the technology perspective, then swipe left and right to explore different sources. FIG. 8 shows an example of a different gesture (e.g., swipe up / down) corresponding to filtering along the perspective dimension. As shown in the example, swiping up may switch to a different perspective / lens (e.g., Technology, business, politics, etc.) within a selected source. It should be noted that a swipe to an opposite direction may reverse a filter. The different directions may correspond to drill up (e.g., reverse the filter and revert back to previous filter result) or drill down (e.g., apply new filters to further narrow the search) from a current filtering result.
[0094] In some cases, a user may filter across different dimension. FIG. 9 shows an example of switching from filtering along the lens dimension to the source dimension. As shown in the example, a user may switch to a different dimension by performing a different gesture (e.g., left / right swipe along the perspective dimension to up / down swipe along the source dimension). A user may be permitted to switch across the different dimensions from any location or intermediate results of filtering without reversing the filtering. Unlike conventional organization of content using a tree-structure (e.g., category branched or divided into subcategories), the dimensions (e.g., perspective, source) can be orthogonal dimensions — that is, dimensions thatare mutually exclusive such that filtering along a selected dimension can get deeper - drilling down (e.g., applying further new topic filters along the topic dimension) as many detail levels as a user prefers. The orthogonality of the multiple dimensions also permits drilling down a media content from any drill level along any dimension and switching to any other dimensions i.e., different dimensions can be available at any given detail level to further drill down the content at the given level without reverting to previous level of drill.
[0095] FIG. 10 shows an example of audio media content search result presented in the form of timestamped segment. As shown in the example, the user query may be entered via the various navigational gestures described above (e.g., selecting from a system-generated perspective, swiping left / right to switch to different perspectives), and the search result may comprise source, the podcast name, episode, perspective / lens, rationale of the perspective (e.g., why the Al-based real-time highlights and explains timestamp and its relevance in the lens), and a deep link to the timestamp. The GUI may display a media playback icon to play the search result (e.g., segment) with auto-synched transcript. The transcript may highlight the relevant portion and user may scroll through the transcript. In some cases, the result may comprise multiple segments and a user may scroll through the timestamps associated with the multiple segments.
[0096] FIG. 11 shows an example of multilingual content presentation feature of the GUI. The GUI may allow for displaying and sharing audio content in any language. As shown in the example, audio content retrieved from a source in an original language can be automatically translated to a user preferred language. The search result may be displayed including the translated podcast name, country, translated episode name and transcript in user preferred language. The media playback option may allow users to playback the audio segment in the original language or in a user selected language. The system herein may provide multilingual retrieval and playback this feature that allows for cross-lingual search, exploration, and sharing of audio content (e.g., podcasts, interviews, news) by translating both metadata and transcripts into the user’s preferred language. It supports media playback in the original language or an AI- translated version, promoting accessibility, multilingual discovery, and global content interoperability. For instance, the system ingests audio content from global sources in various original languages. In addition to the ingestion methods and indexing processes as described above, the system may perform language detection (e.g., using fastText or compact language detectors), perform transcription via ASR (speech-to-text) in the original language. In some cases, the system may use domain-adapted or fine-tuned translation models trained on podcast and conversational corpora. The rest of the ingestion and indexing can be the same as described above. In some cases, the GUI may provide media playback options for multiple playback modessuch as original language, machine-translated audio or dual display mode (shows translated transcript while playing original audio).Computing system
[0097] The search engine, systems and methods as described above may be implemented by a computing system. Referring to FIG. 2, a block diagram is shown depicting an exemplary machine that includes a computer system 200 (e.g., a processing or computing system) within which a set of instructions can cause a device to perform or execute any one or more of the aspects of the search engine of the present disclosure. The components in FIG. 2 are examples only and do not limit the scope of use or functionality of any hardware, software, embedded logic component, or a combination of two or more such components implementing particular embodiments.
[0098] Computer system 200 may include one or more processors 201, a memory 203, and a storage 208 that communicate with each other, and with other components, via a bus 240. The bus 240 may also link a display 232, one or more input devices 233 (which may, for example, include a keypad, a keyboard, a mouse, a stylus, etc.), one or more output devices 234, one or more storage devices 235, and various tangible storage media 236. All of these elements may interface directly or via one or more interfaces or adaptors to the bus 240. For instance, the various tangible storage media 236 can interface with the bus 240 via storage medium interface 226. Computer system 200 may have any suitable physical form, including but not limited to one or more integrated circuits (ICs), printed circuit boards (PCBs), mobile handheld devices (such as mobile telephones or PDAs), laptop or notebook computers, distributed computer systems, computing grids, or servers.
[0099] Computer system 200 includes one or more processor(s) 201 (e.g., central processing units (CPUs) or general-purpose graphics processing units (GPGPUs)) that carry out functions. Processor(s) 201 optionally contains a cache memory unit 202 for temporary local storage of instructions, data, or computer addresses. Processor(s) 201 are configured to assist in execution of computer readable instructions. Computer system 200 may provide functionality for the components depicted in FIG. 2 as a result of the processor(s) 201 executing non-transitory, processor-executable instructions embodied in one or more tangible computer-readable storage media, such as memory 203, storage 208, storage devices 235, and / or storage medium 236. The computer-readable media may store software that implements particular embodiments, and processor(s) 201 may execute the software. Memory 203 may read the software from one or more other computer-readable media (such as mass storage device(s) 235, 236) or from one or more other sources through a suitable interface, such as network interface 220. The software maycause processor(s) 201 to carry out one or more processes or one or more steps of one or more processes described or illustrated herein. Carrying out such processes or steps may include defining data structures stored in memory 203 and modifying the data structures as directed by the software.
[0100] The memory 203 may include various components (e.g., machine readable media) including, but not limited to, a random-access memory component (e.g., RAM 204) (e.g., static RAM (SRAM), dynamic RAM (DRAM), ferroelectric random-access memory (FRAM), phasechange random access memory (PRAM), etc.), a read-only memory component (e.g., ROM 205), and any combinations thereof. ROM 205 may act to communicate data and instructions unidirectionally to processor(s) 201, and RAM 204 may act to communicate data and instructions bidirectionally with processor(s) 201. ROM 205 and RAM 204 may include any suitable tangible computer-readable media described below. In one example, a basic input / output system 206 (BIOS), including basic routines that help to transfer information between elements within computer system 200, such as during start-up, may be stored in the memory 203.
[0101] Fixed storage 208 is connected bidirectionally to processor(s) 201, optionally through storage control unit 207. Fixed storage 208 provides additional data storage capacity and may also include any suitable tangible computer-readable media described herein. Storage 208 may be used to store operating system 209, executable(s) 210, data 211, applications 212 (application programs), and the like. Storage 208 can also include an optical disk drive, a solid- state memory device (e.g., flash-based systems), or a combination of any of the above. Information in storage 208 may, in appropriate cases, be incorporated as virtual memory in memory 203.
[0102] In one example, storage device(s) 235 may be removably interfaced with computer system 200 (e.g., via an external port connector (not shown)) via a storage device interface 225. Particularly, storage device(s) 235 and an associated machine-readable medium may provide non-volatile and / or volatile storage of machine-readable instructions, data structures, program modules, and / or other data for the computer system 200. In one example, software may reside, completely or partially, within a machine-readable medium on storage device(s) 235. In another example, software may reside, completely or partially, within processor(s) 201.
[0103] Bus 240 connects a wide variety of subsystems. Herein, reference to a bus may encompass one or more digital signal lines serving a common function, where appropriate. Bus 240 may be any of several types of bus structures including, but not limited to, a memory bus, a memory controller, a peripheral bus, a local bus, and any combinations thereof, using any of a variety of bus architectures. As an example and not by way of limitation, such architecturesinclude an Industry Standard Architecture (ISA) bus, an Enhanced ISA (EISA) bus, a Micro Channel Architecture (MCA) bus, a Video Electronics Standards Association local bus (VLB), a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, an Accelerated Graphics Port (AGP) bus, HyperTransport (HTX) bus, serial advanced technology attachment (SATA) bus, and any combinations thereof.
[0104] Computer system 200 may also include an input device 233. In one example, a user of computer system 200 may enter commands and / or other information into computer system 200 via 2input device(s) 233. Examples of an input device(s) 233 include, but are not limited to, an alpha-numeric input device (e.g., a keyboard), a pointing device (e.g., a mouse or touchpad), a touchpad, a touch screen, a multi-touch screen, a joystick, a stylus, a gamepad, an audio input device (e.g., a microphone, a voice response system, etc.), an optical scanner, a video or still image capture device (e.g., a camera), and any combinations thereof. In some embodiments, the input device is a Kinect, Leap Motion, or the like. Input device(s) 233 may be interfaced to bus 240 via any of a variety of input interfaces 223 (e.g., input interface 223) including, but not limited to, serial, parallel, game port, USB, FIREWIRE, THUNDERBOLT, or any combination of the above.
[0105] In particular embodiments, when computer system 200 is connected to network 230, computer system 200 may communicate with other devices, specifically mobile devices and enterprise systems, distributed computing systems, cloud storage systems, cloud computing systems, and the like, connected to network 230. Communications to and from computer system 200 may be sent through network interface 220. For example, network interface 220 may receive incoming communications (such as requests or responses from other devices) in the form of one or more packets (such as Internet Protocol (IP) packets) from network 230, and computer system 200 may store the incoming communications in memory 203 for processing. Computer system 200 may similarly store outgoing communications (such as requests or responses to other devices) in the form of one or more packets in memory 203 and communicated to network 230 from network interface 220. Processor(s) 201 may access these communication packets stored in memory 203 for processing.
[0106] Examples of the network interface 220 include, but are not limited to, a network interface card, a modem, and any combination thereof. Examples of a network 230 or network segment 230 include, but are not limited to, a distributed computing system, a cloud computing system, a wide area network (WAN) (e.g., the Internet, an enterprise network), a local area network (LAN) (e.g., a network associated with an office, a building, a campus or other relatively small geographic space), a telephone network, a direct connection between two computing devices, a peer-to-peer network, and any combinations thereof. A network, such as network 230,may employ a wired and / or a wireless mode of communication. In general, any network topology may be used.
[0107] Information and data can be displayed through a display 232. Examples of a display 232 include, but are not limited to, a cathode ray tube (CRT), a liquid crystal display (LCD), a thin film transistor liquid crystal display (TFT-LCD), an organic liquid crystal display (OLED) such as a passive-matrix OLED (PMOLED) or active-matrix OLED (AMOLED) display, a plasma display, and any combinations thereof. The display 232 can interface to the processor(s) 201, memory 203, and fixed storage 208, as well as other devices, such as input device(s) 233, via the bus 240. The display 232 is linked to the bus 240 via a video interface 222, and transport of data between the display 232 and the bus 240 can be controlled via the graphics control 221. In some embodiments, the display is a video projector. In some embodiments, the display is a head-mounted display (HMD) such as a VR headset. In further embodiments, suitable VR headsets include, by way of non-limiting examples, HTC Vive, Oculus Rift, Samsung Gear VR, Microsoft HoloLens, Razer OSVR, FOVE VR, Zeiss VR One, Avegant Glyph, Freefly VR headset, and the like. In still further embodiments, the display is a combination of devices such as those disclosed herein.
[0108] In addition to a display 232, computer system 200 may include one or more other peripheral output devices 534 including, but not limited to, an audio speaker, a printer, a storage device, and any combinations thereof. Such peripheral output devices may be connected to the bus 240 via an output interface 524. Examples of an output interface 524 include, but are not limited to, a serial port, a parallel connection, a USB port, a FIREWIRE port, a THUNDERBOLT port, and any combinations thereof.
[0109] In addition or as an alternative, computer system 200 may provide functionality as a result of logic hardwired or otherwise embodied in a circuit, which may operate in place of or together with software to execute one or more processes or one or more steps of one or more processes described or illustrated herein. Reference to software in this disclosure may encompass logic, and reference to logic may encompass software. Moreover, reference to a computer- readable medium may encompass a circuit (such as an IC) storing software for execution, a circuit embodying logic for execution, or both, where appropriate. The present disclosure encompasses any suitable combination of hardware, software, or both.
[0110] Those of skill in the art will appreciate that the various illustrative logical blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrativecomponents, blocks, modules, circuits, and steps have been described above generally in terms of their functionality.
[0111] The various illustrative functional features, logical blocks, modules, and circuits described in connection with the embodiments disclosed herein may be implemented or performed with a general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.
[0112] The steps of a method or algorithm described in connection with the embodiments disclosed herein may be embodied directly in hardware, in a software module executed by one or more processor(s), or in a combination of the two. A software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art. An exemplary storage medium is coupled to the processor such the processor can read information from, and write information to, the storage medium. In the alternative, the storage medium may be integral to the processor. The processor and the storage medium may reside in an ASIC. The ASIC may reside in a user terminal. In the alternative, the processor and the storage medium may reside as discrete components in a user terminal.
[0113] In accordance with the description herein, suitable computing devices include, by way of non-limiting examples, server computers, desktop computers, laptop computers, notebook computers, sub-notebook computers, netbook computers, netpad computers, set-top computers, media streaming devices, handheld computers, Internet appliances, mobile smartphones, tablet computers, personal digital assistants, video game consoles, and vehicles. Those of skill in the art will also recognize that select televisions, video players, and digital music players with optional computer network connectivity are suitable for use in the system described herein. Suitable tablet computers, in various embodiments, include those with booklet, slate, and convertible configurations, known to those of skill in the art.
[0114] In some embodiments, the computing device includes an operating system configured to perform executable instructions. The operating system is, for example, software, including programs and data, which manages the device’s hardware and provides services forexecution of applications. Those of skill in the art will recognize that suitable server operating systems include, by way of non-limiting examples, FreeBSD, OpenBSD, NetBSD®, Linux, Apple® Mac OS X Server®, Oracle® Solaris®, Windows Server®, and Novell® NetWare®. Those of skill in the art will recognize that suitable personal computer operating systems include, by way of non-limiting examples, Microsoft® Windows®, Apple® Mac OS X®, UNIX®, and UNIX-like operating systems such as GNU / Linux®. In some embodiments, the operating system is provided by cloud computing. Those of skill in the art will also recognize that suitable mobile smartphone operating systems include, by way of non-limiting examples, Nokia® Symbian® OS, Apple® iOS®, Research In Motion® BlackBerry OS®, Google® Android®, Microsoft® Windows Phone® OS, Microsoft® Windows Mobile® OS, Linux®, and Palm® WebOS®. Those of skill in the art will also recognize that suitable media streaming device operating systems include, by way of non-limiting examples, Apple TV®, Roku®, Boxee®, Google TV®, Google Chromecast®, Amazon Fire®, and Samsung® HomeSync®. Those of skill in the art will also recognize that suitable video game console operating systems include, by way of nonlimiting examples, Sony® PS3®, Sony® PS4®, Microsoft® Xbox 360®, Microsoft Xbox One, Nintendo® Wii®, Nintendo® Wii U®, and Ouya®.Non-transitory computer readable storage medium
[0115] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more non-transitory computer readable storage media encoded with a program including instructions executable by the operating system of an optionally networked computing device. In further embodiments, a computer readable storage medium is a tangible component of a computing device. In still further embodiments, a computer readable storage medium is optionally removable from a computing device. In some embodiments, a computer readable storage medium includes, by way of non-limiting examples, CD-ROMs, DVDs, flash memory devices, solid state memory, magnetic disk drives, magnetic tape drives, optical disk drives, distributed computing systems including cloud computing systems and services, and the like. In some cases, the program and instructions are permanently, substantially permanently, semi-permanently, or non-transitorily encoded on the media.Computer program
[0116] In some embodiments, the platforms, systems, media, and methods disclosed herein include at least one computer program, or use of the same. A computer program includes a sequence of instructions, executable by one or more processor(s) of the computing device’s CPU, written to perform a specified task. Computer readable instructions may be implemented as program modules, such as functions, objects, Application Programming Interfaces (APIs),computing data structures, and the like, that perform particular tasks or implement particular abstract data types. In light of the disclosure provided herein, those of skill in the art will recognize that a computer program may be written in various versions of various languages.
[0117] The functionality of the computer readable instructions may be combined or distributed as desired in various environments. In some embodiments, a computer program comprises one sequence of instructions. In some embodiments, a computer program comprises a plurality of sequences of instructions. In some embodiments, a computer program is provided from one location. In other embodiments, a computer program is provided from a plurality of locations. In various embodiments, a computer program includes one or more software modules. In various embodiments, a computer program includes, in part or in whole, one or more web applications, one or more mobile applications, one or more standalone applications, one or more web browser plug-ins, extensions, add-ins, or add-ons, or combinations thereof.Web application
[0118] In some embodiments, a computer program includes a web application. In light of the disclosure provided herein, those of skill in the art will recognize that a web application, in various embodiments, utilizes one or more software frameworks and one or more database systems. In some embodiments, a web application is created upon a software framework such as Microsoft® .NET or Ruby on Rails (RoR). In some embodiments, a web application utilizes one or more database systems including, by way of non-limiting examples, relational, non-relational, object-oriented, associative, and XML database systems. In further embodiments, suitable relational database systems include, by way of non-limiting examples, Microsoft® SQL Server, mySQL™, and Oracle®. Those of skill in the art will also recognize that a web application, in various embodiments, is written in one or more versions of one or more languages. A web application may be written in one or more markup languages, presentation definition languages, client-side scripting languages, server-side coding languages, database query languages, or combinations thereof. In some embodiments, a web application is written to some extent in a markup language such as Hypertext Markup Language (HTML), Extensible Hypertext Markup Language (XHTML), or extensible Markup Language (XML). In some embodiments, a web application is written to some extent in a presentation definition language such as Cascading Style Sheets (CSS). In some embodiments, a web application is written to some extent in a clientside scripting language such as Asynchronous Javascript and XML (AJAX), Flash® Actionscript, Javascript, or Silverlight®. In some embodiments, a web application is written to some extent in a server-side coding language such as Active Server Pages (ASP), ColdFusion®, Perl, Java™, JavaServer Pages (JSP), Hypertext Preprocessor (PHP), Python™, Ruby, Tel,Smalltalk, WebDNA®, or Groovy. In some embodiments, a web application is written to some extent in a database query language such as Structured Query Language (SQL). In some embodiments, a web application integrates enterprise server products such as IBM® Lotus Domino®. In some embodiments, a web application includes a media player element. In various further embodiments, a media player element utilizes one or more of many suitable multimedia technologies including, by way of non-limiting examples, Adobe® Flash®, HTML 5, Apple® QuickTime®, Microsoft® Silverlight®, Java™, and Unity®.
[0119] Referring to FIG. 3, in a particular embodiment, an application provision system comprises one or more databases 300 accessed by a relational database management system (RDBMS) 310. Suitable RDBMSs include Firebird, MySQL, PostgreSQL, SQLite, Oracle Database, Microsoft SQL Server, IBM DB2, IBM Informix, SAP Sybase, SAP Sybase, Teradata, and the like. In this embodiment, the application provision system further comprises one or more application severs 320 (such as Java servers, .NET servers, PHP servers, and the like) and one or more web servers 330 (such as Apache, IIS, GWS and the like). The web server(s) optionally expose one or more web services via app application programming interfaces (APIs) 340. Via a network, such as the Internet, the system provides browser-based and / or mobile native user interfaces.
[0120] Referring to FIG. 4, in a particular embodiment, an application provision system alternatively has a distributed, cloud-based architecture 400 and comprises elastically load balanced, auto-scaling web server resources 410 and application server resources 420 as well synchronously replicated databases 430.Mobile Application
[0121] In some embodiments, a computer program includes a mobile application provided to a mobile computing device. In some embodiments, the mobile application is provided to a mobile computing device at the time it is manufactured. In other embodiments, the mobile application is provided to a mobile computing device via the computer network described herein. One or more features of the search engine such as the deep link may be provided for mobile applications. For instance, the deep link may be a mobile deep link.
[0122] In view of the disclosure provided herein, a mobile application is created by techniques known to those of skill in the art using hardware, languages, and development environments known to the art. Those of skill in the art will recognize that mobile applications are written in several languages. Suitable programming languages include, by way of nonlimiting examples, C, C++, C#, Objective-C, Java™, Javascript, Pascal, Object Pascal, Python™, Ruby, VB.NET, WML, and XHTML / HTML with or without CSS, or combinations thereof.
[0123] Suitable mobile application development environments are available from several sources. Commercially available development environments include, by way of non-limiting examples, AirplaySDK, alcheMo, Appcelerator®, Celsius, Bedrock, Flash Lite, .NET Compact Framework, Rhomobile, and WorkLight Mobile Platform. Other development environments are available without cost including, by way of non-limiting examples, Lazarus, MobiFlex, MoSync, and Phonegap. Also, mobile device manufacturers distribute software developer kits including, by way of non-limiting examples, iPhone and iPad (iOS) SDK, Android™ SDK, BlackBerry® SDK, BREW SDK, Palm® OS SDK, Symbian SDK, webOS SDK, and Windows® Mobile SDK.
[0124] Those of skill in the art will recognize that several commercial forums are available for distribution of mobile applications including, by way of non-limiting examples, Apple® App Store, Google® Play, Chrome WebStore, BlackBerry® App World, App Store for Palm devices, App Catalog for webOS, Windows® Marketplace for Mobile, Ovi Store for Nokia® devices, Samsung® Apps, and Nintendo® DSi Shop.Standalone Application
[0125] In some embodiments, a computer program includes a standalone application, which is a program that is run as an independent computer process, not an add-on to an existing process, e.g., not a plug-in. Those of skill in the art will recognize that standalone applications are often compiled. A compiler is a computer program(s) that transforms source code written in a programming language into binary object code such as assembly language or machine code. Suitable compiled programming languages include, by way of non-limiting examples, C, C++, Objective-C, COBOL, Delphi, Eiffel, Java™, Lisp, Python™, Visual Basic, and VB .NET, or combinations thereof. Compilation is often performed, at least in part, to create an executable program. In some embodiments, a computer program includes one or more executable complied applications.Web Browser Plug-in
[0126] In some embodiments, the computer program includes a web browser plug-in (e.g., extension, etc.). In computing, a plug-in is one or more software components that add specific functionality to a larger software application. Makers of software applications support plug-ins to enable third-party developers to create abilities which extend an application, to support easily adding new features, and to reduce the size of an application. When supported, plug-ins enable customizing the functionality of a software application. For example, plug-ins are commonly used in web browsers to play video, generate interactivity, scan for viruses, anddisplay particular file types. Those of skill in the art will be familiar with several web browser plug-ins including, Adobe® Flash® Player, Microsoft® Silverlight®, and Apple® QuickTime®. In some embodiments, the toolbar comprises one or more web browser extensions, add-ins, or addons. In some embodiments, the toolbar comprises one or more explorer bars, tool bands, or desk bands.Software Modules
[0127] In some embodiments, the platforms, systems, media, and methods disclosed herein include software, server, and / or database modules, or use of the same. In view of the disclosure provided herein, software modules are created by techniques known to those of skill in the art using machines, software, and languages known to the art. The software modules disclosed herein are implemented in a multitude of ways. In various embodiments, a software module comprises a file, a section of code, a programming object, a programming structure, or combinations thereof. In further various embodiments, a software module comprises a plurality of files, a plurality of sections of code, a plurality of programming objects, a plurality of programming structures, or combinations thereof. In various embodiments, the one or more software modules comprise, by way of non-limiting examples, a web application, a mobile application, and a standalone application. In some embodiments, software modules are in one computer program or application. In other embodiments, software modules are in more than one computer program or application. In some embodiments, software modules are hosted on one machine. In other embodiments, software modules are hosted on more than one machine. In further embodiments, software modules are hosted on a distributed computing platform such as a cloud computing platform. In some embodiments, software modules are hosted on one or more machines in one location. In other embodiments, software modules are hosted on one or more machines in more than one location.Databases
[0128] In some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, those of skill in the art will recognize that many databases are suitable for storage and retrieval of medical imaging information. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object-oriented databases, object databases, entity-relationship model databases, associative databases, and XML databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, and Sybase. In some embodiments, a database is internet-based. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devicesln some embodiments, the platforms, systems, media, and methods disclosed herein include one or more databases, or use of the same. In view of the disclosure provided herein, those of skill in the art will recognize that many databases are suitable for storage and retrieval of medical imaging information. In various embodiments, suitable databases include, by way of non-limiting examples, relational databases, non-relational databases, object-oriented databases, object databases, entity -relationship model databases, associative databases, and XML databases. Further non-limiting examples include SQL, PostgreSQL, MySQL, Oracle, DB2, and Sybase. In some embodiments, a database is internetbased. In further embodiments, a database is web-based. In still further embodiments, a database is cloud computing-based. In a particular embodiment, a database is a distributed database. In other embodiments, a database is based on one or more local computer storage devices.
[0129] While preferred embodiments of the present invention have been shown and described herein, it will be obvious to those skilled in the art that such embodiments are provided by way of example only. It is not intended that the invention be limited by the specific examples provided within the specification. While the invention has been described with reference to the aforementioned specification, the descriptions and illustrations of the embodiments herein are not meant to be construed in a limiting sense. Numerous variations, changes, and substitutions will now occur to those skilled in the art without departing from the invention. Furthermore, it shall be understood that all aspects of the invention are not limited to the specific depictions, configurations or relative proportions set forth herein which depend upon a variety of conditions and variables. It should be understood that various alternatives to the embodiments of the invention described herein may be employed in practicing the invention. It is therefore contemplated that the invention shall also cover any such alternatives, modifications, variations or equivalents. It is intended that the following claims define the scope of the invention and that methods and structures within the scope of these claims and their equivalents be covered thereby.
Claims
CLAIMSWHAT IS CLAIMED IS:
1. A computer-implemented method for querying and presenting audio content, the computer-implemented method comprising:(a) ingesting a plurality of audio files from a plurality of sources;(b) converting the ingested plurality of audio files into indexed data, wherein the indexed data comprises a vector representation of a transcription of a segment within an audio file, timestamps associated with the segment, and extracted metadata;(c) augmenting the indexed data by assigning a category to the respective segment with aid of a large language model (LLM); and(d) displaying, on a graphical user interface (GUI), a first filter dimension corresponding to filtering across the plurality of sources and a second filter dimension corresponding to filtering across a plurality of categories.
2. The computer-implemented method of claim 1, wherein the plurality of audio files comprise a plurality of podcasts.
3. The computer-implemented method of claim 1, wherein ingesting the plurality of audio files comprises using a domain-aware semantic crawler to detect media content in audio-format.
4. The computer-implemented method of claim 3, further comprising assessing a credibility of a source from the plurality of sources using a classification model.
5. The computer-implemented method of claim 1, wherein the plurality of sources comprise public websites, manual uploads and content stream sources.
6. The computer-implemented method of claim 1, wherein converting the ingested plurality of audio files into indexed data comprises chunking an audio file into one or more segments and associating each segment with a start timestamp and an end timestamp.
7. The computer-implemented method of claim 1, wherein the extracted metadata comprises at least one of language, duration, source of the audio file.
8. The computer-implemented method of claim 1, wherein the LLM is fine-tuned using training dataset comprising a label of a category and a rationale for assigning the category.
9. The computer-implemented method of claim 8, further comprising evaluating the LLM’s performance for generating a rationale for assigning a category.
10. The computer-implemented method of claim 1, further comprising receiving a search query via the GUI and identifying one or more relevant segments in response to the search query.
11. The computer-implemented method of claim 10, further comprising sorting the one or more relevant segments based at least in part on a category assigned to each of the one or more relevant segments and one or more metadata fields associated with the one or more relevant segments.
12. The computer-implemented method of claim 10, wherein the search query is received by receiving different gestures corresponding to the first filter dimension and the second filter dimension.
13. The computer-implemented method of claim 10, wherein the one or more relevant segments are identified by converting the search query into a vector format and performing a semantic search by processing the search query in the vector format and the indexed data.
14. The computer-implemented method of claim 11, further comprising displaying, via the GUI, the one or more relevant segments with the associated category and an option to share one or more segments selected from the one or more relevant segments.
15. The computer-implemented method of claim 14, further comprising embedding a deep link to each of the one or more relevant segments for the user to choose to play a selected relevant segment via the GUI.
16. The computer-implemented method of claim 15, further comprising displaying a translated transcription of the one or more metadata fields, a translated transcription in a user preferred language.
17. The computer-implemented method of claim 10, further comprising displaying one or more commerce options within the GUI, wherein the one or more commerce options are identified based at least in part on the search query and the one or more relevant segments.
18. A system for querying and presenting audio content, the system comprising: one or more computer processors are individually or collectively programmed to:(a) ingest a plurality of audio files from a plurality of sources;(b) convert the ingested plurality of audio files into indexed data, wherein the indexed data comprises a vector representation of a transcription of a segment within an audio file, timestamps associated with the segment, and extracted metadata;(c) augment the indexed data by assigning a category to the respective segment with aid of a large language model (LLM); and(d) display, on a graphical user interface (GUI), a first filter dimension corresponding to filtering across the plurality of sources and a second filter dimension corresponding to filtering across a plurality of categories.
19. The system of claim 18, wherein the plurality of audio files comprise a plurality of podcasts.
20. The system of claim 18, wherein the plurality of audio files are ingested using a domain- aware semantic crawler to detect media content in audio-format.
21. The system of claim 18, wherein the plurality of audio files are ingested by assessing a credibility of a source from the plurality of sources using a classification model.
22. The system of claim 18, wherein the plurality of sources comprise public websites, manual uploads and content stream sources.
23. The system of claim 18, wherein the indexed data are generated by chunking an audio file into one or more segments and associating each segment with a start timestamp and an end timestamp.
24. The system of claim 18, wherein the extracted metadata comprises at least one of language, duration, source of the audio file.
25. The system of claim 18, wherein the LLM is fine-tuned using training dataset comprising a label of a category and a rationale for assigning the category.
26. The system of claim 25, wherein the LLM’s performance is evaluated for generating a rationale for assigning a category.
27. The system of claim 18, wherein the one or more computer processors are further programmed to receive a search query via the GUI and identify one or more relevant segments in response to the search query.
28. The system of claim 18, wherein the one or more computer processors are further programmed to sort the one or more relevant segments based at least in part on a category assigned to each of the one or more relevant segments and one or more metadata fields associated with the one or more relevant segments.
29. The system of claim 27, wherein the search query is received by receiving different gestures corresponding to the first filter dimension and the second filter dimension.
30. The system of claim 27, wherein the one or more relevant segments are identified by converting the search query into a vector format and performing a semantic search by processing the search query in the vector format and the indexed data.
31. The system of claim 28, wherein the one or more computer processors are further programmed to display, via the GUI, the one or more relevant segments with the associated category and an option to share one or more segments selected from the one or more relevant segments.
32. The system of claim 31, wherein the one or more computer processors are further programmed to embed a deep link to each of the one or more relevant segments for the user to choose to play a selected relevant segment via the GUI.
33. A computer-implemented method for querying and presenting audio content, the computer-implemented method comprising:(a) providing a graphical user interface (GUI) for searching audio content, wherein the GUI comprises a first filter dimension corresponding to filtering across the plurality of sources and a second filter dimension corresponding to filtering across a plurality of categories associated with the audio content gathered from the plurality of sources;(b) receiving a search query from the GUI for filtering the audio content;(c) identifying one or more relevant segments from the audio content in response to the search query, wherein a segment is converted to an indexed data comprising a vector representation of a transcription of the segment within an audio file, and is associated with timestamps, extracted metadata and an assigned category; and(d) displaying the one or more relevant segments as search result within the GUI, wherein each relevant segment is displayed with the assigned category, a playback option or an option to share the relevant segment.
34. The computer-implemented method of claim 33, wherein the assigned category is generated by a large language model (LLM).
35. The computer-implemented method of claim 34, wherein the LLM is fine-tuned using training dataset comprising a label of a category and a rationale for assigning the category.
36. The computer-implemented method of claim 33, wherein the search query is received by receiving different gestures corresponding to the first filter dimension and the second filter dimension.
Citation Information
Patent Citations
Rich semantic tag data augmentation method based on super-large-scale language model
CN117494760A
Systems and methods for organizing and analyzing audio content derived from media files
US20160004773A1
Artificial intelligence (AI) based data processing
US20200272915A1
Data Processing System and Data Processing Method
US20210248481A1