Methods and systems for facilitating enhanced semantic search of a corpus of content

The method and system leverage machine learning and parallel computing to enhance semantic search in large content repositories, addressing data accessibility and user-centricity challenges, resulting in accurate and efficient retrieval of relevant content.

WO2026019844A1PCT designated stage Publication Date: 2026-01-22LEARNERSHAPE LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2025/037777
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-15
Filing Date
2025-07-15
Publication Date
2026-01-22

AI Technical Summary

Technical Problem

Traditional search systems struggle with retrieving contextually relevant information due to static algorithms, lack of user preference adaptation, and data privacy issues, leading to inconclusive results in vast digital content repositories like YouTube.

Method used

A method and system utilizing flexible machine learning techniques to generate keywords, extract content details, and analyze metadata through restricted APIs, incorporating real-time feedback and parallel computing for refined semantic search.

Benefits of technology

Enhances search accuracy and efficiency by providing relevant, user-centric results tailored to specific criteria, overcoming data accessibility limitations and improving content curation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000047_0000
    Figure 00000047_0000
  • Figure 00000048_0000
    Figure 00000048_0000
  • Figure 00000049_0000
    Figure 00000049_0000
Patent Text Reader

Abstract

The present disclosure provides a method for facilitating enhanced semantic search of a corpus of content. Further, the method may include receiving one or more requests for conducting one or more content searches, analyzing the one or more requests, generating one or more queries, extracting two or more contents using one or more application programming interfaces (API) of one or more applications based on the one or more queries, analyzing the two or more contents, generating two or more information of the two or more contents, analyzing the two or more information of the two or more contents using one or more analysis models, identifying one or more contents from the two or more contents, generating one or more results for the one or more content searches, and transmitting the one or more results to the one or more user devices.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] METHODS AND SYSTEMS EOR FACILITATING ENHANCED SEMANTIC

[0002] SEARCH OF A CORPUS OF CONTENT

[0003] FIELD OF DISCLOSURE

[0004] The present disclosure relates to the field of data processing. More specifically, the present disclosure relates to methods and systems for facilitating enhanced semantic search of a corpus of content.

[0005] BACKGROUND

[0006] The field of content curation and semantic search is pivotal in today's information- driven era, where users increasingly rely on efficient tools to navigate vast amounts of data. As digital content continues to grow exponentially, the ability to locate relevant and valuable information swiftly becomes paramount. Traditional methods of information retrieval often struggle with restricted data access, leading to inconclusive or low-quality results that fail to meet user expectations.

[0007] The objective of the disclosed system is to enhance the capabilities of semantic search systems, making them more efficient, scalable, and accurate across diverse domains such as text, images, audio, and video. Traditional search systems often struggle with tasks like retrieving contextually relevant information, adapting to user preferences, and processing data from multiple sources simultaneously.

[0008] Existing systems often rely on static search algorithms, which do not account for the variability of content or user preferences. Additionally, issues like data privacy compliance further complicate content aggregation efforts, often exposing sensitive data without sufficient protection. The lack of transparency in search algorithms further complicates user trust, as users are often unaware of how rankings and recommendations are determined. The field of data processing is technologically important to several industries, business organizations, and / or individuals. In the age of digital information abundance, large content corpora such as YouTube host vast repositories of multimedia content, ranging from videos to textual descriptions. Effective search and retrieval of relevant information from such repositories has become increasingly challenging due to the sheer volume and diversity of content available. Traditional keyword-based search algorithms often fall short in capturing the nuanced semantic relationships and context inherent in multimedia data. Current approaches typically rely heavily on metadata associated with the content, which may not always be publicly accessible or sufficiently descriptive to enable precise search outcomes. Moreover, the rapid expansion of digital content necessitates scalable and efficient methods for content indexing, retrieval, and relevance assessment. Currently, there is no public access to video metadata, other than by extracting such metadata from individual videos. This restriction on the searchability of videos makes it difficult and inconvenient to search the video corpus to identify videos that are relevant for a particular use. This situation is good for streaming platforms (which make money by promoting engagement rather than from curation), but it is inconvenient for those interested in curation through effective semantic search.

[0009] Therefore, methods, systems, or apparatuses that can overcome these challenges, enabling more robust, efficient, and user-centric semantic search capabilities, are required. Therefore, there is a need for improved methods and systems for facilitating enhanced semantic search of a corpus of content.

[0010] SUMMARY OF DISCLOSURE

[0011] This summary is provided to introduce a selection of concepts in a simplified form, that are further described below in the Detailed Description. This summary is not intended to identify key features or essential features of the claimed subject matter. Nor is this summary intended to be used to limit the claimed subject matter’s scope. The present disclosure provides a method for facilitating enhanced semantic search of a corpus of content. Further, the method may include receiving, using a communication device, one or more requests from one or more user devices. Further, the one or more requests may be for conducting one or more content searches. Further, the method may include analyzing, using a processing device, the one or more requests. Further, the method may include generating, using the processing device, one or more queries based on the analyzing of the one or more requests. Further, the method may include extracting, using the processing device, two or more contents using one or more application programming interfaces (API) of one or more applications based on the one or more queries. Further, the method may include analyzing, using the processing device, the two or more contents. Further, the method may include generating, using the processing device, two or more information of the two or more contents based on the analyzing of the two or more contents. Further, the method may include analyzing, using the processing device, the two or more information of the two or more contents using one or more analysis models. Further, the one or more analysis models may be configured for executing one or more assessments for each of the two or more information of the two or more contents. Further, the method may include identifying, using the processing device, one or more contents from the two or more contents based on the analyzing of the two or more information of the two or more contents. Further, the method may include generating, using the processing device, one or more results for the one or more content searches based on the identifying of the one or more contents. Further, the method may include transmitting, using the communication device, the one or more results to the one or more user devices.

[0012] The present disclosure provides a system for facilitating enhanced semantic search of a corpus of content. Further, the system may include a communication device. Further, the communication device may be configured for receiving one or more requests from one or more user devices. Further, the one or more requests may be for conducting one or more content searches. Further, the communication device may be configured for transmitting one or more results to the one or more user devices. Further, the system may include a processing device communicatively coupled with the communication device. Further, the processing device may be configured for analyzing the one or more requests. Further, the processing device may be configured for generating one or more queries based on the analyzing of the one or more requests. Further, the processing device may be configured for extracting two or more contents using one or more application programming interfaces (API) of one or more applications based on the one or more queries. Further, the processing device may be configured for analyzing the two or more contents. Further, the processing device may be configured for generating two or more information of the two or more contents based on the analyzing of the two or more contents. Further, the processing device may be configured for analyzing the two or more information of the two or more contents using one or more analysis models. Further, the one or more analysis models may be configured for executing one or more assessments for each of the two or more information of the two or more contents. Further, the processing device may be configured for identifying one or more contents from the two or more contents based on the analyzing of the two or more information of the two or more contents. Further, the processing device may be configured for generating the one or more results for the one or more content searches based on the identifying of the one or more contents.

[0013] Both the foregoing summary and the following detailed description provide examples and are explanatory only. Accordingly, the foregoing summary and the following detailed description should not be considered to be restrictive. Further, features or variations may be provided in addition to those set forth herein. For example, embodiments may be directed to various feature combinations and sub-combinations described in the detailed description.

[0014] BRIEF DESCRIPTIONS OF DRAWINGS

[0015] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate various embodiments of the present disclosure. The drawings contain representations of various trademarks and copyrights owned by the Applicants. In addition, the drawings may contain other marks owned by third parties and are being used for illustrative purposes only. All rights to various trademarks and copyrights represented herein, except those belonging to their respective owners, are vested in and the property of the applicants. The applicants retain and reserve all rights in their trademarks and copyrights included herein, and grant permission to reproduce the material only in connection with reproduction of the granted patent and for no other purpose.

[0016] Furthermore, the drawings may contain text or captions that may explain certain embodiments of the present disclosure. This text is included for illustrative, non-limiting, explanatory purposes of certain embodiments detailed in the present disclosure.

[0017] Fig. 1 is an illustration of an online platform 100 consistent with various embodiments of the present disclosure.

[0018] Fig. l is a block diagram of a computing device 200 for implementing the methods disclosed herein, in accordance with some embodiments.

[0019] Fig. 3A illustrates a flowchart of a method 300 for facilitating enhanced semantic search of a corpus of content, in accordance with some embodiments.

[0020] Fig. 3B illustrates a continuation of the flowchart of the method 300 for facilitating enhanced semantic search of a corpus of content, in accordance with some embodiments.

[0021] Fig. 4 illustrates a flowchart of a method 400 for facilitating enhanced semantic search of a corpus of content including analyzing, using the processing device 804, the plurality of relevancy scores, in accordance with some embodiments.

[0022] Fig. 5 illustrates a flowchart of a method 500 for facilitating enhanced semantic search of a corpus of content including identifying, using the processing device 804, a plurality of salient keywords from the at least one request, in accordance with some embodiments.

[0023] Fig. 6 illustrates a flowchart of a method 600 for facilitating enhanced semantic search of a corpus of content including determining, using the processing device 804, at least one thematic keyword associated with the at least one request, in accordance with some embodiments.

[0024] Fig. 7 illustrates a flowchart of a method 700 for facilitating enhanced semantic search of a corpus of content including processing, using the processing device 804, each of the plurality of contents using at least one of the plurality of processing resource segments in a parallel configuration, in accordance with some embodiments.

[0025] Fig. 8 illustrates a block diagram of a system 800 for facilitating enhanced semantic search of a corpus of content, in accordance with some embodiments.

[0026] Fig. 9 is a block diagram of a system 900 for facilitating semantic search of a large corpus of content using machine learning, in accordance with some embodiments.

[0027] Fig. 10 is a flow chart of a method 1000 for facilitating semantic search of a large corpus of content using machine learning, in accordance with some embodiments.

[0028] Fig. 11 is a flow chart of a method 1100 for facilitating semantic search of a large corpus of content using machine learning, in accordance with some embodiments.

[0029] DETAILED DESCRIPTION OF DISCLOSURE

[0030] As a preliminary matter, it will readily be understood by one having ordinary skill in the relevant art that the present disclosure has broad utility and application. As should be understood, any embodiment may incorporate only one or a plurality of the abovedisclosed aspects of the disclosure and may further incorporate only one or a plurality of the above-disclosed features. Furthermore, any embodiment discussed and identified as being “preferred” is considered to be part of a best mode contemplated for carrying out the embodiments of the present disclosure. Other embodiments also may be discussed for additional illustrative purposes in providing a full and enabling disclosure. Moreover, many embodiments, such as adaptations, variations, modifications, and equivalent arrangements, will be implicitly disclosed by the embodiments described herein and fall within the scope of the present disclosure.

[0031] Accordingly, while embodiments are described herein in detail in relation to one or more embodiments, it is to be understood that this disclosure is illustrative and exemplary of the present disclosure, and are made merely for the purposes of providing a full and enabling disclosure. The detailed disclosure herein of one or more embodiments is not intended, nor is to be construed, to limit the scope of patent protection afforded in any claim of a patent issuing here from, which scope is to be defined by the claims and the equivalents thereof. It is not intended that the scope of patent protection be defined by reading into any claim limitation found herein and / or issuing here from that does not explicitly appear in the claim itself.

[0032] Thus, for example, any sequence(s) and / or temporal order of steps of various processes or methods that are described herein are illustrative and not restrictive. Accordingly, it should be understood that, although steps of various processes or methods may be shown and described as being in a sequence or temporal order, the steps of any such processes or methods are not limited to being carried out in any particular sequence or order, absent an indication otherwise. Indeed, the steps in such processes or methods generally may be carried out in various different sequences and orders while still falling within the scope of the present disclosure. Accordingly, it is intended that the scope of patent protection is to be defined by the issued claim(s) rather than the description set forth herein.

[0033] Additionally, it is important to note that each term used herein refers to that which an ordinary artisan would understand such term to mean based on the contextual use of such term herein. To the extent that the meaning of a term used herein — as understood by the ordinary artisan based on the contextual use of such term — differs in any way from any particular dictionary definition of such term, it is intended that the meaning of the term as understood by the ordinary artisan should prevail.

[0034] Furthermore, it is important to note that, as used herein, “a” and “an” each generally denotes “at least one,” but does not exclude a plurality unless the contextual use dictates otherwise. When used herein to join a list of items, “or” denotes “at least one of the items,” but does not exclude a plurality of items of the list. Finally, when used herein to join a list of items, “and” denotes “all of the items of the list.”

[0035] The following detailed description refers to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and the following description to refer to the same or similar elements. While many embodiments of the disclosure may be described, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to the elements illustrated in the drawings, and the methods described herein may be modified by substituting, reordering, or adding stages to the disclosed methods. Accordingly, the following detailed description does not limit the disclosure. Instead, the proper scope of the disclosure is defined by the claims found herein and / or issuing here from. The present disclosure contains headers. It should be understood that these headers are used as references and are not to be construed as limiting upon the subjected matter disclosed under the header.

[0036] The present disclosure includes many aspects and features. Moreover, while many aspects and features relate to, and are described in the context of the disclosed use cases, embodiments of the present disclosure are not limited to use only in this context.

[0037] In general, the method disclosed herein may be performed by one or more computing devices. For example, in some embodiments, the method may be performed by a server computer in communication with one or more client devices over a communication network such as, for example, the Internet. In some other embodiments, the method may be performed by one or more of at least one server computer, at least one client device, at least one network device, at least one sensor and at least one actuator. Examples of the one or more client devices and / or the server computer may include, a desktop computer, a laptop computer, a tablet computer, a personal digital assistant, a portable electronic device, a wearable computer, a smart phone, an Internet of Things (loT) device, a smart electrical appliance, a video game console, a rack server, a super-computer, a mainframe computer, mini -computer, micro-computer, a storage server, an application server (e.g. a mail server, a web server, a real-time communication server, an FTP server, a virtual server, a proxy server, a DNS server etc.), a quantum computer, and so on. Further, one or more client devices and / or the server computer may be configured for executing a software application such as, for example, but not limited to, an operating system (e.g. Windows, Mac OS, Unix, Linux, Android, etc.) in order to provide a user interface (e.g. GUI, touch-screen based interface, voice based interface, gesture based interface etc.) for use by the one or more users and / or a network interface for communicating with other devices over a communication network. Accordingly, the server computer may include a processing device configured for performing data processing tasks such as, for example, but not limited to, analyzing, identifying, determining, generating, transforming, calculating, computing, compressing, decompressing, encrypting, decrypting, scrambling, splitting, merging, interpolating, extrapolating, redacting, anonymizing, encoding and decoding. Further, the server computer may include a communication device configured for communicating with one or more external devices. The one or more external devices may include, for example, but are not limited to, a client device, a third party database, public database, a private database and so on. Further, the communication device may be configured for communicating with the one or more external devices over one or more communication channels. Further, the one or more communication channels may include a wireless communication channel and / or a wired communication channel. Accordingly, the communication device may be configured for performing one or more of transmitting and receiving of information in electronic form. Further, the server computer may include a storage device configured for performing data storage and / or data retrieval operations. In general, the storage device may be configured for providing reliable storage of digital information. Accordingly, in some embodiments, the storage device may be based on technologies such as, but not limited to, data compression, data backup, data redundancy, deduplication, error correction, data finger-printing, role based access control, and so on.

[0038] Further, one or more steps of the method disclosed herein may be initiated, maintained, controlled and / or terminated based on a control input received from one or more devices operated by one or more users such as, for example, but not limited to, an end user, an admin, a service provider, a service consumer, an agent, a broker and a representative thereof. Further, the user as defined herein may refer to a human, an animal or an artificially intelligent being in any state of existence, unless stated otherwise, elsewhere in the present disclosure. Further, in some embodiments, the one or more users may be required to successfully perform authentication in order for the control input to be effective. In general, a user of the one or more users may perform authentication based on the possession of a secret human readable secret data (e.g. username, password, passphrase, PIN, secret question, secret answer etc.) and / or possession of a machine readable secret data (e.g. encryption key, decryption key, bar codes, etc.) and / or or possession of one or more embodied characteristics unique to the user (e.g. biometric variables such as, but not limited to, fingerprint, palm-print, voice characteristics, behavioral characteristics, facial features, iris pattern, heart rate variability, evoked potentials, brain waves, and so on) and / or possession of a unique device (e g. a device with a unique physical and / or chemical and / or biological characteristic, a hardware device with a unique serial number, a network device with a unique IP / MAC address, a telephone with a unique phone number, a smartcard with an authentication token stored thereupon, etc.). Accordingly, the one or more steps of the method may include communicating (e.g. transmitting and / or receiving) with one or more sensor devices and / or one or more actuators in order to perform authentication. For example, the one or more steps may include receiving, using the communication device, the secret human readable data from an input device such as, for example, a keyboard, a keypad, a touch-screen, a microphone, a camera and so on. Likewise, the one or more steps may include receiving, using the communication device, the one or more embodied characteristics from one or more biometric sensors.

[0039] Further, one or more steps of the method may be automatically initiated, maintained and / or terminated based on one or more predefined conditions. In an instance, the one or more predefined conditions may be based on one or more contextual variables. In general, the one or more contextual variables may represent a condition relevant to the performance of the one or more steps of the method. The one or more contextual variables may include, for example, but are not limited to, location, time, identity of a user associated with a device (e.g. the server computer, a client device etc.) corresponding to the performance of the one or more steps, environmental variables (e.g. temperature, humidity, pressure, wind speed, lighting, sound, etc.) associated with a device corresponding to the performance of the one or more steps, physical state and / or physiological state and / or psychological state of the user, physical state (e g. motion, direction of motion, orientation, speed, velocity, acceleration, trajectory, etc.) of the device corresponding to the performance of the one or more steps and / or semantic content of data associated with the one or more users. Accordingly, the one or more steps may include communicating with one or more sensors and / or one or more actuators associated with the one or more contextual variables. For example, the one or more sensors may include, but are not limited to, a timing device (e.g. a real-time clock), a location sensor (e.g. a GPS receiver, a GLONASS receiver, an indoor location sensor etc.), a biometric sensor (e.g. a fingerprint sensor), an environmental variable sensor (e.g. temperature sensor, humidity sensor, pressure sensor, etc.) and a device state sensor (e.g. a power sensor, a voltage / current sensor, a switch-state sensor, a usage sensor, etc. associated with the device corresponding to performance of the or more steps).

[0040] Further, the one or more steps of the method may be performed one or more number of times. Additionally, the one or more steps may be performed in any order other than as exemplarily disclosed herein, unless explicitly stated otherwise, elsewhere in the present disclosure. Further, two or more steps of the one or more steps may, in some embodiments, be simultaneously performed, at least in part. Further, in some embodiments, there may be one or more time gaps between performance of any two steps of the one or more steps.

[0041] Further, in some embodiments, the one or more predefined conditions may be specified by the one or more users. Accordingly, the one or more steps may include receiving, using the communication device, the one or more predefined conditions from one or more and devices operated by the one or more users. Further, the one or more predefined conditions may be stored in the storage device. Alternatively, and / or additionally, in some embodiments, the one or more predefined conditions may be automatically determined, using the processing device, based on historical data corresponding to performance of the one or more steps. For example, the historical data may be collected, using the storage device, from a plurality of instances of performance of the method. Such historical data may include performance actions (e.g. initiating, maintaining, interrupting, terminating, etc.) of the one or more steps and / or the one or more contextual variables associated therewith. Further, machine learning may be performed on the historical data in order to determine the one or more predefined conditions. For instance, machine learning on the historical data may determine a correlation between one or more contextual variables and performance of the one or more steps of the method. Accordingly, the one or more predefined conditions may be generated, using the processing device, based on the correlation.

[0042] Further, one or more steps of the method may be performed at one or more spatial locations. For instance, the method may be performed by a plurality of devices interconnected through a communication network. Accordingly, in an example, one or more steps of the method may be performed by a server computer. Similarly, one or more steps of the method may be performed by a client computer. Likewise, one or more steps of the method may be performed by an intermediate entity such as, for example, a proxy server. For instance, one or more steps of the method may be performed in a distributed fashion across the plurality of devices in order to meet one or more objectives. For example, one objective may be to provide load balancing between two or more devices. Another objective may be to restrict a location of one or more of an input data, an output data and any intermediate data therebetween corresponding to one or more steps of the method. For example, in a client-server environment, sensitive data corresponding to a user may not be allowed to be transmitted to the server computer. Accordingly, one or more steps of the method operating on the sensitive data and / or a derivative thereof may be performed at the client device.

[0043] Overview:

[0044] The present disclosure describes methods and systems for facilitating refined semantic search of a large corpus of content using flexible machine learning methods, where there is no public access to the full corpus of content and its metadata in a structured format, but only through a restricted application programming interface (API).

[0045] Further, an example of this problem, and the context in which the disclosed system is first being practiced, is the YouTube™ corpus of Google™ . A large number of YouTube™ videos are publicly available (some are private), but the publicly available videos are only searchable using the YouTube™ web interface or the YouTube™ Data API. Further, there is no public access to video metadata, other than by extracting such metadata from individual videos. Further, the given restriction on the search ability of YouTube™ videos makes it difficult and inconvenient to search the YouTube™ video corpus to identify videos that are relevant for a particular use. This situation is good for Google™ / YouTube™ (who make money by promoting engagement rather than from curation), but it is inconvenient for those interested in curation through effective semantic search.

[0046] Further, in some embodiments, the present disclosure addresses this problem through the following multi-step data pipeline:

[0047] 1. Keywords for potentially relevant content are generated from information on the semantic search target (e.g., content relevant to a particular university course or other topic).

[0048] 2. The keywords are used to extract a “superset” of potentially relevant content using the API for the corpus. (This works for the YouTube™ Data API, which enables keyword search. Other methods may be required for other APIs). The API may be publicly accessible, or made accessible to specific users upon request / application. Generation of the “superset” generally requires multiple calls to the API.

[0049] 3. Content details are extracted from the individual items in the “superset”, using various techniques. For example, items may have a publicly available description, or the full text of a video can be extracted using any of the various widely available software packages.

[0050] 4. Using the content details extracted in step 3, various machine learning and related semantic search methods are used to identify content items from the “superset” that best match a refined set of user-selected search criteria, which are likely to be related to: a. content relevance; b. content quality; c. content level (e.g., basic or advanced), and other criteria. Further, semantic search methods used in this step may include (a) machine learning methods such as various types of neural networks (including generative models), logistic regression, random forest, gradient-boosting, and other and (b) other statistical methods.

[0051] Further, the data pipeline described above is implemented as a series of connected software components currently implemented in the Python programming language. Further, upon submission of the semantic search target via a web user interface or an application programming interface, the target information is passed to a supervisor process accessible as a REST API. Further, the supervisor process is responsible for processing subsequent steps either via a) sub-components in the same process, b) worker processes running on either the same or different compute resources, or c) third-party services, before returning the curated content to the original requestor. In step 1, summarization and topic modelling are used to generate search keywords or phrases. These can then be augmented with additional keywords with the aim of returning relevant and diverse content. In step 2, the large corpus is then searched via an API to generate candidate content options. Content details are then extracted for each piece of content in step 3. As each piece of content can be processed separately, parallelization can be used to reduce the time required to generate the final curated list of content options. Finally, in step 4, the suitability of each piece of content is assessed, with the highest-scoring, or most relevant, content returned to the requestor. The information returned to the requestor includes both the large corpus metadata and additional information generated during the curation process, such as a relevance score and assessments against other criteria.

[0052] Further, in some embodiments, the disclosed system greatly improves the efficiency of the semantic search of a content corpus accessible only through a restricted API.

[0053] In some embodiments, the present disclosure describes methods and systems for facilitating semantic search of a large corpus of content using machine learning. Further, the disclosed system may allow a refined semantic search of a large corpus of content using flexible machine learning methods, where there is no public access to the full corpus of content and its metadata in a structured format, but only through a restricted application programming interface (API). Further, the disclosed system may be configured for performing a first step of using keywords for potentially relevant content generated from information on the semantic search target (e.g., content relevant to a particular university course or other topic). Further, the disclosed system may be configured for performing a second step of using the keywords to extract a “superset” of potentially relevant content using the API for the corpus. (This works for the YouTube Data API, which enables keyword search. Other methods may be required for other APIs). The API may be publicly -accessible or made accessible to specific users upon request / application. Generation of the “superset” generally requires multiple calls to the API. Further, the disclosed system may be configured for performing a third step of extracting content details from individual items in the “superset”, using various techniques. For example, items may have a publicly available description, or the full text of a video can be extracted using any of various widely-available software packages. Further, the disclosed system may be configured for performing a first step of using the content details based on various machine learning and related semantic search methods to identify content items from the “superset” that best match a refined set of user-selected search criteria, which are likely to be related to (a) content relevance and context, (b) content quality, (c) content level (e.g. basic or advanced), and other criteria. These criteria may be combined to produce an overall content score for ranking content for presentation to the user.

[0054] Further, technical methods used in this step may include (a) machine learning methods such as various types of neural networks (including generative models), logistic regression, random forest, gradient-boosting, and others, and (b) other data science and statistical methods. These methods may be applied differently for different search criteria: for example, content relevance and context assessment may use semantic search methods; content quality assessment may use a weighted combination of various datasets on content quality and / or generative models; and content level assessment may use a weighted combination of methods including term frequency, neural networks and others.

[0055] In some embodiments, the system incorporates advanced machine learning techniques to enhance content curation capabilities, particularly when full corpus access is restricted. Further, the given feature addresses the technical problem of limited data accessibility by enabling refined and relevant search results through flexible ML methods, thereby improving efficiency in information retrieval.

[0056] Further, in some embodiments, to enhance relevance during semantic searches, the system may use context-aware keyword generation. Further, the system generates keywords tailored to specific search targets, such as university courses or topics, thereby improving precision and focus in content selection, which is a significant improvement over conventional, broad-based searches.

[0057] In some embodiments, the system may include API-agnostic content aggregation, allowing the system to gather data from various sources regardless of access restrictions. Further, this feature improves comprehensive content collection, reducing reliance on single sources and enhancing adaptability, which is a notable improvement in data handling for restricted environments.

[0058] Further, in some embodiments, the implementation of API-agnostic content aggregation enables the extraction and structuring of metadata and details from individual items, facilitating deeper analysis. Further, improving data organization and accessibility, addressing the technical problem of raw data management, and enhancing analytical capabilities.

[0059] In some embodiments, the system may utilize various ML techniques for semantic matching, allowing customization of content relevance, quality, and level based on user criteria. This feature significantly improves search accuracy, catering to diverse user needs, which is a key improvement in content matching algorithms.

[0060] Further, by leveraging parallel computing, the system may process large data volumes efficiently, reducing time constraints and enhancing performance. Further, parallel computing tackles the technical challenge of handling extensive datasets, optimizing content processing speed.

[0061] In some embodiments, real-time feedback systems may be integrated to refine curation based on user interactions. Further, the real-time feedback system may enhance content selection by leveraging user behavior and preferences, providing dynamic and adaptive curation.

[0062] In some embodiments, the system may support multiple languages, allowing for broader accessibility and relevance in diverse contexts. Further, multiple language support improves global applicability, addressing the technical challenge of language barriers in content retrieval and management.

[0063] Further, implementing dynamic prioritization based on user behavior can optimize content delivery, ensuring that the most relevant and high-quality items are presented first. Further, dynamic prioritization enhances user experience by delivering more pertinent information efficiently.

[0064] In some embodiments, the system may integrate with advanced Al tools like large language models (LLMs) to enhance semantic understanding and content generation. Further, the integration improves content diversity and accuracy, addressing the technical problem of static or outdated data sources.

[0065] In some embodiments, adaptive privacy controls can be added to ensure compliance with data protection regulations. Further, the adaptive privacy controls improve data security, allowing for safe and responsible content aggregation and sharing. Further, by incorporating these features, the system provides a robust and versatile way of handling diverse data sources, improving search accuracy, and adapting to evolving user needs, thus establishing itself as an innovative solution in content curation and management.

[0066] In some embodiments, the system divides available computing resources into multiple processing resource segments. Each segment may represent a distinct allocation of processor cores, virtual machines, threads, containers, or other units of computing capacity. These segments can be assigned to different contents for parallel processing. By dividing processing in this manner, the system achieves faster overall throughput when analyzing large volumes of data or content.

[0067] To optimize how the processing resources are used, the system may monitor a complexity parameter associated with each content. As used herein, a complexity parameter may include, but is not limited to: The length of the content (e.g., number of words in a video transcript or text file), The file size of the content; The number of detected keywords or semantic entities identified during keyword extraction or topic modelling, An estimated processing time, which may be determined using heuristic rules, machine learning models, or historical processing statistics.

[0068] The complexity parameter enables the system to predict how much computing capacity a given content is likely to require.

[0069] In some implementations, the system dynamically adjusts the processing capacity allocated to each processing resource segment based on the monitored complexity parameter. For example, larger or more complex content may be allocated more computing resources (e.g., more CPU cores or memory), while smaller or simpler content is assigned fewer resources. The system may automatically reassign content items among segments to balance the workload and avoid bottlenecks. This dynamic reallocation can occur continuously or at predefined intervals during the processing workflow.

[0070] In some embodiments, the system allows a user to specify a priority level for certain content or batches. For example, a user might indicate that recent videos relevant to a breaking news event should be processed first. The system uses this priority information to override the default processing order or assignment. Higher-priority content can be assigned to faster or less-burdened resource segments to reduce latency.

[0071] Further, the processing device may include any computing system and / or computing device, including physical servers, virtual machines, cloud compute nodes, or combinations thereof, configured to execute the steps described. Further, the communication device may include any computing system and / or computing device comprising network-enabled component capable of sending or receiving data (e.g., network interface card, HTTP client). Further, the processing resource segment may be a logical or physical subdivision of the total computing capacity, which may include one or more processor cores, threads, containers, or virtual instances. Further, the reassign may include an operation of moving a content from one processing resource segment to another to better balance the workload. Further, the complexity parameter may be a measurable or estimated indicator of how computationally demanding it will be to process a given content item.

[0072] Further, the disclosed system facilitates enhanced semantic search of a corpus of content by receiving user requests, generating keyword-based and semantically enriched queries, extracting content using restricted or open APIs, analyzing the extracted content with machine learning models, and returning curated results with relevance scoring and filtering. The system may use one or more processing devices and communication devices, which include any combination of servers, cloud instances, network interfaces, or client-server infrastructure capable of performing the operations.

[0073] The system receives at least one request from at least one user device using a communication device, such as an internet-connected interface or a REST API. A request may include keywords, free-form text, or structured search criteria. The system analyzes the request using natural language processing (NLP) techniques to extract intent, context, or constraints. The system generates at least one query based on the request analysis. This query may include one or more one or more of keywords, phrases, thematic or contextual expansions (e.g., synonyms, related topics). The query is structured to comply with the syntax and endpoint specifications of the target corpus’s API (such as the YouTube™ Data API). The query may consist of multiple API calls if needed to cover a broad superset of potential content. The system uses the generated query to extract a plurality of contents using at least one API. The API may expose endpoints for searching, listing, or retrieving metadata about individual content items. For example, for video content, the API may return video IDs, titles, descriptions, and other metadata. The system analyzes the extracted content to derive information, such as: Transcripts, Extracted text, Tags or keywords, and Metadata such as upload date, length, or popularity metrics. The analysis may include NLP tasks, topic modeling, summarization, keyword extraction, or embedding generation. Further, the one or more analysis models may include natural language processing (NLP) models. Further, the analysis model may be implemented using neural networks, statistical models, or rule-based methods.

[0074] In some embodiments, the analysis of the content information uses pre-defined scoring criteria, such as keyword density, semantic similarity, or topic relevance. The system determines relevancy scores representing how well each content matches the search request. These scores are stored and used to rank and filter content candidates. Further, the relevancy score is a numerical or categorical value indicating the estimated suitability of a content item for the user’s request.

[0075] In some embodiments, the users may specify one or more user-defined search criteria when submitting a request. Further, the one or more user-defined search criteria may include a desired video length range, a preferred language, a content level (beginner, intermediate, advanced), and a date range or freshness.

[0076] In some embodiments, the system may execute topic modeling (e.g., using LDA, BERT -based clustering, or other algorithms) on the request to detect thematic keywords that describe the user’s area of interest. These thematic keywords are added to the query to broaden or refine the search.

[0077] In some embodiments, the queries sent to the API are formatted as API calls that comply with endpoint specifications (e.g., HTTP requests with proper headers, authentication tokens, query parameters). The system may construct these calls dynamically based on the API’s documentation. Further, the API endpoint refers to a specific Uniform Resource Locator (URL), Uniform Resource Identifier (URI), or designated network address within an Application Programming Interface (API) that is configured to receive one or more structured requests and return one or more responses according to predefined rules and protocols. Further, the API endpoint defines a logical entry point for accessing a particular function or resource exposed by an application (i.e., software application, web based application, etc.), service, or database.

[0078] In further embodiments, the analysis model is configured to execute the assessments on the information associated with the contents extracted from the target corpus via the restricted API. As used herein, the assessments refer to a computational operation or process performed by the analysis model to evaluate the characteristics, relevance, quality, or suitability of each content item relative to a user’s request and defined search criteria. For example, the analysis model may execute a relevance assessment to assess the relevance of each content item by comparing the extracted textual or metadata information against the original user request, including any keywords, semantic context, or user-defined criteria. This relevance assessment may involve measuring keyword overlap or semantic similarity using vector embedding, applying natural language inference or semantic matching models, or calculating similarity scores with models such as cosine similarity, BERT -based similarity, or other machine learning algorithms. The output of this relevance assessment typically includes a relevancy score that represents how closely a given content item aligns with the search target.

[0079] Further, the analysis model may execute a content quality assessment to assess the overall quality of the content to determine whether it meets certain standards of usefulness, production value, or trustworthiness. This content quality assessment may include checking the audio or video resolution and clarity when applicable, evaluating the completeness or coherence of extracted transcripts, or detecting undesirable features such as spam, misleading information, or other indicators of low credibility. Such quality assessments can be implemented using statistical models, machine learning classifiers, or heuristic rules configured to flag or filter content accordingly.

[0080] For content intended for educational or professional use, the analysis model may further execute an assessment to assess the level or depth of the content to match it appropriately to a target audience. For example, the model may classify whether the content is introductory, intermediate, or advanced, using topic complexity measures, vocabulary analysis, or trained classification models based on labeled examples. This enables the system to deliver results that align with a user’s knowledge level or instructional need.

[0081] Further, the analysis model may execute additional assessments to determine other factors relevant to content selection. These additional assessments may include estimating how recent or up-to-date the content is by evaluating publication date metadata or other content signals, analyzing engagement metrics such as view counts, likes, or user ratings if made available through the APT, and measuring whether the content sufficiently addresses the intended topic scope. The assessments may be executed sequentially or in parallel, and the resulting scores, labels, or flags may be combined to inform the step of identifying the most suitable content items for inclusion in the final curated results. The analysis model may include pre-trained or fine-tuned neural networks such as BERT, GPT, or other domain-specific transformers, as well as classical machine learning models such as logistic regression or random forests. In some embodiments, rule-based or statistical models may be used for lightweight checks or to complement assessments that are more complex. Multiple models may be orchestrated in a pipeline configuration, wherein the output of one assessment feeds into another or into an ensemble decision process that integrates multiple assessment results.

[0082] FIG. 1 is an illustration of an online platform 100 consistent with various embodiments of the present disclosure. By way of non-limiting example, the online platform 100 may be hosted on a centralized server 102, such as, for example, a cloud computing service. The centralized server 102 may communicate with other network entities, such as, for example, a mobile device 106 (such as a smartphone, a laptop, a tablet computer etc.), other electronic devices 110 (such as desktop computers, server computers etc.), databases 114, and sensors 116 over a communication network 104, such as, but not limited to, the Internet. Further, users of the online platform 100 may include relevant parties such as, but not limited to, end-users, administrators, service providers, service consumers and so on. Accordingly, in some instances, electronic devices operated by the one or more relevant parties may be in communication with the platform.

[0083] A user 112, such as the one or more relevant parties, may access online platform 100 through a web based software application or browser. The web based software application may be embodied as, for example, but not be limited to, a website, a web application, a desktop application, and a mobile application compatible with a computing device 200.

[0084] With reference to FIG. 2, a system consistent with an embodiment of the disclosure may include a computing device or cloud service, such as computing device 200. In a basic configuration, computing device 200 may include at least one processing unit 202 and a system memory 204. Depending on the configuration and type of computing device, system memory 204 may comprise, but is not limited to, volatile (e.g. randomaccess memory (RAM)), non-volatile (e.g. read-only memory (ROM)), flash memory, or any combination. System memory 204 may include operating system 205, one or more programming modules 206, and may include a program data 207. Operating system 205, for example, may be suitable for controlling computing device 200’s operation. In one embodiment, programming modules 206 may include image-processing module, machine learning module. Furthermore, embodiments of the disclosure may be practiced in conjunction with a graphics library, other operating systems, or any other application program and is not limited to any particular application or system. This basic configuration is illustrated in FIG. 2 by those components within a dashed line 208.

[0085] Computing device 200 may have additional features or functionality. For example, computing device 200 may also include additional data storage devices (removable and / or non-removable) such as, for example, magnetic disks, optical disks, or tape. Such additional storage is illustrated in FIG. 2 by a removable storage 209 and a nonremovable storage 210. Computer storage media may include volatile and non-volatile, removable and non-removable media implemented in any method or technology for storage of information, such as computer-readable instructions, data structures, program modules, or other data. System memory 204, removable storage 209, and non-removable storage 210 are all computer storage media examples (i.e., memory storage.) Computer storage media may include, but is not limited to, RAM, ROM, electrically erasable readonly memory (EEPROM), flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium which can be used to store information and which can be accessed by computing device 200. Any such computer storage media may be part of device 200. Computing device 200 may also have input device(s) 212 such as a keyboard, a mouse, a pen, a sound input device, a touch input device, a location sensor, a camera, a biometric sensor, etc. Output device(s) 214 such as a display, speakers, a printer, etc. may also be included. The aforementioned devices are examples and others may be used. Computing device 200 may also contain a communication connection 216 that may allow device 200 to communicate with other computing devices 218, such as over a network in a distributed computing environment, for example, an intranet or the Internet. Communication connection 216 is one example of communication media. Communication media may typically be embodied by computer readable instructions, data structures, program modules, or other data in a modulated data signal, such as a carrier wave or other transport mechanism, and includes any information delivery media. The term “modulated data signal” may describe a signal that has one or more characteristics set or changed in such a manner as to encode information in the signal. By way of example, and not limitation, communication media may include wired media such as a wired network or direct-wired connection, and wireless media such as acoustic, radio frequency (RF), infrared, and other wireless media. The term computer readable media as used herein may include both storage media and communication media.

[0086] As stated above, a number of program modules and data files may be stored in system memory 204, including operating system 205. While executing on processing unit 202, programming modules 206 (e.g., application 220 such as a media player) may perform processes including, for example, one or more stages of methods, algorithms, systems, applications, servers, databases as described above. The aforementioned process is an example, and processing unit 202 may perform other processes. Other programming modules that may be used in accordance with embodiments of the present disclosure may include machine learning applications.

[0087] Generally, consistent with embodiments of the disclosure, program modules may include routines, programs, components, data structures, and other types of structures that may perform particular tasks or that may implement particular abstract data types. Moreover, embodiments of the disclosure may be practiced with other computer system configurations, including hand-held devices, general purpose graphics processor-based systems, multiprocessor systems, microprocessor-based or programmable consumer electronics, application specific integrated circuit-based electronics, minicomputers, mainframe computers, and the like. Embodiments of the disclosure may also be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.

[0088] Furthermore, embodiments of the disclosure may be practiced in an electrical circuit comprising discrete electronic elements, packaged or integrated electronic chips containing logic gates, a circuit utilizing a microprocessor, or on a single chip containing electronic elements or microprocessors. Embodiments of the disclosure may also be practiced using other technologies capable of performing logical operations such as, for example, AND, OR, and NOT, including but not limited to mechanical, optical, fluidic, and quantum technologies. In addition, embodiments of the disclosure may be practiced within a general-purpose computer or in any other circuits or systems.

[0089] Embodiments of the disclosure, for example, may be implemented as a computer process (method), a computing system, or as an article of manufacture, such as a computer program product or computer readable media. The computer program product may be a computer storage media readable by a computer system and encoding a computer program of instructions for executing a computer process. The computer program product may also be a propagated signal on a carrier readable by a computing system and encoding a computer program of instructions for executing a computer process. Accordingly, the present disclosure may be embodied in hardware and / or in software (including firmware, resident software, micro-code, etc.). In other words, embodiments of the present disclosure may take the form of a computer program product on a computer-usable or computer-readable storage medium having computer-usable or computer-readable program code embodied in the medium for use by or in connection with an instruction execution system. A computer-usable or computer-readable medium may be any medium that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device.

[0090] The computer-usable or computer-readable medium may be, for example but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, device, or propagation medium. More specific computer-readable medium examples (a non-exhaustive list), the computer-readable medium may include the following: an electrical connection having one or more wires, a portable computer diskette, a random-access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CD-ROM). Note that the computer-usable or computer-readable medium could even be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, via, for instance, optical scanning of the paper or other medium, then compiled, interpreted, or otherwise processed in a suitable manner, if necessary, and then stored in a computer memory.

[0091] Embodiments of the present disclosure, for example, are described above with reference to block diagrams and / or operational illustrations of methods, systems, and computer program products according to embodiments of the disclosure. The functions / acts noted in the blocks may occur out of the order as shown in any flowchart. For example, two blocks shown in succession may in fact be executed substantially concurrently or the blocks may sometimes be executed in the reverse order, depending upon the functionality / acts involved.

[0092] While certain embodiments of the disclosure have been described, other embodiments may exist. Furthermore, although embodiments of the present disclosure have been described as being associated with data stored in memory and other storage mediums, data can also be stored on or read from other types of computer-readable media, such as secondary storage devices, like hard disks, solid state storage (e.g., USB drive), or a CD- ROM, a carrier wave from the Internet, or other forms of RAM or ROM. Further, the disclosed methods’ stages may be modified in any manner, including by reordering stages and / or inserting or deleting stages, without departing from the disclosure.

[0093] Fig. 3A and Fig. 3B illustrate a flowchart of a method 300 for facilitating enhanced semantic search of a corpus of content, in accordance with some embodiments.

[0094] Accordingly, the method 300 may include a step 302 of receiving, using a communication device 802, one or more requests from one or more user devices 808. Further, the one or more requests may be for conducting one or more content searches. Further, the method 300 may include a step 304 of analyzing, using a processing device 804, the one or more requests. Further, the method 300 may include a step 306 of generating, using the processing device 804, one or more queries based on the analyzing of the one or more requests. Further, the method 300 may include a step 308 of extracting, using the processing device 804, two or more contents using one or more application programming interfaces (API) of one or more applications based on the one or more queries. Further, the method 300 may include a step 310 of analyzing, using the processing device 804, the two or more contents. Further, the method 300 may include a step 312 of generating, using the processing device 804, two or more information of the two or more contents based on the analyzing of the two or more contents. Further, the method 300 may include a step 314 of analyzing, using the processing device 804, the two or more information of the two or more contents using one or more analysis models. Further, the one or more analysis models may be configured for executing one or more assessments for each of the two or more information of the two or more contents. Further, the method 300 may include a step 316 of identifying, using the processing device 804, one or more contents from the two or more contents based on the analyzing of the two or more information of the two or more contents. Further, the method 300 may include a step 318 of generating, using the processing device 804, one or more results for the one or more content searches based on the identifying of the one or more contents. Further, the method 300 may include a step 320 of transmitting, using the communication device 802, the one or more results to the one or more user devices 808. Further, the one or more results may include the one or more contents, and the one or more information of the one or more contents. Further, the one or more contents may include audio content, textual content, video content, audio video content, multimedia content, etc.

[0095] Fig. 4 illustrates a flowchart of a method 400 for facilitating enhanced semantic search of a corpus of content including analyzing, using the processing device 804, the plurality of relevancy scores, in accordance with some embodiments.

[0096] Further, in some embodiments, the analyzing of the two or more information of the two or more contents may include analyzing the two or more information of the two or more contents using one or more pre-defined scoring criteria. Further, the method 400 further may include a step 402 of determining, using the processing device 804, two or more relevancy scores for the two or more contents based on the analyzing the two or more information of the two or more contents using the one or more pre-defined scoring criteria. Further, each of the two or more relevancy scores represents a suitability of each of the two or more contents in relation to the one or more requests. Further, the method 400 further may include a step 404 of analyzing, using the processing device 804, the two or more relevancy scores. Further, the identifying of the one or more contents from the two or more contents may be further based on the analyzing of the two or more relevancy scores.

[0097] In some embodiments, the one or more requests include one or more user-defined search criteria for facilitating the conducting of the one or more content searches. Further, the analyzing of the two or more information of the two or more contents include analyzing the two or more information of the two or more contents using the one or more user-defined search criteria. Further, the identifying of the one or more contents may be further based on the analyzing of the two or more information of the two or more contents using the one or more user-defined search criteria.

[0098] Fig. 5 illustrates a flowchart of a method 500 for facilitating enhanced semantic search of a corpus of content including identifying, using the processing device 804, a plurality of salient keywords from the at least one request, in accordance with some embodiments.

[0099] Further, in some embodiments, the method 500 further may include a step 502 of executing, using the processing device 804, one or more summarization operations on the one or more requests. Further, in some embodiments, the method 500 further may include a step 504 of identifying, using the processing device 804, two or more salient keywords from the one or more requests based on the executing of the one or more summarization operations on the one or more requests. Further, the generating of the one or more queries may be further based on the identifying of the two or more salient keywords.

[0100] Fig. 6 illustrates a flowchart of a method 600 for facilitating enhanced semantic search of a corpus of content including determining, using the processing device 804, at least one thematic keyword associated with the at least one request, in accordance with some embodiments.

[0101] Further, in some embodiments, the method 600 further may include a step 602 of executing, using the processing device 804, one or more topic modelling operations on the one or more requests. Further, in some embodiments, the method 600 further may include a step 604 of determining, using the processing device 804, one or more thematic keywords associated with the one or more requests based on the executing of the one or more topic modelling operations. Further, the generating of the one or more queries may be further based on the determining of the one or more thematic keywords.

[0102] In some embodiments, the method 300 may further include obtaining, using the processing device 804, one or more additional keywords based on the analyzing of the one or more requests. Further, the generating of the one or more queries may be further based on the one or more additional keywords. Further, the generating of the one or more queries includes augmenting one or more keywords comprised in the one or more queries with the one or more additional keywords.

[0103] Fig. 7 illustrates a flowchart of a method 700 for facilitating enhanced semantic search of a corpus of content including processing, using the processing device 804, each of the plurality of contents using at least one of the plurality of processing resource segments in a parallel configuration, in accordance with some embodiments.

[0104] Further, in some embodiments, the method 700 further may include a step 702 of dividing, using the processing device 804, a processing resource into two or more processing resource segments. Further, each of the two or more processing resource segments may be associated with a processing capacity. Further, two or more processing capacities of the two or more processing resource segments may be same. Further, two or more processing capacities of the two or more processing resource segments may be different. Further, two or more processing capacities of the two or more processing resource segments may be in one or more proportions of one or more of the two or more processing capacities. Further, in some embodiments, the method 700 further may include a step 704 of assigning, using the processing device 804, each of the two or more contents to a respective one of the two or more processing resource segments. Further, in some embodiments, the method 700 further may include a step 706 of processing, using the processing device 804, each of the two or more contents using the respective one of the two or more processing resource segments in a parallel configuration based on the assigning. Further, the generating of the two or more information of the two or more contents may be further based on the processing of each of the two or more contents using the respective one of the two or more processing resource segments in the parallel configuration.

[0105] In further embodiments, the method 700 may include determining, using the processing device 804, a complexity parameter associated with each of the plurality of contents. Further, the complexity parameter may include at least one of a length of a content, a file size of the content, a number of detected keywords for the content, and an estimated processing time for the content. Further, the method 700 may include adjusting, using the processing device 804, an allocation of the processing capacity for each of the plurality of processing resource segments based on the complexity parameter. Further, the method 700 may include reassigning, using the processing device 804, one or more of the two or more contents among the two or more processing resource segments based on the allocation of the processing capacity to improve overall processing efficiency in the generating of the two or more information.

[0106] In an embodiment, the method 700 may include prioritizing, using the processing device 804, one or more of the two or more contents for processing based on a user- defined priority level. Further, the priority level is used to override an assignment of one or more of the two or more contents to one or more of the two or more processing resource segments.

[0107] In some embodiments, the one or more queries include one or more API calls structured to comply with one or more endpoint specifications associated with the one or more APIs. Further, the extracting of the two or more contents may be further based on the one or more API calls. In some embodiments, the method 300 may further include determining, using the processing device 804, one or more attributes associated with each of the two or more contents based on the analyzing of the two or more information of the two or more contents using the one or more user-defined search criteria. Further, the one or more attributes include one or more of a relevance, a quality, and a level associated with each of the two or more contents. Further, the identifying of the one or more contents may be further based on the determining of the one or more attributes.

[0108] In some embodiments, the method 300 may further include transmitting, using the communication device 802, the one or more API calls to one or more endpoints associated with the one or more APIs. Further, the extracting of the two or more contents may be further based on the transmitting of the one or more API calls to the one or more endpoints.

[0109] Fig. 8 illustrates a block diagram of a system 800 for facilitating enhanced semantic search of a corpus of content, in accordance with some embodiments.

[0110] Accordingly, the system 800 may include a communication device 802. Further, the communication device 802 may be configured for receiving one or more requests from one or more user devices 808. Further, the one or more requests may be for conducting one or more content searches. Further, the communication device 802 may be configured for transmitting one or more results to the one or more user devices 808. Further, the system 800 may include a processing device 804 communicatively coupled with the communication device 802. Further, the processing device 804 may be configured for analyzing the one or more requests. Further, the processing device 804 may be configured for generating one or more queries based on the analyzing of the one or more requests. Further, the processing device 804 may be configured for extracting two or more contents using one or more application programming interfaces (API) of one or more applications based on the one or more queries. Further, the processing device 804 may be configured for analyzing the two or more contents. Further, the processing device 804 may be configured for generating two or more information of the two or more contents based on the analyzing of the two or more contents. Further, the processing device 804 may be configured for analyzing the two or more information of the two or more contents using one or more analysis models. Further, the one or more analysis models may be configured for executing one or more assessments for each of the two or more information of the two or more contents. Further, the processing device 804 may be configured for identifying one or more contents from the two or more contents based on the analyzing of the two or more information of the two or more contents. Further, the processing device 804 may be configured for generating the one or more results for the one or more content searches based on the identifying of the one or more contents.

[0111] Further, in some embodiments, the analyzing of the two or more information of the two or more contents may include analyzing the two or more information of the two or more contents using one or more pre-defined scoring criteria. Further, the processing device 804 may be further configured for determining two or more relevancy scores for the two or more contents based on the analyzing the two or more information of the two or more contents using the one or more pre-defined scoring criteria. Further, each of the two or more relevancy scores represents a suitability of each of the two or more contents in relation to the one or more requests. Further, the processing device 804 may be further configured for analyzing the two or more relevancy scores. Further, the identifying of the one or more contents from the two or more contents may be further based on the analyzing of the two or more relevancy scores.

[0112] In some embodiments, the one or more requests include one or more user-defined search criteria for facilitating the conducting of the one or more content searches. Further, the analyzing of the two or more information of the two or more contents includes analyzing the two or more information of the two or more contents using the one or more user-defined search criteria. Further, the identifying of the one or more contents may be further based on the analyzing of the two or more information of the two or more contents using the one or more user-defined search criteria.

[0113] Further, in some embodiments, the processing device 804 may be further configured for executing one or more summarization operations on the one or more requests. Further, the processing device 804 may be further configured for identifying two or more salient keywords from the one or more requests based on the executing of the one or more summarization operations on the one or more requests. Further, the generating of the one or more queries may be further based on the identifying of the two or more salient keywords.

[0114] Further, in some embodiments, the processing device 804 may be further configured for executing one or more topic modelling operations on the one or more requests. Further, the processing device 804 may be further configured for determining one or more thematic keywords associated with the one or more requests based on the executing of the one or more topic modelling operations. Further, the generating of the one or more queries may be further based on the determining of the one or more thematic keywords.

[0115] In some embodiments, the processing device 804 may be further configured for obtaining one or more additional keywords based on the analyzing of the one or more requests. Further, the generating of the one or more queries may be further based on the one or more additional keywords. Further, the generating of the one or more queries includes augmenting one or more keywords comprised in the one or more queries with the one or more additional keywords.

[0116] Further, in some embodiments, the processing device 804 may be further configured for dividing a processing resource into two or more processing resource segments. Further, each of the two or more processing resource segments may be associated with a processing capacity. Further, the processing device 804 may be further configured for assigning each of the two or more contents to a respective one of the two or more processing resource segments. Further, the processing device 804 may be further configured for processing each of the two or more contents using the respective one of the two or more processing resource segments in a parallel configuration based on the assigning. Further, the generating of the two or more information of the two or more contents may be further based on the processing of each of the two or more contents using the respective one of the two or more processing resource segments in the parallel configuration.

[0117] In some embodiments, the one or more queries include one or more API calls structured to comply with one or more endpoint specifications associated with the one or more APIs. Further, the extracting of the two or more contents may be further based on the one or more API calls.

[0118] In some embodiments, the processing device 804 may be further configured for determining one or more attributes associated with each of the two or more contents based on the analyzing of the two or more information of the two or more contents using the one or more user-defined search criteria. Further, the one or more attributes include one or more of a relevance, a quality, and a level associated with each of the two or more contents. Further, the identifying of the one or more contents may be further based on the determining of the one or more attributes.

[0119] In some embodiments, the communication device 802 may be further configured for transmitting the one or more API calls to one or more endpoints associated with the one or more APIs. Further, the extracting of the two or more contents may be further based on the transmitting of the one or more API calls to the one or more endpoints.

[0120] In some embodiments, the level associated with each of the two or more information includes one or more of a basic level and an advanced level.

[0121] In some embodiments, the one or more APIs include at least one multimedia contentretrieval API. Further, the two or more content associated with the at least one multimedia content-retrieval API include one or more video corpora representing a structured-collection of two or more video data associated with two or more videos. Further, the two or more information associated with the one or more video corpora include one or more of a video title and a video description associated with each of the two or more videos.

[0122] In some embodiments, the two or more contents include one or more contentsupersets associated with the two or more contents. Further, the analyzing of the two or more contents includes analyzing of the content-superset. Further, the generating of the two or more information may be further based on the analyzing of the content-superset.

[0123] In some embodiments, the two or more contents include two or more request-based metadata associated with the two or more contents. In some embodiments, the one or more analysis models include one or more machine learning (ML) models. Further, the one or more ML models include one or more of a neural network and gradient boosting.

[0124] In some embodiments, the one or more analysis models may include one or more machine learning models comprising one or more neural networks, one or more generative models, a logistic regression model, a random forest model, and a gradientboosting model, and one or more statistical models.

[0125] In some embodiments, the generating of the two or more information of the two or more contents may be further based on one or more external software modules.

[0126] In some embodiments, the method 300 may further include storing, using a storage device 806, each instance of the one or more results with a corresponding instance of the one or more requests to maintain each of the one or more results and the one or more requests in an associated manner.

[0127] In some embodiments, the generating of the two or more relevancy scores for the two or more contents includes generating one or more relevancy scores for the one or more contents. Further, the one or more results for the one or more content searches include the one or more relevancy scores.

[0128] Further, in some embodiments, the analyzing of the two or more information of the two or more contents using the one or more analysis models may include extracting two or more semantic features from each of the two or more information. Further, the analyzing of the two or more information of the two or more contents using the one or more analysis models may include analyzing the two or more semantic features of each of the two or more information using one or more logistic regression models comprised in the one or more analysis models. Further, the analyzing of the two or more information of the two or more contents using the one or more analysis models may include generating a relevance probability for each of the two or more information based on the analyzing of the two or more semantic features of each of the two or more information using the one or more logistic regression models. Further, the identifying of the one or more contents from the two or more contents may be further based on the relevance probability associated with each of the two or more information.

[0129] Further, in some embodiments, the analyzing of the two or more information of the two or more contents using the one or more analysis models may include generating a semantic feature vector for each of the two or more information. Further, the semantic feature vector represents a semantic relation associated with two or more informationkeywords comprised in each of the two or more information. Further, the analyzing of the two or more information of the two or more contents using the one or more analysis models may include analyzing the semantic feature vector of each of the two or more information using a random forest model comprised in the one or more analysis models. Further, the analyzing of the two or more information of the two or more contents using the one or more analysis models may include generating a plurality relevance scores for each of two or more information using two or more decision trees associated with the random forest model based on the analyzing of the semantic feature vector of each of the two or more information. Further, the analyzing of the two or more information of the two or more contents using the one or more analysis models may include aggregating the two or more relevance scores of each of the two or more information. Further, the analyzing of the two or more information of the two or more contents using the one or more analysis models may include generating an aggregated relevance score for each of the two or more information based on the aggregating. Further, the identifying of the one or more contents from the two or more contents may be further based on the aggregated relevance score associated with each of the two or more information.

[0130] Fig. 9 is a block diagram of a system 900 for facilitating semantic search of a large corpus of content using machine learning, in accordance with some embodiments. Accordingly, the system may include a processing device 906 configured for obtaining at least one input data from at least one user device. Further, the at least one user device may include a smartphone, a tablet, a laptop, and so on. Further, the at least one user device may be associated with at least one user. Further, the at least one user may include an individual, an institution, and an organization that may want to view at least one content. Further, the processing device 906 may be configured for generating at least one keyword based on at least one input data using at least one machine learning model.

[0131] Further, the processing device 906 may be configured for extracting at least one metadata based on the at least one keyword. Further, the processing device 906 may be configured for processing the at least one metadata using at least one second machine learning model. Further, the processing device 906 may be configured for obtaining at least one detailed content data associated with at least one content based on the processing of the at least one meta data. Further, the at least one detailed content data may include an audio content, an audio video content, a visual content, and a textual content. Further, at least one user may want to view the at least one content based on the at least one keyword. Further, the obtaining may include receiving the at least one detailed content from at least one device associated with at least one second user. Further, the at least one second user may include an individual, an institution, and an organization that may host and / or provide the at least one content.

[0132] Further, the processing device 906 may be configured for analyzing the at least one detailed content data using at least one third machine learning model. Further, the at least one third machine learning model may be configured for performing semantic search using natural language processing (NLP) techniques to understand content meaning and context. Further, the processing device 906 may be configured for identifying at least one matching content from the at least one detailed content data based on the analyzing.

[0133] Further, the system may include a communication device 902 configured for transmitting the at least one matching content data associated with the at least one matching content to the at least one user device.

[0134] Further, the system may include a storage device 904 configured for storing the at least one matching content and the at least one input data.

[0135] Fig. 10 is a flow chart of a method 1000 for facilitating semantic search of a large corpus of content using machine learning, in accordance with some embodiments. Accordingly, the method may include a step 1002 of obtaining, using a processing device 906, at least one input data from at least one user device. Further, the at least one user device may include a smartphone, a tablet, a laptop, and so on. Further, the at least one user device may be associated with at least one user. Further, the at least one user may include an individual, an institution, and an organization that may want to view at least one content.

[0136] Further, the method 1000 may include a step 1004 of generating, using the processing device 906, at least one keyword based on at least one input data using at least one machine learning model.

[0137] Further, the method 1000 may include a step 1006 of extracting, using the processing device 906, at least one metadata based on the at least one keyword.

[0138] Further, the method 1000 may include a step 1008 of processing, using the processing device 906, the at least one metadata using at least one second machine learning model.

[0139] Further, the method 1000 may include a step 1010 of obtaining, using the processing device 906, at least one detailed content data associated with at least one content based on the at least one metadata and / or other criteria. Further, the at least one detailed content data may include an audio, a visual, and / or a textual content. Further, at least one user may want to view the at least one content based on the at least one keyword and / or other criteria. Further, the obtaining may include receiving the at least one detailed content from at least one device associated with at least one second user. Further, the at least one second user may include an individual, an institution, and an organization that may host and / or provide the at least one content. Further, in an instance, the at least one user may be a viewer of the at least one detailed content. Further, the at least one second user may be different than the at least one user. Further, in an instance, the at least one second user may include a university official that may use the disclosed system to provide content (or the at least one detailed content) to the at least one user such as students.

[0140] Further, the method 1000 may include a step 1012 of analyzing, using the processing device 906, the at least one detailed content data using at least one third machine learning model. Further, the at least one third machine learning model may be configured for performing semantic search using natural language processing (NLP) techniques to understand content meaning and context. Further, the method 1000 may include a step 1014 of identifying, using the processing device 906, at least one matching content from the at least one detailed content data based on the analyzing.

[0141] Further, the method 1000 may include a step 1016 of transmitting, using a communication device 902, the at least one matching content data associated with the at least one matching content to the at least one user device.

[0142] Further, the method 1000 may include a step 1018 of storing, using a storage device 904, the at least one matching content and the at least one input data.

[0143] Fig. 11 is a flow chart of a method 1100 for facilitating semantic search of a large corpus of content using machine learning, in accordance with some embodiments. Accordingly, the method 1100 may include a step 1102 of analyzing, using the processing device 906, the at least one matching content data using at least one fourth machine learning model. Further, the method 1100 may include a step 1104 of generating, using the processing device 906, at least one content score associated with the at least one matching content based on the analyzing of the at least one matching content data. Further, the at least one content score may represent likeness of the at least one matching content that the at least one user may want to view. Further, the method 1100 may include a step 1106 of generating, using the processing device 906, a result based on the at least one content score. Further, the result may include the at least one matching content organized / arranged in a sequence based on the at least one content score. Further, the method 1100 may include a step 1108 of transmitting, using the communication device 902, the result to the at least one user device.

[0144] Although the invention has been explained in relation to its preferred embodiment, it is to be understood that many other possible modifications and variations can be made without departing from the spirit and scope of the invention as hereinafter claimed.

Claims

CLAIMSWhat is claimed is:

1. A method for facilitating enhanced semantic search of a corpus of content, the method comprising: receiving, using a communication device, at least one request from at least one user device, wherein the at least one request is for conducting at least one content search; analyzing, using a processing device, the at least one request; generating, using the processing device, at least one query based on the analyzing of the at least one request; extracting, using the processing device, a plurality of contents using at least one application programming interface (API) of at least one application based on the at least one query; analyzing, using the processing device, the plurality of contents; generating, using the processing device, a plurality of information of the plurality of contents based on the analyzing of the plurality of contents; analyzing, using the processing device, the plurality of information of the plurality of contents using at least one analysis model, wherein the at least one analysis model is configured for executing at least one assessment for each of the plurality of information of the plurality of contents; identifying, using the processing device, at least one content from the plurality of contents based on the analyzing of the plurality of information of the plurality of contents; generating, using the processing device, at least one result for the at least one content search based on the identifying of the at least one content; andtransmitting, using the communication device, the at least one result to the at least one user device.

2. The method of claim 1, wherein the analyzing of the plurality of information of the plurality of contents comprises analyzing the plurality of information of the plurality of contents using at least one pre-defined scoring criterion, wherein the method further comprises: determining, using the processing device, a plurality of relevancy scores for the plurality of contents based on the analyzing the plurality of information of the plurality of contents using the at least one pre-defined scoring criterion, wherein each of the plurality of relevancy scores represents a suitability of each of the plurality of contents in relation to the at least one request; and analyzing, using the processing device, the plurality of relevancy scores, wherein the identifying of the at least one content from the plurality of contents is further based on the analyzing of the plurality of relevancy scores.

3. The method of claim 1, wherein the at least one request comprises at least one user- defined search criterion for facilitating the conducting of the at least one content search, wherein the analyzing of the plurality of information of the plurality of contents comprises analyzing the plurality of information of the plurality of contents using the at least one user-defined search criterion, wherein the identifying of the at least one content is further based on the analyzing of the plurality of information of the plurality of contents using the at least one user-defined search criterion.

4. The method of claim 1 further comprising: executing, using the processing device, at least one summarization operation on the at least one request; and identifying, using the processing device, a plurality of salient keywords from the at least one request based on the executing of the at least one summarization operation on the at least one request, wherein the generating of the at least one query is further based on the identifying of the plurality of salient keywords.

5. The method of claim 1 further comprising: executing, using the processing device, at least one topic modelling operation on the at least one request; and determining, using the processing device, at least one thematic keyword associated with the at least one request based on the executing of the at least one topic modelling operation, wherein the generating of the at least one query is further based on the determining of the at least one thematic keyword.

6. The method of claim 1 further comprising obtaining, using the processing device, at least one additional keyword based on the analyzing of the at least one request, wherein the generating of the at least one query is further based on the at least one additional keyword, wherein the generating of the at least one query comprises augmenting at least one keyword comprised in the at least one query with the at least one additional keyword.

7. The method of claim 1 further comprising: dividing, using the processing device, a processing resource into a plurality of processing resource segments, wherein each of the plurality of processing resource segments is associated with a processing capacity; assigning, using the processing device, each of the plurality of contents to a respective one of the plurality of processing resource segments; and processing, using the processing device, each of the plurality of contents using the respective one of the plurality of processing resource segments in a parallel configuration based on the assigning, wherein the generating of the plurality of information of the plurality of contents is further based on the processing of each of the plurality of contents using the respective one of the plurality of processing resource segments in the parallel configuration.

8. The method of claim 1, wherein the at least one query comprises at least one API call structured to comply with at least one endpoint specification associated with the at leastone API, wherein the extracting of the plurality of contents is further based on the at least one API call.

9. The method of claim 3 further comprising determining, using the processing device, at least one attribute associated with each of the plurality of contents based on the analyzing of the plurality of information of the plurality of contents using the at least one user- defined search criterion, wherein the at least one attribute comprises at least one of a relevance, a quality, and a level associated with each of the plurality of contents, wherein the identifying of the at least one content is further based on the determining of the at least one attribute.

10. The method of claim 8 further comprising transmitting, using the communication device, the at least one API call to at least one endpoint associated with the at least one API, wherein the extracting of the plurality of contents is further based on the transmitting of the at least one API call to the at least one endpoint.

11. A system for facilitating enhanced semantic search of a corpus of content, the system comprising: a communication device configured for: receiving at least one request from at least one user device, wherein the at least one request is for conducting at least one content search; and transmitting at least one result to the at least one user device; and a processing device communicatively coupled with the communication device, wherein the processing device is configured for: analyzing the at least one request; generating at least one query based on the analyzing of the at least one request; extracting a plurality of contents using at least one application programming interface (API) of at least one application based on the at least one query;analyzing the plurality of contents; generating a plurality of information of the plurality of contents based on the analyzing of the plurality of contents; analyzing the plurality of information of the plurality of contents using at least one analysis model, wherein the at least one analysis model is configured for executing at least one assessment for each of the plurality of information of the plurality of contents; identifying at least one content from the plurality of contents based on the analyzing of the plurality of information of the plurality of contents; and generating the at least one result for the at least one content search based on the identifying of the at least one content.

12. The system of claim 11, wherein the analyzing of the plurality of information of the plurality of contents comprises analyzing the plurality of information of the plurality of contents using at least one pre-defined scoring criterion, wherein the processing device is further configured for: determining a plurality of relevancy scores for the plurality of contents based on the analyzing the plurality of information of the plurality of contents using the at least one pre-defined scoring criterion, wherein each of the plurality of relevancy scores represents a suitability of each of the plurality of contents in relation to the at least one request; and analyzing the plurality of relevancy scores, wherein the identifying of the at least one content from the plurality of contents is further based on the analyzing of the plurality of relevancy scores.

13. The system of claim 11, wherein the at least one request comprises at least one user- defined search criterion for facilitating the conducting of the at least one content search, wherein the analyzing of the plurality of information of the plurality of contents comprises analyzing the plurality of information of the plurality of contents using the atleast one user-defined search criterion, wherein the identifying of the at least one content is further based on the analyzing of the plurality of information of the plurality of contents using the at least one user-defined search criterion.

14. The system of claim 11, wherein the processing device is further configured for: executing at least one summarization operation on the at least one request; and identifying a plurality of salient keywords from the at least one request based on the executing of the at least one summarization operation on the at least one request, wherein the generating of the at least one query is further based on the identifying of the plurality of salient keywords.

15. The system of claim 11, wherein the processing device is further configured for: executing at least one topic modelling operation on the at least one request; and determining at least one thematic keyword associated with the at least one request based on the executing of the at least one topic modelling operation, wherein the generating of the at least one query is further based on the determining of the at least one thematic keyword.

16. The system of claim 11, wherein the processing device is further configured for obtaining at least one additional keyword based on the analyzing of the at least one request, wherein the generating of the at least one query is further based on the at least one additional keyword, wherein the generating of the at least one query comprises augmenting at least one keyword comprised in the at least one query with the at least one additional keyword.

17. The system of claim 11, wherein the processing device is further configured for: dividing a processing resource into a plurality of processing resource segments, wherein each of the plurality of processing resource segments is associated with a processing capacity;assigning each of the plurality of contents to a respective one of the plurality of processing resource segments; and processing each of the plurality of contents using the respective one of the plurality of processing resource segments in a parallel configuration based on the assigning, wherein the generating of the plurality of information of the plurality of contents is further based on the processing of each of the plurality of contents using the respective one of the plurality of processing resource segments in the parallel configuration.

18. The system of claim 11, wherein the at least one query comprises at least one API call structured to comply with at least one endpoint specification associated with the at least one API, wherein the extracting of the plurality of contents is further based on the at least one API call.

19. The system of claim 13, wherein the processing device is further configured for determining at least one attribute associated with each of the plurality of contents based on the analyzing of the plurality of information of the plurality of contents using the at least one user-defined search criterion, wherein the at least one attribute comprises at least one of a relevance, a quality, and a level associated with each of the plurality of contents, wherein the identifying of the at least one content is further based on the determining of the at least one attribute.

20. The system of claim 18, wherein the communication device is further configured for transmitting the at least one API call to at least one endpoint associated with the at least one API, wherein the extracting of the plurality of contents is further based on the transmitting of the at least one API call to the at least one endpoint.

Citation Information

Patent Citations

  • Systems for controllable summarization of content

    US12008332B1

  • Machine learning-based relationship association and related discovery and search engines

    US20190354544A1