System for and method of generating context specification for retrieval and content distribution which are made in context fashion

The system addresses the challenge of irrelevant content distribution by using machine learning to calculate co-occurrence probabilities and context scores, ensuring precise contextual targeting and enhanced advertising performance.

JP2025174972APending Publication Date: 2025-11-28STACKADAPT INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025134470
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2020-02-19
Filing Date
2025-08-12
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing content distribution systems lack an effective and efficient means to specify the context in which content should appear and deliver it to context-relevant websites, leading to irrelevant content distribution and competing incentives between content generators and publishers.

Method used

A system utilizing machine learning techniques and algorithms to preprocess large amounts of Internet content, determining optimal campaign placement opportunities by calculating co-occurrence probabilities and context scores, enabling real-time bid selection for improved relevance and performance.

Benefits of technology

Enhances the relevance of content distribution by accurately matching advertisements with contextually relevant web pages, improving advertising campaign performance through precise contextual targeting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025174972000001_ABST
    Figure 2025174972000001_ABST
Patent Text Reader

Abstract

To provide, according to the present invention, a system for and a method of generating a campaign and efficiently calculating a bid for posting a campaign data to an internet data.SOLUTION: According to an embodiment of the present invention, based on a campaign term and beacon term, a context score of a campaign data can be calculated. The context score can be used to identify an internet content having high page score. If a page score of a specific internet content exceeds a predetermined threshold value, the system can bid a campaign based on a disclosed algorithm using, as an input, an input performance score, the context score, the page score, a campaign budget, and other parameters. Accordingly, if being given to parameters disclosed in the embodiment, the system quickly and effectively calculate an optimum bid for a specific campaign.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to U.S. Provisional Application No. 62 / 978,746, filed February 19, 2020, which is incorporated by reference in its entirety.

[0002] This application relates generally to systems and methods for generating contextual search parameters and delivering highly contextual content over a distributed network. More specifically, this application relates to enabling content providers to efficiently identify contextually relevant content publishers and distributors for particular content. [Background technology]

[0003] Prior art systems use keyword targeting to enable content distributors to match content from a content campaign based on a set of keywords. For example, a search engine may sell advertisements that appear when a search request contains specified keywords. This basic method provides a relatively poor experience, where the content distributed is often irrelevant to the search. For example, a content generator wanting to distribute information related to apple pies and bakeries may want to search and publish content on websites based on the keyword "apple," but because such sites also contain the keyword "apple," the content may mistakenly be posted on technology or business websites that discuss Apple®, a technology company of the same name.

[0004] A second method of specifying the context in which an advertisement should be placed is through some pre-established categorization associated with the content to be distributed, such as Interactive Advertising Bureau (IAB) categories. In this method, publishers and content providers categorize domains or articles into pre-defined IAB categories. A content generator (e.g., an advertiser) then selects the categories in which the content generator wants to collocate its content. However, this method limits the content generator to choosing only from pre-defined categories, creating competing incentives between the content generator and the content publisher / provider. As a result, the content publisher / provider may want to place their articles or other materials in as many or as few categories as possible, which may overly limit or overly satisfy the distribution of the collocated content belonging to the content generator. Other than specifying a large list of keywords or selecting IAB categories, there is no known way to specify the context in which a content generator's content (e.g., advertisements) should appear. Additionally, once such content is specified by the content generator, there is no known way to efficiently distribute such content to the specified context.

[0005] Therefore, what is needed is a more effective and efficient means for specifying the context in which content should appear, and for efficiently delivering content to context-relevant websites for collocation and presentation. Summary of the Invention [Problem to be solved by the invention]

[0006] The systems and methods disclosed herein are intended to address the deficiencies in the art discussed above, but may also provide additional or alternative advantages. As described herein, the systems and methods may use contextual terms to improve correlation between campaign content and Internet content. These systems and methods allow advertisers to specify the context in which they want their ads to appear and deliver the ads to sites that have that context.

[0007] The improved systems and methods disclosed herein include one or more machine learning techniques and algorithms that execute complex algorithms to preprocess a large number (e.g., thousands or millions) of Internet content, such as websites, videos, or audio files, to determine the best opportunity for a campaign to bid to deliver an impression. Contextual or campaign terms can be keywords or phrases or other means of identifying content, such as the digital fingerprint of an audio or visual file. For example, an image of a hurricane may have a digital fingerprint that indicates it is a hurricane and be assigned the keyword "hurricane." Campaign content can include advertisements, such as those selling consumer electronics and magazine subscriptions, or promotional materials, such as political campaigns or calls to action. Current systems are unable to efficiently and effectively connect Internet content with large amounts of campaign content. Internet content can include online videos, news articles, blog posts, real-time video feeds, online forums, social networks, online magazines, games, mobile software applications, or Internet of Things (IoT) devices, such as smart displays, that display content. Much of this Internet content, such as new videos, blog posts, articles, or forum content, can be generated in real time, and the computerized systems and methods of this application can combine such real-time impression opportunities with campaigns. The improved systems and methods disclosed in this application enable greater relevance and better results by storing information in a database that can provide real-time decision-making between campaigns and placements adjacent to Internet content. Improved matches between advertisements and the context in which they appear are likely to result in higher advertising campaign performance. [Means for solving the problem]

[0008] In one embodiment, a computer-implemented method includes applying, by a computer, a machine learning model to a set of first context terms received from a client device to output a set of beacon terms from a corpus database, the machine learning model being trained on a plurality of corpus terms stored in the corpus database and determining a plurality of co-occurrence probabilities corresponding to the plurality of corpus terms; calculating, by the computer, a plurality of page scores for a plurality of corpus web pages stored in the corpus database based on the set of beacon terms and the set of first context terms; identifying, by the computer, a set of context web pages, among the plurality of corpus web pages, having page scores that meet a threshold; and applying, by the computer, the machine learning model to the set of first context terms received from the client device to output a set of beacon terms from a corpus database, the machine learning model being trained on a plurality of corpus terms stored in the corpus database to determine a plurality of co-occurrence probabilities corresponding to the plurality of corpus terms. and applying the updated set of beacon terms to the first set of context terms and the second set of context terms to output an updated set of beacon terms; calculating, by the computer, one or more updated page scores for one or more corpus web pages stored in the corpus database based on the updated set of beacon terms, the first set of context terms, and the second set of context terms; updating, by the computer, the set of context web pages based on the one or more updated page scores; and storing, by the computer, campaign data for the user including the updated set of beacon terms and the set of context web pages in the campaign database, wherein the campaign data is configured to perform a real-time bid selection operation for one or more available web pages during a real-time bid selection operation.

[0009] In another embodiment, the system includes: a corpus database including a non-transitory storage medium configured to store at least a portion of a plurality of Internet content data corresponding to uniform resource locators (URLs) and page context scores; a campaign database including a non-transitory storage medium configured to store campaign data for a plurality of users, the campaign data configured to perform a real-time bid selection operation for one or more available web pages during a real-time bid selection operation; and a server including a processor, the processor applying a machine learning model to a set of first context terms received from a client device to output a set of beacon terms from the corpus database, the machine learning model being trained on a plurality of corpus terms stored in the corpus database to determine a plurality of co-occurrence probabilities corresponding to the plurality of corpus terms; and outputting the set of beacon terms stored in the corpus database based on the set of beacon terms and the set of first context terms. calculating a plurality of page scores for a plurality of corpus web pages stored in the corpus database; identifying a set of context web pages in the plurality of corpus web pages having page scores that meet a threshold; applying a machine learning model to the first set of context terms and the second set of context terms received from the client device to output an updated set of beacon terms; calculating one or more updated page scores for one or more corpus web pages stored in the corpus database based on the updated set of beacon terms, the first set of context terms, and the second set of context terms; updating the set of context web pages based on the one or more updated page scores; storing campaign data of the user including the set of beacon terms and the set of context web pages in a campaign database, the campaign data being configured to perform a real-time bid selection operation for one or more available web pages during a real-time bid selection operation;and a server configured to perform the steps of:

[0010] In another embodiment, a computer-readable medium including machine-executable program instructions, wherein execution of the program instructions by one or more processors of a computer system causes the one or more processors to: apply a machine learning model to a set of first context terms received from a client device to output a set of beacon terms from a corpus database, wherein the machine learning model is trained on a plurality of corpus terms stored in the corpus database to determine a plurality of co-occurrence probabilities corresponding to the plurality of corpus terms; calculate a plurality of page scores for a plurality of corpus web pages stored in the corpus database based on the set of beacon terms and the set of first context terms; identify, among the plurality of corpus web pages, a set of context web pages having page scores that meet a threshold; and apply the machine learning model to the client device to output a set of beacon terms from a corpus database, wherein the machine learning model is trained on a plurality of corpus terms stored in the corpus database to determine a plurality of co-occurrence probabilities corresponding to the plurality of corpus terms. a computer-readable medium for causing a user to: apply a set of beacon terms to a first set of context terms and a second set of context terms received from a client device to output an updated set of beacon terms; calculate one or more updated page scores for one or more corpus web pages stored in a corpus database based on the updated set of beacon terms, the first set of context terms, and the second set of context terms; update the set of context web pages based on the one or more updated page scores; and store campaign data for the user including the set of beacon terms and the set of context web pages in a campaign database, the campaign data being configured to perform a real-time bid selection operation for one or more available web pages during a real-time bid selection operation.

[0011] Both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the claimed invention. The accompanying drawings, which constitute a part of this specification, illustrate embodiments of the invention and, together with the description, explain the invention. [Brief explanation of the drawings]

[0012] [Figure 1] 1 illustrates components of a distributed computer system for distributing content, according to one embodiment. [Figure 2] 1 shows a flowchart for generating a corpus of recent internet content documents from which content can be posted. [Figure 3] 1 shows a flowchart for identifying campaign context and correlating campaign content with third-party internet content. [Figure 4] 1 shows a flowchart of a method according to one embodiment. [Figure 5] 1 illustrates an exemplary web page that allows a user to manage and update the context-building behavior of the system, according to one embodiment. [Figure 6] 1 illustrates an exemplary web page that allows a user to manage and update the context-building behavior of the system, according to one embodiment. [Figure 7] 1 illustrates an exemplary web page that allows a user to manage and update the context-building behavior of the system, according to one embodiment. [Figure 8] 1 illustrates an exemplary web page that allows a user to manage and update the context-building behavior of the system, according to one embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0013] Reference will now be made to various illustrated embodiments, specific language being used herein to describe the same. Nevertheless, it will be understood that no limitations on the claims or the present disclosure are intended thereby. Alterations and further modifications of the features of the invention shown herein, and additional applications of the principles of the subject matter shown herein that would occur to those skilled in the relevant art and in possession of this disclosure, are to be considered within the scope of the subject matter disclosed herein. The present disclosure will now be described in detail with reference to the embodiments illustrated in the drawings that form a part hereof. Other embodiments may be used and / or other changes may be made without departing from the spirit or scope of the present disclosure. The exemplary embodiments described in the detailed description are not meant to limit the subject matter presented herein.

[0014] It should be understood that the embodiments described herein are merely examples for purposes of illustrating the techniques, technical components, and processes disclosed herein. In particular, the various embodiments described herein contemplate advertising and real-time bidding (RTB) implementations of the disclosed techniques and features. However, some embodiments may implement aspects of the disclosed techniques for other purposes or contexts, such as for building and deploying search queries in real-time data retrieval and archives, or for querying digitized data libraries.

[0015] FIG. 1 illustrates components of a distributed computer system 100 for delivering campaign content for display with Internet content, according to one embodiment. The illustrated system 100 may include a web server 101, databases 105, 106, and 107, an administrator device 109, a distributed client 111, a third-party content server 113, and a real-time bidding (RTB) server 114. An embodiment may include additional or alternative components, or omit certain components from those of FIG. 1 , and still be within the scope of this disclosure. Certain components of system 100 may be embodied in multiple computing devices. For example, web server 101 (or other server) is shown as a single computing device but may include any number of computing devices. Additionally or alternatively, certain components may be integrated and embodied in the same computing device. For example, corpus database 105 may be hosted on the same computing device as web server 101.

[0016] The web server 101 executes software programming to crawl the third-party content server 113 to extract and download various types of web page data, including content and metadata. The web server 101 associates content with specific context terms, such as keywords or beacon terms. The web page data or content may include, for example, metadata, fingerprints for the web page's media, web page or server identifiers, content tags, or content containing or associated with such context terms. The web server 101 may, for example, identify known image or video content fingerprints that are pre-correlated with the context terms. As an example, Internet content on the third-party content server 113 may include a known image of a map, and the web server 101 may associate that content with maps generally or with a specific map, e.g., Asia or North America, depending on the image. In this way, the web server 101 may later correlate campaign content with map content if they are related to each other by having a context score above a predetermined threshold, as described further below.

[0017] The web server 101 may also host a website accessible to end users, such as at the distributed client 111. The website may allow users to define and execute campaigns in accordance with embodiments of the present disclosure. For example, the web server 101 may correlate campaign content as a function of various internet content. The web server 101 may be any computing device equipped with a processor and non-transitory machine-readable storage capable of performing the various tasks and processes described herein. Non-limiting examples of such computing devices may include a workstation computer, a laptop computer, a server computer, a laptop computer, etc. Although the exemplary system 100 includes a single web server 101, some embodiments of the web server 101 may include any number of computing devices operating in a distributed computing environment.

[0018] The web server 101 may execute software applications (e.g., Apache®, Microsoft IIS®) configured to host websites that may generate and serve various web pages to the client devices 111. The client-facing websites may be used to generate and access data stored on the system databases 105, 106, 107 of the system 100, or to execute various instructions from the client devices 111, the administrator device 109, or another device in the system 100. In some implementations, the web server 101 may be configured to require user authentication based on a set of user authentication credentials (e.g., username, password, biometrics, cryptographic certificate). In such implementations, the web server 101 may access the system databases 105, 106, 107 configured to store user credentials, which the web server 101 may be configured to reference to determine whether an entered set of credentials (intended to authenticate the user) matches the appropriate set of credentials that identify and authenticate the user. Similarly, in some implementations, the web server 101 may generate and send software code for a web page to the client device 111 based on the user's role within the system 100 (e.g., administrator, campaign content provider, or Internet content provider).

[0019] During operation, web server 101 (or other computing devices of system 100) executes software programming for training and deploying one or more machine learning models and associated machine learning algorithms. The machine learning operations may include machine learning techniques and algorithms executed by any processor, such as various types of neural networks (e.g., convolutional neural networks (CNNs), deep neural networks (DNNs)), linear regression, logistic regression, k-means, k-nearest neighbors (kNNs), or support vector machines (SVMs), among others. A crawler program executed by web server 101 automatically traverses any number of URLs and downloads web page data (e.g., content, metadata) for the web pages. Web server 101 stores some or all of the web page data in corpus database 105. Web server 101 trains machine learning models on the corpus of web page data to identify and generate various statistical associations between terms, phrases, metadata, or other information indicative of the nature or context of each particular web page. The machine learning models determine co-occurrence and other statistical context data for various types of web page data in corpus database 105. For example, web server 101 can apply any number of natural language processing and vectorized machine learning algorithms to the web page content or metadata to generate feature vectors for various corpus terms in order to extract embeddings representing various statistical measures of the corpus terms. These similar algorithms can be applied to input context terms (e.g., in-context terms, out-of-context terms) to extract embeddings for the user's context terms, which web server 101 can use to determine distances from corpus terms and generate context scores for each of the user's context terms based on that distance. User feedback and / or additional context terms can be incorporated to adjust the embeddings and / or adjust the weightings of the various algorithms.The web server 101 may run additional or alternative natural language processing and vectorized machine learning algorithms on the corpus web pages to generate a page score, which represents, for example, the number of instances in which a particular beacon term (e.g., a term extracted from a corpus database that has a short distance from the user's in-context term) occurs within the web page content and / or the in-context term occurs within the web page content.

[0020] Once trained, the machine learning model is ready to take in a set of terms / phrases from an end user building a contextualized campaign or query. The web server 101 receives the set terms / phrases from the client device 111 and applies the trained machine learning model to the set of input terms. The machine learning model determines the co-occurrence probability (and / or other statistical measure) of the input terms co-occurring with corpus terms. The machine learning model then outputs corpus terms / phrases that have a probability that meets a co-occurrence threshold. The web server 101 may then present these terms / phrases to the end user via a GUI, such as a web page presented on the browser of the client device 111.

[0021] The end user can send feedback or other instructions to the web server 101, which the web server 101 can use to further train and develop a machine learning model for the particular end user. The web server 101 receives feedback from the end user indicating whether particular terms should be given more or less weighting (e.g., in-context, out-of-context). The web server 101 then reapplies the trained machine learning model to each set of input terms received from the client device 111 and adjusts the scored weights assigned to the terms according to the user feedback. When the machine learning model is trained and tuned to the end user's context, the web server 101 then applies the machine learning model to a bid stream of URLs for available web pages received from the RTB server 114, as described in further detail in Figures 2-4. In some embodiments, these machine learning operations are performed to train and execute a machine learning model to identify context web pages for context terms generated based on the input terms.

[0022] System databases 105, 106, and 107 may be hosted on computing devices that include non-transitory, machine-readable storage media and are capable of performing the various tasks described herein. As shown in FIG. 1 , system databases 105, 106, and 107 may be accessed by web server 101 via one or more networks. System databases 105, 106, and 107 may be hosted on the same physical computing device that functions as web server 101 and / or provides additional or alternative functionality (e.g., application server, authentication server). System 100 may include any number of public and / or private networks having various hardware and software components configured to interconnect components of system 100 and host data communications. Non-limiting examples of such networks may include a local area network (LAN), a wireless local area network (WLAN), a metropolitan area network (MAN), a wide area network (WAN), and the Internet. Communications over the network may be performed according to various communication protocols, such as Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), and IEEE communication protocols.

[0023] Corpus database 105 stores web pages and associated metadata. Corpus database 105 may be continuously updated at given intervals by a crawler software program executed by web server 101 or other device. In some embodiments, corpus database 105 may be updated during or after a bid time, where web server 101 or other device detects a previously unseen uniform resource locator (URL), triggers a crawler routine to scrape the web page at that URL, and stores the web page and associated metadata in corpus database 105. Each URL or web page may be stored with a timestamp, URL, and content tag for later processing. Web server 101 may access and query data records in corpus database 105 when performing various processes to build contextualized searches for content generator end-users, as described herein.

[0024] The context database 106 stores end-user-specific information related to a contextualized search after the end user constructs the contextualized search. For example, the web server 101 stores the context beacon phrase score (or other data value) in the context database 106 after it is calculated or generated by the web server 101. At bidding time, the web server 101 accesses the end user's data records to determine how to compete for a particular URL published by the RTB server 114. In some implementations, the data records may be moved from the context database 106's hard disk to the web server's 101's memory to improve access speed.

[0025] The cache database 107 stores frequently requested web pages and associated metadata, which the web server 101 accesses during bidding to determine how to compete for a particular URL. The data records for web pages stored in the cache database 107 may be "content-only" versions that have various forms of media, third-party content, or other non-content-related data removed. The data records for web pages may also include indicators of beacon phrases on the web page and frequency scores for the beacon phrases. The cache database 107 is generated by software routines that pull particular web pages stored in the corpus database 105 and then convert them into content-only versions that are stored in the cache database 107. For example, if the RTB server 114 publishes a URL a certain number of times or if the URL is unseen, the software routine loads the URL into a queue for preprocessing for conversion and storage in the cache database 107, thereby allowing the URL to be quickly accessed by the web server 111 during a later bidding period to perform the various processes described herein.

[0026] The administrator device 109 executes various software programming, enabling an administrator-user to maintain and improve the web server 101. The administrator device may be any computing device equipped with hardware (e.g., processor, non-transitory storage medium) and software components and capable of performing the various tasks and processes described herein. In operation, the administrator device 109 may initiate operations to build a corpus of web page data or documents for use in future campaigns, as shown in FIG. 2, either manually (according to user input) or automatically (according to user configuration). The administrator may also configure the administrator device 109 to perform various tasks that assist the administrator in maintaining quality control. For example, the administrator may review the association of corpus web page data or documents and context terms to ensure proper correlation. This process may also be performed via programmatic means in certain embodiments of the present disclosure.

[0027] The RTB server 114 may be one or more computing devices in the RTB system that host a web portal or other web-based external service that publishes and manages competitions between content generators to vie for the opportunity to distribute content on various web page URLs. In some embodiments, the RTB server 114 transmits or otherwise publishes URLs of particular web pages that are received by the web server 101. The content generator (e.g., an end user) uses the client device 111 to generate a context specification that informs the web server 101 of URLs of interest to the content generator. The web server 101 then automatically initiates transactions with the RTB server 114 for those URLs that have higher contextual relevance to the content generator based on the context specification of that particular content generator. The content is then forwarded to the third-party server 113 that hosts the URLs for publication and display on the Internet.

[0028] While the RTB server 114 in the exemplary embodiment herein is associated with an ad-centric RTB service, it should be understood that the RTB service is merely exemplary and non-limiting. Other embodiments may include any third-party external web service that exposes information (e.g., an API service) and instructions (e.g., API requests) for performing various tasks described herein. Similarly, an API service (e.g., Bidstream) may be any remote data publishing service that generates and exposes data to subscribing computing devices for data consumption and for executing processes associated with the API service. An API request (e.g., a bid request) may include computer-implemented instructions for executing one or more processes associated with the API service, such as collecting and responding to the requested data or generating a GUI for displaying the exposed data.

[0029] In an example embodiment, a bid stream is a data stream of published URLs available for bids to content generators interested in posting campaign content on those URLs, and a bid request may be computer-implemented instructions on a computer to display and / or distribute those URLs and collect bid inputs corresponding to the published URLs, which thereby triggers the computer to perform the various processes described herein to generate and submit to the API service the data entered via the API request.

[0030] 2 shows a flowchart of steps performed by a web server to build a corpus of documents (e.g., web pages) for future campaign content postings, according to an embodiment of an exemplary method 200. While method 200 is described with respect to a single computing device and a single database, it should be understood that any number of computing devices may be involved in other embodiments, including additional or alternative computing devices from the web server and corpus database. It should also be understood that particular embodiments may provide additional or alternative steps, or omit certain steps from those of method 200, and still be within the scope of the present disclosure.

[0031] In step 201, a computer hosting a web server application (e.g., web server 101) identifies an API service (e.g., BidStream) associated with a real-time bidding (RTB) system to call and access. Using the selected API, the computer receives and sends API requests (e.g., bid requests) to and from the RTB server and / or client computers. As mentioned, an API service may be any remote data publishing, querying, and / or archiving executable service that generates and exposes data to subscribing computing devices, which consume the exposed data and perform various processes associated with the API service. An API request may include computer-implemented instructions for performing one or more processes associated with the API service, such as collecting and responding to requested data or generating a GUI to display the exposed data.

[0032] In method 200, the bid stream is a data stream of published URLs available for bidding to content generators interested in posting campaign content at those URLs, and the bid request may be computer-implemented instructions to a computer that display and / or distribute those URLs and collect bid inputs corresponding to the published URLs, thereby triggering the computer to perform various tasks described herein.

[0033] In step 202, the computer samples web pages associated with the bid requests by scraping the web pages located at URLs published in the bid stream. This internet-centric approach allows for efficient scraping and identification of keywords, beacon terms, or fingerprints associated with third-party content to generate future scores associated with the campaign content.

[0034] In step 203, the scraping algorithm may create an individual corpus snapshot for each URL, including various types of web page data for the web page corresponding to the URL. In step 204, the web server stores the individual corpus snapshot in a corpus database (e.g., corpus database 105). Those skilled in the art will appreciate that there may be one or more computer-implemented techniques for capturing or scraping web page content and storing the content in a corpus database. In some implementations, the content may be updated according to crawler software routines that programmatically traverse URLs or web pages and download the web page content according to the crawler software's algorithmic logic. In some implementations, the corpus database may be updated by a computing device of the system using data captured during the live bidding process, where the web pages to be updated or URLs to such web pages may be received from the bidding system via the bid stream.

[0035] 3 shows a flowchart of steps performed for identifying campaign context and correlating campaign content with third-party Internet content, according to an embodiment of an exemplary method 300. While method 300 is described with reference to particular computing devices and databases, it should be understood that any number of computing devices may be involved in other embodiments, including computing devices in addition to or alternative to those mentioned herein. It should also be understood that particular embodiments may provide additional or alternative steps, or omit certain steps from those of method 300, and still be within the scope of the present disclosure.

[0036] In a first step 301, a computer hosts a website that receives input from an end user's (e.g., a content generator's) client computer to define a new content distribution campaign or update a previously generated campaign. The input may include input phrases, which may be in-context phrases and / or out-of-context phrases, where each phrase may include any number of words. This step 301 may include identifying specific beacon phrases associated with the campaign, which are phrases contextually related to the input phrase that are automatically identified based on the input phrase. The system may receive in-context terms that the user indicated are within the context of a desired search. The input may typically include keywords (one or more words) entered by the user, such as is the case in method 300. However, in some implementations, the input from the user may include a website URL or a fingerprint for media (e.g., audio, visual, audiovisual) content that has a high context that correlates with the campaign content.

[0037] Users often want campaign content to be featured in a particular context. For example, an insurance company may determine that their campaign is most effective if featured next to articles primarily about severe and stormy weather. As another example, a backpacking company may want to feature campaign content on pages related to adventure travel. As used herein, contextualization is the process of featuring campaign content in third-party internet content that has an associated context. As an example, campaign contexts, or input of words or phrases within a context, may include "winter storm," "ice conditions," and "flood warning."

[0038] Input received from the user in current step 301 or a later step may also specify out-of-context terms (as described further below), which are terms that are out of context with the campaign. Based on the algorithmic scores discussed herein for the input phrases (e.g., in-context phrases, out-of-context phrases), the computer determines or updates a list of identified beacon phrases that are contextually relevant to the user's campaign.

[0039] The system calculates and assigns context scores to certain terms, such as making the word "heat" more important than the word "hurricane." In some implementations, the system may calculate and assign context scores for out-of-context terms to make them more or less important. Context scores may be decimal numbers and may be assigned automatically or based on user preferences. In some situations, the system may apply a default context score or manually assign context scores that are the same for similar terms. For example, the terms "heat" and "hot" may receive the same context score.

[0040] Based on the in-context and out-of-context terms received from the user, the algorithm of step 302 may calculate a context score for content in corpus database 105 based on probabilistic implications, word co-occurrence, and context scores. The algorithm of step 302 may use those context scores to find high-context websites. As shown in Figure 5, a web page 500 of a website hosted by a computer includes an input that allows a user to enter a set of context phrases. Web page 500 includes input boxes that allow a user to enter in-context phrases and out-of-context phrases.

[0041] Referring back to FIG. 3 , the algorithm in step 302 uses the input phrase to calculate word probabilities, word co-occurrences, and probabilistic implications between words. The system may determine campaign context scores for terms, such as phrases, keywords, or fingerprints, based on the probabilistic implications and word co-occurrences. The system uses these scores to generate context scores for web pages and identify high-context websites. In step 302, the system queries the corpus database 105 to determine word probabilities, word co-occurrences, and / or probabilistic implications between words to identify beacon phrases associated with the campaign. Based on such queries to the corpus database 105, the computer may generate beacon phrase and page context scores and, in some implementations, retrieve URLs and other electronic content with high context scores.

[0042] Continuing with the weather and adventure travel example, the results of step 302 may be as follows:

[0043] In the weather example, where the user inputs "winter storm," "ice conditions," and "flood warning" as in-context phrases, the computer uses these in-context phrases to query pre-stored web page content in the corpus database 105 and calculates beacon phrases based on the corpus content. In this example, the computer returns the following beacon phrases: "weather service," "national weather," "nws," "storm warning," "snowfall," "snow accumulation," "coastal flooding," and "wind gusts."

[0044] In the adventure travel example, where the user inputs "flight," "adventure," "trip," and "hotel" as in-context phrases, the computer uses these in-context phrases to query pre-stored web page content in corpus database 105 and calculates beacon phrases based on the corpus content. In this example, the computer returns the following beacon phrases: "skyscan," "airfare," "connections," "booking.com," "rebook," "hostel," "expedia," "tsa precheck," "icelandair," "carry on," "itinerary," and "ryanair."

[0045] In step 304, the computer identifies high context web pages based on the beacon phrase and the input phrase, and then generates a user feedback display that shows the user the resulting URLs that have high context for the campaign.

[0046] For example, in the weather example, after the computer queries the corpus database 105 and calculates the beacon phrases, phrase scores, and page context scores, the computer returns, via a user feedback display, the following list of URLs of web pages that were calculated as having high context scores: https: / / www.mlive.com / weather / 2018 / 01 / heres_a_snowfall_tally_on_mich.html https: / / www.chicagotribune.com / news / breaking / ct-first-snowfall-chicago-2016-htmlstory.html https: / / www.masslive.com / weather / 2016 / 11 / these_are_the_10_snowiest_citi.html https: / / www.express.co.uk / showbiz / tv-radio / 1151582 / snowfall-season-3-how-many-episodes-are-in-snowfall-fx-series-damson-idris https: / / www.express.co.uk / showbiz / tv-radio / 1151561 / Snowfall-season-3-cast-Who-is-in-the-cast-of-Snowfall-FX-series-Damson-Idris https: / / www.theactivetimes.com / snow / n / 14-cities-get-most-snowfall https: / / www.tripsavvy.com / does-it-ever-snow-in-memphis-2321876 https: / / www.denverpost.com / 2019 / 09 / 08 / colorado-weather-september-snowfall-denver / amp / https: / / www.mlive.com / weather / 2018 / 05 / and_the_winner_of_michigans_wi.html https: / / minecraft.gamepedia.com / snowfall https: / / seat42f.com / tv-review-snowfall.html https: / / www.denverpost.com / 2019 / 05 / 04 / colorado-weather-front-range-late-season-snow https: / / www.denverpost.com / 2019 / 05 / 10 / denver-weather-below-average-snowfall

[0047] Similarly, in the adventure travel example, after the computer queries the corpus database 105 and calculates the beacon phrases, phrase scores, and page context scores, the computer returns, via a user feedback display, the following list of URLs of web pages that were calculated as having high context scores: https: / / www.annees-de-pelerinage.com / the-best-hotels-in-machu-picchu-for-any-budget https: / / www.drinkteatravel.com / train-to-machu-picchu-tickets / https: / / www.forbes.com / sites / geoffwhitmore / 2018 / 04 / 03 / how-to-book-aer-lingus-award-flights-to-ireland-for-cheap https: / / traveltips.usatoday.com / closest-airport-machu-picchu-109221.html http: / / www.travelfuntu.com / insider-info / airports-that-offer-free-city-tours https: / / www.whereverwriter.com / 15-things-machu-picchu https: / / traveltips.usatoday.com / closest-airport-machu-picchu-109221.html https: / / www.thebrokebackpacker.com / best-hostels-in-cinque-terre-italy / https: / / www.thebrokebackpacker.com / best-hostels-in-koh-lanta-thailand / https: / / travel-made-simple.com / layover-long-enough /

[0048] Once the user is presented with a user feedback display including the resulting campaign contextualization results (e.g., beacon phrases, high-context webpage URLs, scores), the user may refine the campaign through a web portal. The user's refinement feedback is entered into the computer via a GUI on the user's client device, allowing the user to, for example, select, deselect, or otherwise enter inputs indicating in-context phrases (high-scoring context phrases), computer-identified beacon phrases, or out-of-context phrases (low-scoring context phrases). The user's GUI may also include inputs for selecting, deselecting, or otherwise entering inputs indicating website URLs with high or low context scores, respectively.

[0049] Referring to FIG. 6, a web page 600 displays the context phrases identified based on the user input (shown in FIG. 5) and the associated URLs of the context web pages identified by the computer based on the user's previous input.

[0050] As illustrated by FIG. 3 , the contextualization campaign building process can be iterative. In particular, through iterations of the previous steps 301-304, a user can refine and confirm aspects of the contextualized campaign data, including a set of in-context phrases, a set of beacon phrases, a set of out-of-context phrases, a set of high-context webpage URLs, and various calculated scores that together define the contextualized search parameters the user would like to deploy for the user's content delivery campaign. In some cases, there may be a predetermined number of iterations, and in some cases, the user can iterate until the user is satisfied. The final campaign data can be stored in the context database 106 (sometimes referred to as the “campaign database”).

[0051] For example, in the weather example, the user may refine the campaign by entering out-of-context phrases. In this example, the user may determine that heat advisories and hurricanes are not relevant forms of extreme weather conditions, so the user may enter "heat," "heat advisory," and "hurricane" as out-of-context phrases. The computer re-queries the corpus database 105 using the user-selected input phrases to generate revised context beacon phrases and high-context web pages. In this example, the computer generates the following list of updated context beacon phrases: "snowfall," "snowstorm," "winter," "icy," "snow accumulation," "snow," "snowy," "caltran," "spotter," and "commute." The computer also creates and displays the following updated list of URLs to the user: https: / / www.mlive.com / weather / 2018 / 01 / heres_a_snowfall_tally_on_mich.html https: / / www.chicagotribune.com / news / breaking / ct-first-snowfall-chicago-2016-htmlstory.html https: / / www.masslive.com / weather / 2016 / 11 / these_are_the_10_snowiest_citi.html https: / / www.express.co.uk / showbiz / tv-radio / 1151582 / snowfall-season-3-how-many-episodes-are-in-snowfall-fx-series-damson-idris https: / / www.express.co.uk / showbiz / tv-radio / 1151561 / Snowfall-season-3-cast-Who-is-in-the-cast-of-Snowfall-FX-series-Damson-Idris https: / / www.theactivetimes.com / snow / n / 14-cities-get-most-snowfall https: / / www.tripsavvy.com / does-it-ever-snow-in-memphis-2321876 https: / / www.denverpost.com / 2019 / 09 / 08 / colorado-weather-september-snowfall-denver / amp / https: / / www.mlive.com / weather / 2018 / 05 / and_the_winner_of_michigans_wi.html https: / / minecraft.gamepedia.com / snowfall https: / / seat42f.com / tv-review-snowfall.html https: / / www.denverpost.com / 2019 / 05 / 04 / colorado-weather-front-range-late-season-snow https: / / www.denverpost.com / 2019 / 05 / 10 / denver-weather-below-average-snowfall

[0052] In the travel example, the user may also refine the campaign by entering out-of-context phrases. In this example, the user may determine that TSA Precheck and airline websites are irrelevant, so the user may enter “TSA Precheck,” “Precheck,” “TSA,” “airline,” and “airline hub” as out-of-context phrases. The computer re-queries the corpus database 105 using the user-selected input phrases to generate revised context beacon phrases and high-context web pages. In this example, the computer generates the following list of updated context beacon phrases: “Skyscan,” “airfare,” “connections,” “booking.com,” “rebooking,” “hostel,” “Expedia,” “Icelandair,” “carry-on,” “itinerary,” and “Ryanair.” The computer further creates and displays the following updated list of URLs to the user: https: / / www.annees-de-pelerinage.com / the-best-hotels-in-machu-picchu-for-any-budget https: / / www.drinkteatravel.com / train-to-machu-picchu-tickets / https: / / www.forbes.com / sites / geoffwhitmore / 2018 / 04 / 03 / how-to-book-aer-lingus-award-flights-to-ireland-for-cheap https: / / traveltips.usatoday.com / closest-airport-machu-picchu-109221.html http: / / www.travelfuntu.com / insider-info / airports-that-offer-free-city-tours https: / / www.whereverwriter.com / 15-things-machu-picchu https: / / traveltips.usatoday.com / closest-airport-machu-picchu-109221.html https: / / traveltips.usatoday.com / closest-airport-machu-picchu-109221.html https: / / www.thebrokebackpacker.com / best-hostels-in-cinque-terre-italy / https: / / www.thebrokebackpacker.com / best-hostels-in-koh-lanta-thailand / https: / / www.thebrokebackpacker.com / best-hostels-in-koh-lanta-thailand https: / / travel-made-simple.com / layover-long-enough /

[0053] The user may indicate via the GUI that the iterative campaign contextualization building process is complete. In response, the computer is instructed to store the final campaign data (e.g., in-context terms, beacon terms, scores, high context page URLs) in the context database 106. In some implementations, the previous iterative steps 301-304 may be re-implemented later by the user to further improve or refine the campaign contextualization, even after the campaign has launched or otherwise evolved. Thus, in such implementations, storing the campaign data in the context database 106 does not imply that the campaign is immutable.

[0054] Referring to Figure 7, web page 700 again displays an input box, allowing the user another opportunity to update / refine the context phrases and context web pages. The computer displays web page 700 in response to the computer receiving instructions from the client device via the GUI display to perform another iteration to generate the context phrases and context web pages. In Figure 8, web page 800 updates and displays the context phrases identified by the computer based on the updated user input (shown in Figure 8) and the associated URLs of the context web pages identified by the computer based on the user's previous input.

[0055] In some embodiments, a computing device (e.g., a computer hosting a website) generates context data based on user input and feedback received from the client device, and the context data is displayed to the user on the client device's GUI (e.g., browser, web page). Additional information about the specified context (e.g., context data) includes, for example, but not limited to, a score calculated based on the user input, terms identified based on the user input, portions of web pages with positive context scores, or a list of web pages in a particular context score range. The computer (or other computing device) generates context data about various aspects of the user's campaign, terms / phrases, web pages (e.g., context web pages, corpus web pages, cached web pages), and / or other aspects of the system. The context data for each web page may include any statistical information extracted or calculated by the computer at any point in method 300, or before, during, or after method 300. The computer updates data records in various databases of the system according to the calculated or recalculated context data.

[0056] Referring back to FIG. 3 , in optional step 305, certain content stored in corpus database 105 may be stored in cache database 107. The content stored in cache database 107 may include the most frequently requested web pages (or other data content) that are subject to competition, as such information (e.g., number of requests) is determined by the computer or received from an RTB server. By minimizing the amount of data in cache database 107, the computer can more efficiently perform processes performed during bidding, such as steps 306-309 below. The content stored in cache database 107 may be automatically or manually selected, for example, by administrator selections entered as configuration into the computer directly or from an administrator computing device. For example, the computer may select a particular web page when the web page has been requested from RTB a certain threshold number of times during bidding, and / or the computer may automatically select a web page or content when such web page is not in cache database 107. To further improve computational efficiency when performing the bidding process, the web pages may be transformed or otherwise stripped of unnecessary data when they are stored in cache database 107. It should be understood that optional step 305 may be performed before, during, or after method 300.

[0057] In step 306, at bidding time, the computer receives a bid stream of bid requests from the RTB service's servers and initiates the user's content distribution campaign. The bid stream contains one or more URLs of web pages for which content generators, such as the user of process 300, compete to host content. In steps 307-309 below, the computer executes an automated process based on the user's contextualized campaign data to compete on the user's behalf and distribute the user's content to third-party servers hosting the desired URLs.

[0058] In step 307, using the same or similar algorithm as in the previous step 302 (e.g., calculating a fast, linear, vector product between the campaign's context score and the occurrence of terms on the page), and using the campaign data (e.g., in-context phrases, out-of-context phrases, beacon phrases) stored in the context database 106, the computer calculates beacon phrase scores and page context scores for the web pages published in the bid stream by RTB. The computer matches the URLs (of the available web pages) published in the bid stream with the URLs in the cache database 107 to quickly calculate web page context scores for the available web pages, identify which web pages have the highest context scores, and indicate which web pages are most contextually relevant to the user's campaign data.

[0059] In current step 307, by executing the same or similar algorithm that generated the beacon phrase score (as in previous steps 307-309), the computer may calculate a page score for the web page URL published by RTB and a context score for each of the context phrases on the particular web page, where calculating the context score may include calculating the probability, geometric mean, and expected number of occurrences of words co-occurrence versus actual occurrence. The "contextual beacon phrase score" may be calculated, at least in part, based on the probabilistic implications, geometric mean, and co-occurrence number of phrases in the user's campaign data across web pages in the cache database 107 or otherwise published by the RTB service. The "context score" for a particular web page may be calculated, at least in part, based on the number of times the beacon phrase and / or in-context phrase occur on the web page, the fraction of instances each word appears on the web page, and the beacon phrase score of the beacon phrase. In some implementations, the computer quickly calculates the web page context scores in real time at the time of bidding, while the computer may be more computationally intensive (and slower) before bidding time when the campaign is deployed for competition. The result of the current step 307 is that the computer generates web page scores for particular web pages stored in the cache database 107. The web pages may be "ordered" by their web page context scores.

[0060] Based on the user's input phrase and the beacon phrase, the final calculated web page score targets specific keywords or demographics, and thus the web page score can be customized for each campaign. In some cases, the web page scores can be grouped for specialized use in certain verticals, such as automotive, insurance, consumer electronics, toys, healthcare, and government. These grouped web page scores can be reused for similar campaigns.

[0061] In step 308, the computer bids on behalf of the user on the top X% of available web page URLs, or otherwise those with web page scores above a predefined threshold score. The threshold percentage may be predefined by the user, an administrator, or automatically by an algorithm (highest percentage of budget consumed) intended to consume the entire budget. The computer may generate a GUI that displays those web pages with high context scores to the end user, and the computer delivers their content for publication. The computer may implement any number of additional or alternative bid volume thresholds to control or limit the number (volume) of submitted bids.

[0062] A campaign typically has a specified budget, and the length of the campaign and the budget may determine the X% value (or other criteria that serve as the bid quantity threshold). In step 309, the system bids only on context pages that meet the bid threshold requirement. An algorithm may automatically select a percentile for spending the budget on most in-context pages, but the X% value may also be manually specified. Thus, when the system receives a bid request, it may read the page score of the Internet content associated with the bid request and determine whether the page score of the Internet content exceeds the X% percentile for each campaign (or other priority criteria). If the page score exceeds the bid threshold, the system may bid to place the campaign content on the Internet content.

[0063] In some cases, a conflict may arise when two campaigns on the system bid for the same placement. In these situations, the system may use different approaches. In one embodiment, the system may use a round-robin arbitration method, giving each campaign an opportunity to bid for the placement. In another embodiment, the bid amount may vary depending on the delta between the page score and each campaign's X%. The best campaign for placement wins. Embodiments also include hybrid approaches. For example, if campaign A is winning much more frequently than campaign B, campaign B may be given an opportunity to win after a predetermined number of wins for campaign A. For example, campaign B will make a higher winning bid after campaign A has nine wins. Other embodiments include determining the bid price based on other aspects of the bid request / campaign combination, such as a predicted click-through rate, and selecting the campaign with the highest bid price.

[0064] In addition to the page score, in some embodiments, the computer (or other computing device of the system) receives or calculates a performance score that defines the likelihood that a campaign content posting will be viewed, clicked through, or engaged with. The click-through rate represents the number or rate / ratio at which users click on the campaign content, while engagement may be determined by other means, such as whether a user scrolls through the content, views the content for a predetermined amount of time (e.g., 15 seconds), clicks on something in the campaign content, or views a predetermined portion of the campaign content, such as additional audio or video content. The computer or other device of the system is configured to track this information using various cookies or other programming on the third-party host server and / or receive this tracked information from the bidding server or third-party host server. Thus, the computer or other device of the system may calculate a performance score for a bid request before placing a bid. The computer submits a bid when the performance score meets a preset threshold performance score. The performance score may be a measure of the level of interest that viewers of the Internet content have in the campaign content, and in some cases, the system may bid higher or lower on the bid request based on the performance score. For example, the higher the performance score, the more willing the system is to bid on that score, and the lower the performance score, the less willing the system is to bid on that page. In this manner, the determination of a high or low bid price may function as a shifted performance score threshold. The system may also consider many other factors related to the bid request, such as the time of the request, the location of the server to which the bid request was sent, the IP address of the bid requester, the bid requester's past performance scores and payments, the likelihood that the bid requester will accept the bid, the payment terms of the bid, etc.

[0065] FIG. 4 illustrates a method 400 according to one embodiment. Step 401 includes receiving a first set of terms from a client device, the first set including one or more in-context terms. An embodiment may also allow for receiving out-of-context terms at this step. The method further includes, in step 402, calculating multiple context scores based on the first set of terms. The context score may represent the importance or lack of importance of a particular term compared to other terms. In step 403, the method may identify a set of beacon terms in a corpus database based on the context scores, each beacon term having a context score exceeding a predetermined threshold. The beacon terms may be related to the terms received in step 401. The method may further include step 404, during which the method may calculate multiple page scores from preprocessed pages stored in the corpus database. As described above, the page score determines the importance of a particular piece of Internet content for a campaign based on the context and the page score. Step 404 includes identifying a set of context web pages in the corpus database having page scores that meet a threshold. As mentioned above, this may be a set of URLs that the system has identified as having a particular relevance. Next, in step 405, the method may transmit data for displaying the set of beacon terms and the set of context web pages on the client device. A user of the client device may evaluate whether the set of context web pages is accurate. If not, in step 405, the system may receive a second set of context terms from the client device, the second set including one or more out-of-context terms. The system may recalculate the page score based on the first set of context terms and the second set of context terms, thereby generating an updated page score, as shown in step 406. Based on this recalculation, step 407 may update the set of beacon terms and the set of context web pages based on the updated page score.Step 408 may include storing the set of context web pages and the set of campaign terms in a campaign database (sometimes referred to as a "context database"), where the set of campaign terms includes each phrase and each beacon term in the context.

[0066] The foregoing method descriptions and process flow diagrams are provided merely as illustrative examples and are not intended to require or imply that the steps of the various embodiments must be performed in the order presented. The steps in the foregoing embodiments may be performed in any order. Words such as "then" and "next" are not intended to limit the order of the steps. These words are used merely to guide the reader through the method description. Although the process flow diagrams may describe operations as a sequential process, many of the operations may be performed in parallel or simultaneously. Additionally, the order of operations may be rearranged. A process may correspond to a method, a function, a procedure, a subroutine, a subprogram, etc. When a process corresponds to a function, process termination may correspond to a return of the function to the calling function or the main function.

[0067] The various illustrative logic blocks, modules, circuits, and algorithm steps described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, the various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software depends on the particular application and design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each particular application, but such implementation decisions should not be interpreted as causing a departure from the scope of the present disclosure or the claims.

[0068] Computer software-implemented embodiments may be implemented in software, firmware, middleware, microcode, hardware description languages, or any combination thereof. A code segment or machine-executable instruction may represent a procedure, a function, a subprogram, a program, a routine, a subroutine, a module, a software package, a class, or any combination of instructions, data structures, or program statements. A code segment may be coupled to another code segment or a hardware circuit by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc. may be passed, forwarded, or transmitted via any suitable means including memory sharing, message passing, token passing, network transmission, etc.

[0069] The actual software code or specialized control hardware used to implement these systems and methods does not limit the claimed features or this disclosure. Accordingly, the operation and behavior of the systems and methods have been described without reference to specific software code, with the understanding that software and control hardware may be designed to implement the systems and methods based on the description herein.

[0070] If implemented in software, the functions may be stored as one or more instructions or code on a non-transitory computer-readable or processor-readable storage medium. The steps of a method or algorithm disclosed herein may be embodied in a processor-executable software module, which may reside on a computer-readable or processor-readable storage medium. Non-transitory computer-readable or processor-readable media include both computer storage media and tangible storage media that facilitate transfer of a computer program from one place to another. Non-transitory processor-readable storage media may be any available medium that can be accessed by a computer. By way of example, and not limitation, such non-transitory processor-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage, or any other tangible storage medium that can be used to store desired program code in the form of instructions or data structures and that can be accessed by a computer or processor. Disk and disc, as used herein, include compact discs (CDs), laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically while discs reproduce data optically with a laser. Combinations of the above should also be included within the scope of computer-readable media. Furthermore, the operations of a method or algorithm may reside as one or any combination or set of code and / or instructions on a non-transitory processor-readable medium and / or computer-readable medium, which may be incorporated into a computer program product.

[0071] The previous description of the disclosed embodiments is provided to enable any person skilled in the art to make or use the embodiments described herein and variations thereof. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other embodiments without departing from the spirit or scope of the subject matter disclosed herein. Thus, the present disclosure is not intended to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the following claims and the principles and novel features disclosed herein.

[0072] While various aspects and embodiments have been disclosed, other aspects and embodiments are contemplated. The various aspects and embodiments disclosed are for purposes of illustration and are not intended to be limiting, with the true scope and spirit being indicated by the following claims.

Claims

1. 1. A computer-implemented method comprising: applying, by a computer, a machine learning model to the set of first context terms received from the client device to output a set of beacon terms from a corpus database, the machine learning model being trained on a plurality of corpus terms stored in the corpus database to determine a plurality of co-occurrence probabilities corresponding to the plurality of corpus terms; calculating, by the computer, a plurality of page scores for a plurality of corpus web pages stored in the corpus database based on the set of beacon terms and the first set of context terms; identifying, by the computer, a set of context web pages in the plurality of corpus web pages that have page scores that meet a threshold; applying, by the computer, the machine learning model to the first set of context terms and the second set of context terms received from the client device to output an updated set of beacon terms; calculating, by the computer, one or more updated page scores for one or more corpus web pages stored in the corpus database based on the updated set of beacon terms, the first set of context terms, and the second set of context terms; updating, by the computer, the set of context web pages based on the one or more updated page scores; and storing, by the computer, campaign data of the user including the set of beacon terms and the set of context web pages in a campaign database, the campaign data being configured to execute the real-time bid selection operation of one or more available web pages during a real-time bid selection operation.

2. applying the machine learning model to the first set of context terms to output the set of beacon terms from the corpus database; 2. The method of claim 1, comprising extracting, by the computer, one or more embeddings for the first set of context terms, wherein the set of beacon terms includes one or more corpus terms in the corpus database having feature vectors that satisfy a threshold distance from the one or more embeddings for the first set of context terms.

3. 3. The method of claim 2, further comprising: calculating, by the computer, a context score for each context term in the first set of context terms based on at least one of word probability, word co-occurrence, and inter-word probabilistic implications, wherein each embedding of each context term to identify the one or more beacon terms is based on the context score of the context term.

4. setting, by the computer, each context score as a default score for each context term; The method of claim 3 , further comprising: adjusting, by the computer, the context scores of the context terms according to one or more user inputs.

5. The method of claim 1 , wherein the first set of context terms includes at least one out-of-context term.

6. generating, by the computer, context data for the set of context web pages based on one or more user inputs received from the client device; The method of claim 1 , further comprising transmitting, by the computer, the contextual data to the client device for display on a graphical user interface (GUI) of the client device.

7. The method of claim 1 , wherein the computer receives each context term from the client device via a query formulation web page.

8. requesting, by the computer, web page data for a web page according to a uniform resource locator (URL), the web page including page content and metadata; 3. The method of claim 2, further comprising: storing, by the computer, the web page data in the corpus database as a preprocessed page, the preprocessed page being stored in a data record for the preprocessed page, the data record including the page content and at least a portion of the metadata, a timestamp, the URL, and one or more content tags.

9. receiving, by the computer, from a bidding server an availability list of one or more available web pages requesting bids from a bidding system; 2. The method of claim 1, further comprising: calculating, by the computer, a real-time page score for each available web page in the availability list based at least in part on the occurrence of one or more campaign terms, including the updated set of beacon terms, and one or more in-context terms within the available web page.

10. The method of claim 9 , further comprising identifying, by the computer, a bid list of web pages that includes each of the available web pages in the availability list that meets the bid threshold.

11. 1. A system comprising: a corpus database including a non-transitory storage medium configured to store at least a portion of a plurality of Internet content data corresponding to uniform resource locators (URLs) and page context scores; a campaign database including a non-transitory storage medium configured to store campaign data of a plurality of users, the campaign data configured to execute a real-time bid selection operation of one or more available web pages during the real-time bid selection operation; 1. A server including a processor, the processor comprising: applying a machine learning model to the set of first context terms received from the client device to output a set of beacon terms from a corpus database, the machine learning model being trained on a plurality of corpus terms stored in the corpus database to determine a plurality of co-occurrence probabilities corresponding to the plurality of corpus terms; and calculating a plurality of page scores for a plurality of corpus web pages stored in the corpus database based on the set of beacon terms and the first set of context terms; identifying a set of context web pages in the plurality of corpus web pages having page scores that meet a threshold; applying the machine learning model to the first set of context terms and the second set of context terms received from the client device to output an updated set of beacon terms; calculating one or more updated page scores for one or more corpus web pages stored in the corpus database based on the updated set of beacon terms, the first set of context terms, and the second set of context terms; updating the set of context web pages based on the one or more updated page scores; and a server configured to: store in the campaign database user campaign data including the set of beacon terms and the set of context web pages, the campaign data configured to execute the real-time bid selection operation of one or more available web pages during a real-time bid selection operation.

12. The server: generating context data for the set of context web pages based on one or more user inputs received from the client device; and The system of claim 11 , further configured to transmit the context data to the client device for display on a graphical user interface (GUI) of the client device.

13. applying the machine learning model to the first set of context terms to output the set of beacon terms from the corpus database; 12. The system of claim 11, comprising extracting one or more embeddings for the first set of context terms, wherein the set of beacon terms includes one or more corpus terms in the corpus database having feature vectors that satisfy a threshold distance from the one or more embeddings for the first set of context terms.

14. The server: receiving from a bidding server an availability list of one or more available web pages requesting bids from the bidding system; and 12. The system of claim 11, further configured to calculate a real-time page score for each available web page in the availability list based at least in part on some occurrences of one or more campaign terms including the updated set of beacon terms and one or more in-context terms within the available web page.

15. A computer-readable medium containing machine-executable program instructions, wherein execution of the program instructions by one or more processors of a computer system causes the one or more processors to: applying a machine learning model to a set of first context terms received from a client device to output a set of beacon terms from a corpus database, the machine learning model being trained on a plurality of corpus terms stored in the corpus database to determine a plurality of co-occurrence probabilities corresponding to the plurality of corpus terms; calculating a plurality of page scores for a plurality of corpus web pages stored in the corpus database based on the set of beacon terms and the first set of context terms; identifying a set of context web pages in the plurality of corpus web pages that have page scores that meet a threshold; applying the machine learning model to the first set of context terms and the second set of context terms received from the client device to output an updated set of beacon terms; calculating one or more updated page scores for one or more corpus web pages stored in the corpus database based on the updated set of beacon terms, the first set of context terms, and the second set of context terms; updating the set of context web pages based on the one or more updated page scores; and storing the user's campaign data, including the set of beacon terms and the set of contextual web pages, in a campaign database, the campaign data being configured to execute the real-time bid selection operation for one or more available web pages during a real-time bid selection operation.

16. the one or more processors:

16. The computer-readable medium of claim 15, further performing the step of extracting one or more embeddings for the first set of context terms, wherein the set of beacon terms includes one or more corpus terms in the corpus database having feature vectors that satisfy a threshold distance from the one or more embeddings for the first set of context terms.

17. the one or more processors:

17. The computer-readable medium of claim 16, further comprising: calculating a context score for each context term in the first set of context terms based on at least one of word probability, word co-occurrence, and inter-word probabilistic implications, wherein each embedding of each context term to identify the one or more beacon terms is based on the context score of the context term.

18. the one or more processors: setting each context score as a default score for each context term; and adjusting the context scores of the context terms according to one or more user inputs.

19. the one or more processors: generating context data for the set of context web pages based on one or more user inputs received from the client device; 16. The computer-readable medium of claim 15, further comprising: transmitting the contextual data to the client device for display on a graphical user interface (GUI) of the client device.

20. the one or more processors: receiving, by the computer, from a bidding server an availability list of one or more available web pages that request bids from a bidding system; 16. The computer-readable medium of claim 15, further comprising: calculating, by the computer, a real-time page score for each available web page in the availability list based at least in part on the occurrence of one or more campaign terms including the updated set of beacon terms and one or more in-context terms within the available web page.