Automated system and method for entity-specific proposition generation using artificial intelligence

An AI-driven system automates entity-specific proposition generation through web scraping and validation, addressing inefficiencies in existing methods by providing timely and relevant insights.

WO2025175343A1PCT designated stage Publication Date: 2025-08-28HEATSEEKER TECH PTY LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/AU2025/050136
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-02-19
Filing Date
2025-02-19
Publication Date
2025-08-28

AI Technical Summary

Technical Problem

Entities face challenges in navigating vast information landscapes to extract relevant insights efficiently, leading to resource-intensive and outdated content generation processes.

Method used

An automated system using artificial intelligence to generate entity-specific propositions by web scraping, crawling, and structuring queries with large language models, incorporating multiple agent constructs for validation and content package generation.

Benefits of technology

Facilitates efficient and cost-effective generation of entity-specific propositions, ensuring relevance and viability through AI-driven validation, reducing manual effort and time lag.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure AU2025050136_28082025_PF_FP_ABST
    Figure AU2025050136_28082025_PF_FP_ABST
Patent Text Reader

Abstract

Some embodiments relate to a system and an automated computer-implemented method for entity-specific proposition generation using artificial intelligence. An example method comprises: receiving a seed dataset from one or more entity-specific devices and extracting one or more seed URLs from the seed dataset; performing a web scrape on the one or more seed URLs to generate a first dataset; extracting one or more scraped URLs from the first dataset and performing a web crawl on the one or more scraped URLs to generate a second dataset; structuring an entity-specific query based on the first and second datasets; receiving, by providing the entity-specific query to a first large language model, one or more key data points in relation to the specific entity; and generating one or more entity-specific propositions based on the key data points.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] "Automated system and method for entity-specific proposition generation using artificial intelligence"

[0002] Technical Field

[0003] [1] Embodiments of this disclosure generally relate to using artificial intelligence to generate propositions. In particular, embodiments of this disclosure relate to entityspecific proposition generation using artificial intelligence.

[0004] Background

[0005] [2] It is difficult for entities, such as businesses, to navigate an ever-increasing information landscape in which a vast amount of information is available. A large number of manhours is generally required to filter through this data to determine relevancy, interpret it, and extract insights from which the business can learn and grow, making it a resource and cost intensive endeavour.

[0006] [3] Standard methods of conducting entity -specific research for content generation involve assembling a team of individuals to gather data, analyse the data for relevant information, and extract key points, from which a person, or persons, can infer insights relevant to that entity. These insights can then be used by the entity to create content and determine how best to disseminate information. However, in the time that it takes for the data to be gathered, analysed, and to have insights inferred from it by the team of people, the themes or context from which the data was gathered (and propositions formed) can change, resulting in the disseminated information no longer being as impactful as it otherwise could be.

[0007] [4] It is desired to address or ameliorate one or more shortcomings or disadvantages of prior proposition generation methods and systems, or to at least provide a useful alternative thereto. [5] Throughout this specification the word "comprise", or variations such as "comprises" or "comprising", will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers or steps, but not the exclusion of any other element, integer or step, or group of elements, integers or steps.

[0008] [6] Any discussion of documents, acts, materials, devices, articles or the like which has been included in the present specification is not to be taken as an admission that any or all of these matters form part of the prior art base or were common general knowledge in the field relevant to the present disclosure as it existed before the priority date of each of the appended claims.

[0009] Summary

[0010] [7] Some embodiments relate to an automated computer-implemented method for entity-specific proposition generation using artificial intelligence, the method may comprise: receiving a seed dataset from one or more entity -specific devices and extracting one or more seed URLs from the seed dataset; performing a web scrape on the one or more seed URLs to generate a first dataset; extracting one or more scraped URLs from the first dataset and performing a web crawl on the one or more scraped URLs to generate a second dataset; structuring an entity- specific query based on the first and second datasets; receiving, by providing the entity- specific query to a first large language model, one or more key data points in relation to the specific entity; and generating one or more entity -specific propositions based on the key data points.

[0011] [8] The method may further comprise: providing instructions to a second large language model to generate multiple agent constructs, wherein the multiple agent constructs are each configured with different modelled behaviour parameters; subsequently, structuring a validation query to the second large language model based on the one or more entity -specific propositions; receiving, by providing the validation query to the second large language model, proposition validation data based on the interaction of the multiple agent constructs with the one or more entity- specific propositions indicative of the validity of each of the one or more entity- specific propositions.

[0012] [9] The different modelled behaviour parameters may include one or more of: age, gender, ethnicity, education, income, employment status, relationship status, personality, and location. The different modelled behaviours may be determined, in part, on one or more of the seed dataset, the first dataset, and the second dataset.

[0013]

[0010] The method may further comprise: generating one or more content packages based on the generated entity- specific propositions and proposition validation data; and publishing the one or more content packages to one or more platforms.

[0014]

[0011] Each of the one or more content packages may be platform specific content packages. The content packages may be generated for a target user. The one or more content packages may be generated based on one or more of the target user’s demographics, industry, and location. The one or more content packages may include one or more of: text, image, audio, and video.

[0015]

[0012] The method may further comprise, prior to publishing the one or more content packages: providing the one or more content packages to the second large language model; subsequently, structuring a query to the second large language model based on the one or more content packages; receiving, by providing the query to the second large language model, content package validation data based on the interaction of the multiple agent constructs with the one or more content packages indicative of the validity of each of the one or more content package; and wherein the publishing includes publishing the one or more content packages to one or more platforms based on the content package validation data.

[0016]

[0013] The method may further comprise: retrieving performance metrics of the one or more content packages from the one or more platforms; and generating performance results in relation to the one or more entity- specific propositions based on the retrieved performance metrics.

[0014] Performing the web scrape may extract markup language code from the one or more seed URLs to generate the first dataset. Performing the web crawl may extract markup language code from the one or more scraped URLs to generate the second dataset.

[0017]

[0015] The method may further comprise, prior to structuring the entity -specific query: applying retrieval-augmentation generation by providing one or more of the seed dataset, the first dataset, and the second dataset to the first large language model.

[0018]

[0016] The method may further comprise, prior to structuring the entity -specific query: providing one or more frameworks to the first large language model; and wherein structuring the entity- specific query is further based on the provided one or more framework. The one or more frameworks may include one or more of: double diamond, PESTEL, and design thinking.

[0019]

[0017] Some embodiments relate to a system for automated entity- specific proposition generation using artificial intelligence, the system comprising a server, the server may comprise: processing circuitry; memory accessible to the processing circuitry; a communications module; wherein the memory stores instructions that, when executed by the processing circuitry, cause the processing circuitry to perform the method as previously described.

[0020]

[0018] The system may further comprise a large language model module, wherein the entity-specific query may be provided to the large language model module. Instructions to generate multiple agent constructs may be provided to the large language model module, and wherein the validation query may be provided to the large language model module.

[0021]

[0019] The system may further comprise one or more entity- specific devices, wherein the one or more entity- specific devices are in communication with the server via the communications module.

[0020] Some embodiments relate to a non-transitory computer readable medium having one or more computer program instructions configured to perform a method, the method may comprise: receiving a seed dataset from one or more entity- specific devices and extracting one or more seed URLs from the seed dataset; performing a web scrape on the one or more seed URLs to generate a first dataset; extracting one or more scraped URLs from the first dataset and performing a web crawl on the one or more scraped URLs to generate a second dataset; structuring an entity- specific query based on the first and second datasets; receiving, by providing the entity- specific query to a first large language model, one or more key data points in relation to the specific entity; and generating one or more entity -specific propositions based on the key data points.

[0022] Brief Description of Drawings

[0023]

[0021] Embodiments of the present disclosure will now be described in further detail, by way of non-limiting example, with reference to the accompanying drawings, in which:

[0024]

[0022] Figure 1 shows a block diagram of an entity- specific proposition generation system, according to some embodiments;

[0025]

[0023] Figure 2 shows a block diagram of a proposition generation server, according to some embodiments;

[0026]

[0024] Figure 3 shows a process flow diagram of a method of generating entityspecific propositions using artificial intelligence, according to some embodiments;

[0027]

[0025] Figure 4, shows a process flow diagram of a method of validating one or more generated entity- specific propositions, according to some embodiments;

[0028]

[0026] Figure 5 shows a process flow diagram of a method of generating entityspecific propositions using artificial intelligence including retrieval-augmentation generation, according to some embodiments; and

[0027] Figure 6 shows a process flow diagram of a method of generating entityspecific propositions using artificial intelligence including framework application, according to some embodiments; and

[0029]

[0028] Figure 7 shows a block diagram of an example computer system according to some embodiments.

[0030] Description of Embodiments

[0031]

[0029] Embodiments of this disclosure generally relate to using artificial intelligence to generate propositions. In particular, embodiments of this disclosure relate to entityspecific proposition generation using artificial intelligence.

[0032]

[0030] Propositions relating to a specific entity, such as a person, a company, an organisation, a non-profit organisation, or a government body, for example, may be used by individuals or employees within the entity to promote growth and implement positive change in their industry, field, or target user base. A plurality of propositions may be generated at the same time, wherein each proposition is unique within the generated plurality of propositions and includes an associated description.

[0033]

[0031] Propositions may provide individuals with information pertaining to their own businesses, organisations, competitors, and / or exemplars within their industry, field, or target user base. This information can then be used by the individual or employee as a basis for content generation which can then aid to drive change within their industry, field, or target user base. For example, a non-profit may generate one or more propositions in relation to increasing cancer awareness in a target user base to generate content packages relating to sun cancer awareness, analyse the performance (or likely performance) of the content packages using artificial intelligence and then utilise the well performing content packages to drive awareness with the public.

[0034]

[0032] A proposition may be or include one or more of the following components: a title summarising the proposition’s purpose or outcome; entity metadata specifying the entity’s context, such as type, industry, location, or demographics, for example; core recommendation / proposal being an actionable statement addressing the entity’s needs or objective; supporting evidence including key data points or analytical insights justifying the recommendation, which may include trends, benchmarks, or specific findings, for example; validation results, which may include artificial intelligence assessment outcomes, such as confidence scores and / or agent construct validations; and content packages, which may include text, image, audio, and / or video content, associated with the proposition for immediate implementation.

[0035]

[0033] For example, a proposition may include a title reading “Implement Video Marketing Campaign Targeting Millennials”, provide a core recommendation to “Focus on creating short-form videos for platforms like Instagram and TikTok to drive engagement with millennial audiences”, include supporting evidence showing that “Behavioural analysis indicates a 65% preference for video content among target demographics” and “Competitor analysis shows similar strategies yielding 25% higher engagement rates”, and validate the results through artificial intelligence assessment outcomes, such as “Confidence Score: 90%, with high approval from simulated user personas aged 25-35”.

[0036]

[0034] Propositions may be utilised to digest a large amount of data and provide key areas of exploration or focus within the data. For example, research data may be digested and analysed by an artificial intelligence model, propositions as to potential research pathways in areas of scientific endeavour including medical technologies are then generated by the artificial intelligence model, and wherein the propositions are generated in relation to pathways not yet explored or not yet indicated to be unviable by the digested research data.

[0037]

[0035] Manually gathering data, analysing its contents, and extracting key data points to infer insights and generate propositions is an arduous and expensive process.

[0038] Therefore, it may be beneficial to provide a process for automatically and efficiently performing this process in a cost-effective manner. Further, validating these insights and testing them within the market by producing content packages, prior to implementing them, to ensure their viability, would also be a resource and cost intensive process.

[0039]

[0036] Referring to Figure 1, there is shown an entity- specific proposition generation system 10, according to some embodiments. The entity -specific proposition generation system 10 comprises a proposition generation server 100. The entity-specific proposition generation system 10 may comprise an external data store 105. Proposition generation server 100, hereinafter referred to as server 100, comprises processing circuitry 102 and memory 104 accessible to the processing circuitry 102. The processing circuitry 102 may be configured to access data stored in memory 104, to execute instructions stored in memory 104, and to read and write data to and from memory 104. Processing circuitry 102 may comprise one or more microprocessors, microcontrollers, central processing units (CPUs), application specific instruction set processors (ASIPs), or other computer processor capable of reading and executing instruction code. Processing circuitry 102 may be distributed across multiple computing devices.

[0040]

[0037] Memory 104 may comprise one or more volatile or non-volatile memory types, such as RAM, ROM, EEPROM, or flash, for example. Memory 104 may be configured to store executable applications or instruction code for execution by the processing circuitry 102. Memory 104 may store executable software code modules, such as seed data retrieval module 120 and framework integration module 150, for execution by processing circuitry 102, to be further described below in relation to Figure 2. In some embodiments, software code modules or data described as part of memory 104 may be stored in the external data store 105 rather than in memory 104.

[0041]

[0038] Memory 104 stores executable software code modules and data, including: seed data retrieval module 120, seed dataset 122, web scrape module 124, first data set 126, web crawl module 128, second dataset 130, query structure module 132, key data points 134, proposition generation module 136, generated propositions 138, agent construct generation module 140, validation module 142, content generation, validation, and publishing module 144, content package performance analysis module 146, retrieval-augmentation generation module 148, framework integration module 150 and API integration module 152. In some embodiments, memory 104 stores communications module 106 and / or LLM module 108. The functions and use of these modules and data are described further below in relation to example processes / methods 300, 400, 500 and 600.

[0042]

[0039] To facilitate communication with external and / or remote devices, server 100 further comprises a communications module 106. Communications module 106 may allow for wired and / or wireless communication between server 100 and external computing devices and components, such as client server 210, client device(s) 211, external web pages 215, and / or one or more external large language models (LLM) 220. Communications module 106 may facilitate communication via Bluetooth, USB, Wi-Fi, Ethernet, or via a telecommunications network, for example. According to some embodiments, communication module 106 may facilitate communication with external devices and systems via a network 200.

[0043]

[0040] Network 200 may comprise one or more local area networks or wide area networks that facilitate communication between elements of Figure 1. For example, according to some embodiments, network 200 may include the Internet and subnetworks to communicate between computing devices via the Internet. However, network 200 may comprise at least a portion of any one or more networks having one or more nodes that transmit, receive, forward, generate, buffer, store, route, switch, process, or a combination thereof, etc. one or more messages, packets, signals, some combination thereof, or so forth. Network 200 may include, for example, one or more of: a wireless network, a wired network, an internet, an intranet, a public network, a packet-switched network, a circuit-switched network, an ad hoc network, an infrastructure network, a public- switched telephone network (PSTN), a cable network, a cellular network, a satellite network, a fibre-optic network, or some combination thereof.

[0044]

[0041] Server 100 may be in communication with a client server 210. Client server 210 may be accessible to server 100 via network 200 using communications module 106. Server 100 may query the client server 210 and / or retrieve data, such as client data and / or seed data, from client server 210 via network 200 using communications module 106.

[0045]

[0042] Server 100 may be in communication with one or more client devices 211. The one or more client devices 211 may be accessible to server 100 via network 200 using communications module 106. Server 100 may query the one or more client devices 211 and / or retrieve data, such as client data and / or seed data, from the one or more client devices 211 via network 200 using communications module 106. The client server 210 and the one or more client devices 211 may be considered entity -specific devices, for example.

[0046]

[0043] Server 100 may be in communication with one or more external web pages 215. The one or more external web pages 215 may be accessible to server 100 via network 200 using communications module 106. Server 100 may access the one or more external web pages 215 and / or retrieve data, such as client data and / or competitor data, from the one or more external web pages 215 via network 200 using communications module 106.

[0047]

[0044] Server 100 may be in communication with one or more external LLMs 220. The one or more external LLMs 220 may be accessible to server 100 via network 200 using communications module 106. Server 100 may query the one or more external LLMs and receive data, such as key data points in relation to a specific entity, from the one or more external LLMs 220 via network 200 using communications module 106. The one or more external LLMs 220 may be readily available / accessible LLMs, such as ChatGPT, Gemini, LaMDA, Cohere, Antropic, Falcon, Llama2, BERT, and Mistral, for example.

[0048]

[0045] In some embodiments, server 100 further includes an LLM module 108. That is, server 100 may query LLM module 108 instead of, or in combination with, querying the one or more external LLMs 220. LLM module 108 may include one or more LLMs accessible by processing circuitry 102. That is, processing circuitry 102 may execute instructions stored in memory 104 to structure a query to LLM module 108.

[0049]

[0046] The LLM module 108 and / or the one or more external LLMs 220 may be trained on a combination of general-purpose datasets and domain specific datasets. General-purpose datasets may include publicly available corpora such as books, articles, encyclopedias, and online content to establish broad language understanding, for example. Domain specific datasets may include industry reports, proprietary organisational datasets, and structured entity -specific inputs (e.g., seed data, product information, competitor insights) obtained by the proposition generation server 100. The proposition generation server 100 may further enhance the contextual knowledge of the LLM module 108 and / or the one or more external LLMs 220 by providing it web scraped data and web crawled data from web pages associated with the relevant entity, such as websites, social media, and competitor platforms. The web scraped data and web crawled data is initially in an unstructured form, but the proposition generation server 100 uses an external LLM 220 to turn it into structured data, which is then stored in memory 104 or data store 105.

[0050]

[0047] Referring to Figure 2, there is shown a block diagram of the proposition generation server 100, according to some embodiments. As previously described in relation to Figure 1, server 100 comprises processing circuitry 102, memory 104, and communications module 106. Further, server 100 may further comprise LLM module 108. Memory 104 may store modules accessible and executable by processing circuitry 102 to perform entity-specific proposition generation in relation to method 300, 400, 500, and 600, to be described in detail below. Memory 104 may store data in one or more formats, including, but not limited to: CSV, JSON, AVRO, and XML, for example.

[0051]

[0048] Referring to Figure 3, there is shown a process flow diagram of a method 300 of generating entity- specific propositions using artificial intelligence, according to some embodiments. At step 302 of method 300, processing circuitry 102 executes instructions, stored in memory 104, to cause server 100 to communicate with the client server 210 and / or the one or more client devices 211 via communications module 106.

[0052]

[0049] Server 100 communicates with the client server 210 and / or the one or more client devices 211 to receive data associated with the specific entity. Processing circuitry 102 may process the data received via the communications module 106 and store it in memory 104 as seed dataset 122, for example. Seed dataset 122 comprises data associated with the specific entity for which propositions are to be generated using method 300. For example, the seed dataset 122 may comprise raw data specific to the entity such as URLs pointing to the entity’s own content, reports, or user-provided datasets. In some embodiments, processing circuitry 102 may generate the seed dataset 122 from the received data associated with the specific entity and store the seed dataset 122 in memory 104 in a particular format. This format may include, but is not limited to: CSV, JSON, AVRO, and XML, for example.

[0053]

[0050] The seed dataset 122 may comprise entity -specific data relating to corporate information, such as mission statements, value propositions, organisational charts, and / or annual reports, for example. The seed dataset 122 may comprise entity -specific data relating to product or service details, such as descriptions, specifications, pricing, and / or customer reviews, for example. The seed dataset 122 may comprise entityspecific data relating to website content, such as text, metadata, and / or links from the specific entity’s website or landing page, for example. The seed dataset 122 may comprise entity- specific data relating to social media data, such as posts, follower demographics, engagement metrics, and / or campaign histories, for example.

[0054]

[0051] The seed dataset 122 may comprise customer data relating to customer feedback, such as survey results, support tickets, and / or reviews, for example. The seed dataset 122 may comprise customer data relating to user demographics, such as age, location, preferences and / or purchase histories, for example. The seed dataset 122 may comprise customer data relating to behavioural analytics, such as website heatmaps, clickstream data, and / or app usage patterns, for example.

[0052] The seed dataset 122 may comprise market data relating to competitor information, such as details scraped from competitor websites, marketing materials, and / or product reviews, for example. The seed dataset 122 may comprise market data relating to industry trends, such as insights from public reports, blogs, and / or news articles relevant to the specific entity, for example. The seed dataset 122 may comprise market data relating to geographic data, such as regional and / or local market insights tied to the specific entity’s operating areas, for example.

[0055]

[0053] The seed dataset 122 may comprise operational data relating to internal documents, such as business plans, process flow diagrams, and / or internal project reports, for example. The seed dataset 122 may comprise operational data relating to sales data, such as customer relationship management (CRM) exports, revenue figures, and / or lead conversion rates, for example. The seed dataset 122 may comprise operational data relating to performance metrics, such as key performance indicators (KPIs) related to marketing, product, and / or service delivery, for example.

[0056]

[0054] The seed dataset 122 may comprise third party data relating to API feeds, such as inputs from third party analytics tools (e.g., Google analytics) and / or advertising networks, for example. The seed dataset 122 may comprise third party data relating to partnership data, such as information shared by partners or affiliates of the entity, for example. The seed dataset 122 may comprise third party data relating to public databases, such as relevant entries from government or industry databases (e.g., patents and / or regulatory filings), for example.

[0057]

[0055] After processing the received data and storing it in memory 104, processing circuitry 102 executes instructions to extract one or more seed URLs from the seed dataset 122. That is, processing circuitry 102 processes the seed dataset 122 to identify and extract URLs contained therein, for example. The one or more seed URLs include addresses of a resource on the Web associated with the specific entity. For example, the one or more seed URLs may include the specific entities web page(s) / website(s) and / or social media pages, for example.

[0056] To perform step 302, processing circuitry 102 may execute seed data retrieval module 120 stored in memory 104. Executing the seed dataset retrieval module 120 may cause server 100 to communicate with the client server 210 and / or the one or more client devices 211, receive data associated with the specific entity, process the received data, and extract one or more seed URLs from the received data.

[0058]

[0057] At step 304, processing circuitry 102 executes instructions, stored in memory 104, to cause server 100 to perform a web scrape on the one or more seed URLs extracted from seed dataset 122 at step 302. A web scrape, or web scraping, involves fetching one or more web pages and extracting from it. Fetching is the downloading of each web page to be scraped, and once fetch, extraction occurs. Extraction of each web page may involve parsing the web page, searching the web page for specific data, reformatting the data contained therein, and / or copying the data to another location, such as memory 104.

[0059]

[0058] Server 100 accesses one or more external web pages 215 specific to the extracted one or more seed URLs via network 200 using communications module 106. In accessing the one or more external web pages 215 specific to the one or more seed URLs, server 100 fetches the one or more external web pages 215 and extracts the content of the one or more external web pages 215 accessed to generate a first dataset 126. In some embodiments, server 100 extracts the underlying markup language code of the one or more external web pages 215 accessed to generate the first dataset 126. The first dataset 126 comprises all data accessed by server 100 in performing the web scrape on the one or more seed URLs and may not be filtered prior to generating the first dataset 126.

[0060]

[0059] In some embodiments, processing circuitry 102 may generate the first dataset 126 from the extracted contents of the one or more external web pages 215 and store the first dataset 126 in memory 104 in a particular format. This format may include, but is not limited to: CSV, JSON, AVRO, and XML, for example.

[0060] To perform step 304, processing circuitry 102 may execute web scrape module 124 stored in memory 104. Executing the web scrape module 124 may cause server 100 to access one or more external web pages 215 specific to the one or more seed URLs, copy the contents of the accessed external web pages 215, and generate a first dataset 126 based on the extracted contents. In some embodiments, the web scrape module 124 is configured to perform seed-based filtering. Seed-based filtering involves configuring the web scrape module 124 to curate the data obtained from the one or more accessed external web pages 215 such that the first dataset 126 is generated from data directly related to the specific entity. The seed -based filtering may be performed based on the seed dataset 122. That is, the relevance of data accessed on the one or more external web pages 215 may be determined based on the seed dataset 122, for example.

[0061]

[0061] At step 306, processing circuitry 102 executes instructions, stored in memory 104, to identify and extract one or more scraped URLs from the first dataset 126. That is, the first dataset 126 is processed and scraped URLs extracted from performing the web scrape at step 304 are identified and extracted from the first dataset 126, for example. Processing circuitry 102 then executes instructions to cause server 100 to perform a web crawl on the one or more scraped URLs extracted from the first dataset 126.

[0062]

[0062] A web crawl involves providing one or more seed URLs to a web crawling bot, such as web crawl module 128. In some embodiments, web crawl module 128 may include an off-the-shelf software package, such as ScrapingBee, for example. In some embodiments, web crawl module 128 may be implemented in a browser framework, such as Selenium, for example. The web crawling bot is configured to access the seed URLs and performs a web scrape on each accessed web page of the respective seed URL. From the accessed web pages, the web crawler identifies hyperlinks contained thereon and accesses the URLs contained within the hyperlinks. The web crawling bot then performs a web scrape on the web pages of the URLs identified in the hyperlinks and repeats the process.

[0063] In some embodiments, to minimise the volume of irrelevant data obtained through web crawling, the web crawl module 128 may be configured to perform focused crawling. Focused crawling utilises entity -specific parameters to prioritise content from highly relevant domains, such as competitors or industry sites, for example, while excluding irrelevant sources. The relevant domains are determined based on entity- specific data.

[0063]

[0064] In some embodiments, the web crawling bot may be configured by processing circuitry 102 to include limitations on the web pages by implementing crawler policies, such as a selection policy, a re-visit policy, a politeness policy, and a parallelisation policy, for example. In some embodiments, the web crawling bot may be a focused crawling bot. That is, the web crawling bot may be configured to focus on accessing and extracting data from web pages that are similar to the web pages of the seed URLs, for example. A focused web crawling bot may provide the advantage of only extracting data considered relevant to the specific entity and reducing overall data size of extracted data. In some embodiments, the web crawling bot is configured to be any one of: a path-ascending crawler, a focused crawler, an academic focused crawler, or a semantic focused crawler.

[0064]

[0065] Server 100 accesses one or more external web pages 215 specific to the extracted one or more scraped URLs via network 200 using communications module 106. In accessing the one or more external web pages 215 specific to the one or more scraped URLs, server 100 extracts the content of the one or more external web pages 215 accessed to generate a second dataset 130. In some embodiments, server 100 extracts the underlying markup language code of the one or more external web pages 215 accessed to generate the second dataset 130. The second dataset 130 comprises all data accessed by server 100 in performing the web crawl on the one or more scraped URLs and may not be filtered prior to generating the second dataset 130.

[0065]

[0066] In some embodiments, processing circuitry 102 may generate the second dataset 130 from data extracted by the web crawling bot and store the second dataset 130 in memory 104 in a particular format. This format may include, but is not limited to: CSV, JSON, AVRO, and XML, for example.

[0066]

[0067] To perform step 306, processing circuitry 102 may execute web crawl module 128 stored in memory 104. Executing the web crawl module 128 may cause server 100 to access one or more external web pages 215 specific to the one or more scraped URLs, copy the contents of the accessed external web pages 215, and generate a second dataset 130 based on the extracted contents.

[0067]

[0068] In some embodiments, after generation of each of the first dataset 126 and the second dataset 130, data conversion and categorisation of the respective datasets is performed. That is, data contained in the first dataset 126 and / or the second dataset 130 may be categorised into structured formats, such as by topic or relevance, for example. Categorisation of the first dataset 126 and / or the second dataset 130 may improve contextual understanding of the data by the LLM module 108 and / or the one or more external LLMs 220, for example.

[0068]

[0069] In some embodiments, after generation of each of the first dataset 126 and the second dataset 130, data filtering is performed on the respective datasets. The data filtering may be performed to retain data considered most relevant to the specific entity. The data filtering may be based on keyword and topic matching or relevance scoring based on data contained in the seed dataset 122, for example. In some embodiments, the categorised and filtered data of the first dataset 126 and / or the second dataset 130 is enriched / augmented using external data sources to enhance context and accuracy.

[0069]

[0070] At step 308, processing circuitry 102 executes instructions, stored in memory 104, to structure an entity- specific query, to one or more external LLMs 220, based on the first dataset 126 and the second dataset 130. In some embodiments, structuring the query to the one or more external LLMs 220 may include providing the one or more external LLMs 220 with the first dataset 126 and / or the second dataset 130. Processing circuitry 102 may structure the query in relation to the specific entity for which propositions are being generated using method 300. That is, the query may be structured differently depending on the specific entity for which the method 300 is being performed. Processing circuitry 102 may structure the query as a text phrase, such as “summarise the provided data relating to the specific entity in view of the entities organisational model and current value propositions”, or the like, for example.

[0070]

[0071] Server 100, upon structuring the entity -specific query, communicates, or provides, the query to the one or more external LLMs 220. In some embodiments, the processing circuitry 102 provides the structured query to the LLM module 108. In response to receiving the entity -specific query, the respective LLM, external LLM 220 and / or LLM module 108, provides an LLM-generated response, being one or more key data points 134, in relation to the specific entity. The one or more key data points 134 are based, at least in part, on the first dataset 126, the second dataset 130, and the structured query. Server 100 receives the one or more key data points 134 in relation to the specific entity, stores them in memory 104, and proceeds to step 310.

[0071]

[0072] To perform step 308, processing circuitry 102 may execute query structure module 132 stored in memory 104. Executing the query structure module 132 may cause server 100 to structure a query specific to the entity for which propositions are to be generated, provide the query to an LLM, such as external LLM 220 or LLM module 108, and receive from the LLM key data points 134 relating to the specific entity. In some embodiments, formulating the structured query includes incorporating frameworks such as PESTEL (Political, Economic, Social, Technological, Environmental, and Legal) or design thinking.

[0072]

[0073] At step 310, processing circuitry 102 executes instructions, stored in memory 104, to generate one or more entity -specific propositions 138 based on the key data points 134 received at step 308. In some embodiments, the one or more entity -specific generated propositions 138 may be generated in relation to the specific entity and / or entities related to the specific entity. That is, the generated propositions 138 may be generated in relation to the specific entity, competitors of the entity, or exemplary entities in the same field as specific entity, for example.

[0074] To perform step 310, processing circuitry 102 may execute proposition generation module 136 stored in memory 104. Executing the proposition generation module 136 may cause server 100 to structure a proposition query, provide the proposition query to the external LLM 220 and / or the LLM module 108. In response to receiving the proposition query, the respective LLM provides an LLM-generated response, being one or more propositions 138 in relation to the specific entity.

[0073]

[0075] In some embodiments, the proposition query includes components relating to the specific entity’s context, experimental data (e.g., engagement metrics, performance scores, or audience segmentation), and a desired LLM output format. The proposition query may further include a pain point, being an issue that the specific entity is experiencing. The proposition query may further include instructions for the LLM to incorporate external data sources, such as Google Trends application programming interface (API), for example. The proposition query may further include instructions for the LLM to structure the LLM generated response in accordance with a particular framework, such as SWOT (strengths, weaknesses, opportunities, and threats) or PESTEL, for example.

[0074]

[0076] An example of a proposition query instruction for a SaaS company may be “Provide two actionable recommendations to improve adoption rates. Structure the output in JSON format and include confidence scores. Optionally, validate results using Linkedln company data API to analyse small business trends.” This example proposition query may include experimental and trial data, for example. An example LLM generated response to this proposition query may be “add a feature for seamless integration with popular tools like Slack and Trello” or “create video tutorials showcasing how to set up integrations effortlessly.” The LLM generated response may also include a suggested target audience and a confidence score, for example. The following shows the described example proposition query in JSON format:

[0075] {

[0076] “entity”: “A B2B SaaS company providing workflow automation tools for small businesses ”, “experiment_data”: {

[0077] “trial_data”: {

[0078] “conversion_rate”: 15,

[0079] “pain_point”: “Integration challenges”, “preferred feature”: “One-click setup” }

[0080] },

[0081] “instruction”: “Provide two actionable recommendations to improve adoption rates. Structure the output in JSON format and include confidence scores. Optionally, validate results using Linkedln company data API to analyse small business trends.” }

[0082]

[0077] The following shows the described example LLM generated response to the proposition query in JSON format:

[0083] [

[0084] {

[0085] “title”: “Introduce One-Click Integration”,

[0086] “description”: “Add a feature for seamless integration with popular tools like Slack and Trello ”,

[0087] “target_audience”: “Small businesses with existing software stacks”, “confidence_score”: 0.92

[0088] },

[0089] {

[0090] “title”: “Launch Educational Campaigns”,

[0091] “description”: “Create video tutorials showcasing how to set up integrations effortlessly.”,

[0092] “target audience”: “Small business owners unfamiliar with automation tools”,

[0093] “confidence_score”: 0.87

[0094] }

[0095] ]

[0078] In some embodiments, prior to generating one or more entity specific propositions 138, the key data points 134 are structured and organised based on established frameworks, such as PESTEL and design thinking. The key data points 134 may further be enhanced using calculated values such as confidence scores and other metrics that quantify their relevance and impact, for example. In some embodiments, related key data points contained in key data points 134 are combined to enhance context and provide more insightful information.

[0096]

[0079] Referring to Figure 4, there is shown a process flow diagram of a method 400 of validating one or more generated entity -specific propositions, according to some embodiments. In some embodiments, method 400 may be performed immediately following generation of the entity -specific propositions 138 in method 300. That is, on completion of method 300, processing circuitry 102 executes instructions in memory 104 to perform method 400, for example. In some embodiments, method 400 may be performed independently of method 300. That is, processing circuitry 102 may execute instructions in memory 104 to perform method 400 on previously generated propositions 138, for example.

[0097]

[0080] At step 402, processing circuitry 102 executes instructions, stored in memory 104, to provide instructions to the external LLM 220 to generate multiple agent constructs. In some embodiments, the processing circuitry 102 provides the instructions to the LLM module 108. In some embodiments, the external LLM 220 and / or the LLM module 108 are instructed to generate tens, hundreds, or thousands of agent constructs.

[0098]

[0081] An agent construct is an artificial element within an LLM configured to simulate a human being, for example. Each of the multiple agent constructs are configured, by the respective LLM, with different modelled behaviour parameters. The different modelled behaviour parameters may include, but are not limited to: age, gender, ethnicity, education, income, employment status, occupation, occupation role, job goals, job challenges, preferred communication methods, work interests, relationship status, personality, and location. In some embodiments, the different modelled behaviour parameters are determined based, at least in part, on one or more of: the seed dataset 122, the first dataset 126, and the second dataset 130.

[0099]

[0082] In some embodiments, the external LLM 220 or the LLM module 108 for generating the multiple agent constructs is trained on one or more of: public text corpora (e.g., books, articles, encyclopedias, and general web content for broad language comprehension), conversational datasets (e.g., dialogue datasets relating to human-like interactions), demographic data (e.g., public datasets with demographic details such as census data and industry reports), behavioural data (e.g., consumer behaviour patterns, trends, and psychographics), and domain specific data (e.g., industry specific datasets).

[0100]

[0083] In some embodiments, the external LLM 220 or the LLM module 108 for generating the multiple agent constructs is further trained on market experiment data, such as engagement metrics, preference trends, and performance scores. The market experiment data may enhance the ability of the agent constructs to simulate realistic customer reactions, for example. In some embodiments, the external LLM 220 or the LLM module 108 for generating the multiple agent constructs is further trained on behavioural and demographic parameters, such as purchase patterns, content engagement, feature requests, feedback, age, gender, income, education, personality, and location, for example.

[0101]

[0084] In some embodiments, the multiple agent constructs are generated based on one or more of: the seed dataset 122, the first dataset 126, and the second dataset 130. That is, the multiple agents may be generated based on the specific entity for which propositions have been generated, for example. Utilising one or more of the seed dataset 122, the first dataset 126, and the second dataset 130 may enable the respective LLM, external LLM 220 and / or LLM module 108, to generate multiple agent constructs to simulate the characteristics and behaviours of users of the services / goods of the specific entity, for example.

[0085] To perform step 402, processing circuitry 102 may execute agent construct generation module 140 stored in memory 104. Executing the agent construct generation module 140 may cause server 100 to determine different modelled behaviour parameters relevant to the entity for which propositions have been generated and provide instructions including the determined different modelled behaviour parameters to the respective LLM.

[0102]

[0086] At step 404, processing circuitry 102 executes instructions, stored in memory 104, to structure a validation query to the external LLM 220 based on the one or more entity-specific generated propositions 138. Processing circuitry 102 may structure the validation query in relation to the specific entity for which the propositions were generated using method 300. That is, the validation query may be structured differently depending on the specific entity for which the generated propositions are being validated. In some embodiments, external reporting documentation is provided with the validation query to the respective LLM, external LLM 220 and / or LLM module 108, to further enhance the quality of the LLM-generated responses. For example, external industry reports or trend descriptions, such as Gartner Special Reports or IBISWorld reports, are provided to the respective LLM, thereby improving the relevance and accuracy of the LLM-generated responses.

[0103]

[0087] Server 100, upon structuring the validation query, communicates, or provides, the query to the external LLM 220. In some embodiments, the processing circuitry 102 provides the structured query to the LLM module 108. In response to receiving the validation query, the external LLM 220 provides an LLM-generated proposition validation response, being or including proposition validation data, to server 100. The proposition validation data is based on the interaction of the multiple agent constructs with the one or more entity -specific generated propositions 138. The proposition validation data may be indicative of the validity of each of the one or more entityspecific generated propositions 138. That is, the proposition validation data may indicate generated propositions related to the specific entity that are valid and should be investigated further, for example. The proposition validation data may indicate the relevance of each of the one or more entity -specific generated propositions 138 to the multiple agent constructs, for example. The validation query may be structured to request that the returned proposition validation data includes an engagement score or confidence score to indicate the predicted effectiveness of the or each proposition, for example. The proposition validation data may indicate the potential impact of each of the one or more entity- specific generated propositions 138 on the multiple agent constructs, for example. Server 100, upon receiving the proposition validation data, stores the data in memory 104, and proceeds to step 406.

[0104]

[0088] To perform step 404, processing circuitry 102 may execute validation module 142 stored in memory 104. Executing validation module 142 may cause server 100 to structure a validation query based on the one or more generated propositions, provide the query to an LLM, such as external LLM 220 or LLM module 108, and receive from the respective LLM a proposition validation response, including proposition validation data. The proposition validation response may be based on the interaction of the multiple agent constructs with the one or more entity- specific propositions, for example. In some embodiments, the received LLM proposition validation data is further validated against known outcomes, such as historical engagement data, to further train the external LLM 220 or LLM module 108.

[0105]

[0089] At step 406, processing circuitry 102 executes instructions, stored in memory 104, to generate one or more content packages based on the proposition validation data and respective one or more generated entity -specific propositions 138. Content packages may only be generated for propositions considered valid based on the validation data. That is, content packages may only be generated for propositions with sufficient validation scores or confidence levels, for example. In some embodiments, generation of the content packages is further based on one or more of: the seed dataset 122, the first dataset 126, and the second dataset 130. A content package may provide means of presenting a validated proposition to one or more users of a platform. A content package may include one or more of: text, image, audio, and video.

[0106]

[0090] In some embodiments, the one or more content packages are generated for a target user, which is not a specific unique person but is a specified class of persons. That is, a user, or group of users, may be identified as relevant to the specific entity and / or the generated proposition and the content package may be targeted to them, for example. Generation of the one or more content packages in respect of a target user may take into account various factors including, but not limited to: demographics, industry, and / or location.

[0107]

[0091] In some embodiments, the target user is automatically determined based on a data analysis of the seed dataset 122, the first dataset 126, and the second dataset 130. The data analysis may determine the target user based on demographics, location, and industry, for example. In some embodiments, the target user is further determined based on behavioural insights obtained through engagement metrics and / or interaction data. In some embodiments, the target user is further determined using one or more LLM’s to classify and prioritise user segments based on past engagement and likelihood of conversion or interaction. In some embodiments, the target user is further determined through dynamic matching of target users and personas of agent constructs generated at step 402 and used to validate the entity- specific proposition at step 404. In some embodiments, the target user is manually determined by received user input. A user may specify a particular audience, industry, demographics, and / or location for the target user. A user may further refine the automatically determined target user through manual interaction via a user interface.

[0108]

[0092] A platform may be a content sharing platform. A platform may be a social media platform, for example, such as Linkedln, Instagram, TikTok, or Facebook. A platform may be a website / web page of the specific entity or associated with the specific entity, for example. The content packages may be generated for general use across a number of different platforms. In some embodiments, the one or more content packages are platform specific content packages. That is, they are specifically generated for use on a particular platform. Multiple content packages relating to a single validated proposition may be generated for use on different platforms. For example, three content packages may be generated for a first validated proposition, each of the three content packages being platform specific to different platforms, such as Linkedln, TikTok, Facebook, and Instagram.

[0093] In some embodiments, the one or more content packages are structured according to the desired platform and target user (audience) using predefined templates and / or best practice frameworks. The predefined templates may include layouts optimised for text, images, audio, or video, for example. The best practice frameworks may include messaging frameworks tailored to the validated proposition data’s context, such as benefits driven messaging or problem-solution structures, for example. In some embodiments, the one or more content packages are integrated with additional data, such as performance scores, engagement rates, market trends, customer testimonials, and / or visual assets generated through artificial intelligence.

[0109]

[0094] In some embodiments, server 100 may utilise external LLM 220 and / or LLM module 108 to generate the one or more content packages. That is, external LLM 220 and / or LLM module 108 may be utilised by server 100, at least in part, to generate the content packages, such as the text, images, audio, or video, for example. The external LLM 220 and / or the LLM module 108 may be utilised to determine the target user and the various factors taken into account during generation of the one or more content packages.

[0110]

[0095] An example content package specifically generated for Linkedln (as one example platform) proposing a new workflow feature may include a body of text describing the workflow feature, an artificially generated image including a branded graphic and feature integration steps, and a short-form video demonstrating the workflow feature. In some embodiments, generative Al models are used to generate the one or more content packages.

[0111]

[0096] Processing circuitry 102, upon generating the one or more content packages, executes instructions stored in memory 104 to provide the one or more generated content packages to the respective large language model by which the multiple agent constructs were generated. Processing circuitry 102, then structures and provides a query to the respective large language model. In response to receiving the query, the respective large language model provides an LLM-generated response, being content package validation data, to server 100.

[0097] The content package validation data is based on the interaction of the multiple agent constructs with the one or more generated content packages. The received content package validation data may be indicative of the validity of each of the one or more generated content packages. That is, the content package validation data may indicate content packages that are valid and recommended to be published and / or trialled on a platform(s), for example. Server 100, upon receiving the content package validation data, stores the data in memory 104 and proceeds to step 408.

[0112]

[0098] Interaction of the multiple agent constructs may include simulated feedback, wherein the agent constructs are programmed to simulate user behaviour, such as clicking, viewing, or reacting to content (e.g., positive feedback, ignoring, or disengaging), for example. The interactions of the multiple agents may be tracked and scored based on engagement metrics (e.g., click-through rates and time spend on content) and sentiment analysis (e.g., positive / negative responses to content messaging).

[0113]

[0099] In some embodiments, the content packages are tested with the multiple agent constructs in various scenarios that mimic real-world conditions, such as mobile browsing and desktop ads, for example. The validation data may be generated based on the scored interactions of the multiple agents with the one or more content packages. The scored interactions may be aggregated across multiple agents to reduce bias, for example.

[0114]

[0100] In some embodiments, validation data obtained through interaction of the multiple agent constructs with the one or more content packages is used to iteratively refine the one or more content packages. That is, the one or more content packages may be validated and subsequently automatically regenerated based on the validation data before being re-validated until a satisfactory validity is obtained, for example.

[0115]

[0101] At step 408, processing circuitry 102, upon validating the one or more generated content packages related to the validated propositions, executes instructions stored in memory 104 to publish the one or more validated content packages to their respective platform(s). That is, the validated content packages are published to the platform(s) on which they were generated for. For example, if a first content package is generated for Linkedln and Instagram and a second content package is generated for Facebook, the first content package will only be published to Linkedln and Instagram and the second content package will only be published to Facebook. To publish the one or more validated content packages to a platform, server 100 may interface with the respective platform via a URL or an API. For example, processing circuitry 102, executing instructions stored in memory 104, may configure server 100 to communicate using communications module 106 with a platform API via network 200. Processing circuitry 102 may configure server 100 to send a request, such as an HTTP request, to the platform API to publish the one or more validated content packages.

[0116]

[0102] In some embodiments, server 100 may utilise external LLM 220 and / or LLM module 108 to publish the one or more validated content packages to their respective platforms. That is, external LLM 220 and / or LLM module 108 may be utilised by server 100, at least in part, to interface with the various platforms and publish the validated content packages thereon. In some embodiments, the external LLM 220 and / or the LLM module 108 may automatically publish the one or more content packages once they are validated. In some embodiments, the one or more generated content packages are published by the external LLM 220 and / or the LLM module 108 without being validated by the multiple agent constructs.

[0117]

[0103] To perform steps 406 and 408, processing circuitry 102 may execute content generation, validation, and publishing module 144 stored in memory 104. Executing content generation, validation, and publishing module 144 may cause server 100 to generate one or more content packages based on the entity -specific generated propositions and the corresponding validation data. The content packages are intended for a target audience relevant to the specific entity. The content generation, validation, and publishing module 144 may validate the one or more generated content packages using the multiple agent constructs, and publish the one or more validated content packages on their respective designated platforms.

[0104] In some embodiments, the content generation, validation, and publishing module 144 comprises three separate modules stored in memory 104, a content generation module, a content validation module, and a content publishing module. Executing the content generation module may cause server 100 to generate one or more content packages based on the entity -specific generated propositions. Executing content validation module may cause server 100 provides the generated one or more content packages to the respective large language model by which the multiple agent constructs were generated to generate validation data and validate the one or more content packages. Executing the content publishing module may cause server 100 to publish the one or more validated content packages to their respective platform(s).

[0118]

[0105] At step 410, upon publishing the one or more content packages on their respective platforms, processing circuitry 102 executes instructions stored on memory 104 to retrieve performance metrics related to each content package from each of the platforms to which the one or more content packages were published. Performance metrics may include one or more of, but not limited to: user interactions, impressions, video or image views, listens, clicks, click through rate, likes, comments, shares, follows, conversion rates, cost per acquisition, return on spend, and engagement rate.

[0119]

[0106] Server 100 may retrieve the performance metrics by accessing the respective platform(s) using a URL or an API 222 via network 200. For example, processing circuitry 102, executing instructions stored in memory 104, may configure server 100 to communicate using communications module 106 with a platform API 222 via network 200. Processing circuitry 102 may configure server 100 to send a request, such as an HTTP request, to the platform API 222 requesting performance metrics and in response the platform provides a response message containing the performance metrics. The performance metrics may be provided by the platform in a JSON or XML format, for example.

[0120]

[0107] Upon retrieving the performance metrics for each content package, processing circuitry 102 executes instructions stored in memory 104 to perform an analysis on the performance metrics in respect of each of the content packages. In some embodiments, the analysis includes performing advertising analysis and / or statistical analysis on the performance metrics. In some embodiments, the retrieved performance metrics are analysed in reference to minimum benchmarks to determine the viability of an entity - specific generated proposition.

[0121]

[0108] In performing the analysis on the retrieved performance metrics, processing circuitry 102, executing instructions stored in memory 104, generates performance results for each of the generated propositions for which one or more content packages were generated and published. The performance results may include a content package score by which each of the one or more content packages are graded based on the performance metrics. The performance results may include a proposition score by which the generated propositions are graded based on the performance of their respective one or more content packages. That is, a generated proposition having two respective content packages may be graded based on the individual content package scores of the two respective content packages, for example. The performance results may include a recommendation for language choice. The performance results may include a recommendation for targeting a different user base. The performance results may include a summary of the performance of the content package.

[0122]

[0109] To perform step 410, processing circuitry 102 may execute content package performance analysis module 146 stored in memory 104. Executing content package performance analysis module 146 may cause server 100 to retrieve the performance metrics from the relevant platform(s), perform an analysis on the retrieved performance metrics, and generate performance results in relation to the entity -specific generated propositions and their respective generated content packages. In some embodiments, the analysis on the retrieved performance metrics by performance analysis module 146 is performed, at least in part, using Al driven analysis to identify patterns and anomalies in the performance metrics.

[0123]

[0110] In some embodiments, retrieval of the performance metrics from the relevant platform(s) involves passing the retrieved metrics through a standardised API integration framework, such as by server 100 executing API integration module 152. Execution of API integration module 152 may cause server 100 to normalise and format the retrieved performance metrics from the platform(s) and transform it into a consistent and cohesive structure. Execution of API integration module 152 may cause server 100 to map inconsistent metrics retrieved from different frameworks to a standardised metric. That is, where one platform records a particular metric differently to another platform, the API integration module standardises these metrics such that they are comparable, for example.

[0124]

[0111] Referring to Figure 5, there is shown a process flow diagram of a method 500 of generating entity- specific propositions using artificial intelligence including retrieval-augmentation generation (RAG), according to some embodiments. Steps 502, 504, 506, 508, and 510 of method 500 correspond to steps 302, 304, 306, 308, and 310 of method 300. That is, performing step 502 of method 500 is equivalent to performing step 302 of method 300 as previously described, for example. Similarly, performing step 510 of method 500 is equivalent to performing step 310 of method 300 as previously described, for example.

[0125]

[0112] Retrieval-augmented generation is an Al framework for improving the quality of LLM-generated responses by grounding the model on external sources of knowledge to supplement the LLM’s internal representation of information. For example, by supplying one or more of the seed dataset 122, the first dataset 126, and the second dataset 130 to the respective LLM, external LLM 220 and / or LLM module 108, the resulting key data points and generated entity- specific propositions may have a higher level of accuracy and / or relevance to the specific entity.

[0126]

[0113] Method 500 further comprises step 507 for applying RAG to the respective LLM, external LLM 220 and / or LLM module 108. Step 507 is performed by processing circuitry 102, executing instructions stored in memory 104, after step 506 and prior to step 508. At step 507, upon generating the second dataset at step 506, processing circuitry 102 executes instructions stored on memory 104 to apply RAG to the respective LLM, external LLM 220 and / or LLM module 108. In applying RAG to the respective LLM, entity -specific data contained in the seed dataset 122, the first dataset 126, and / or the second dataset 130 are provided to the respective LLM to provide the LLM with context of the specific entity and enhance its ability to provide relevant and greater quality responses to queries.

[0127]

[0114] In some embodiments, processing circuitry 102, executing instructions stored in memory 104, transforms one or more of the seed dataset 122, the first dataset 126, and the second dataset 130 into a format ingestible by the respective LLM. Upon applying RAG to the respective LLM, external LLM 220 and / or LLM module 108, method 500 proceeds to step 508, as previously described in relation to Figure 3. In some embodiments, applying RAG to the respective LLM includes refining one or more of the seed dataset 122, the first dataset 126, and the second dataset 130 prior to being provided to the respective LLM. Refining of the datasets may include removing data considered irrelevant to the specific entity’s goals or propositions, for example. The determination of data relevance may be based on entity -specific data (e.g. generated from web-scraping, etc) stored in data store 105 or memory 104, for example.

[0128]

[0115] To perform step 507, processing circuitry 102 may execute retrievalaugmentation generation module 148 stored in memory 104. Executing retrievalaugmentation generation module 148 may cause server 100 to, if required, transform entity-specific data, such as the seed dataset 122, the first dataset 126, and the second dataset 130, into a format ingestible to the respective LLM and provide the transformed entity-specific data to the LLM.

[0129]

[0116] Referring to Figure 6, there is shown a process flow diagram of a method 600 of generating entity- specific propositions using artificial intelligence including framework application, according to some embodiments. At step 602, processing circuitry 102, executing instructions stored in memory 104, provide one or more frameworks to external LLM 220 and / or LLM module 108. The one or more frameworks may influence the respective LLM’s LLM-generated response in relation to the key data points 134 and or the generated one or more propositions 138. In some embodiments, the one or more frameworks include, but are not limited to: PESTEL, double diamond, and design thinking.

[0130]

[0117] To perform step 602, processing circuitry 102 may execute framework integration module 150 stored in memory 104. Executing framework integration module 150 may cause server 100 to provide the one or more frameworks to an LLM, such as external LLM 220 or LLM module 108. In some embodiments, external LLM 220 and / or LLM module 108 are already trained to provide outputs in relation to the one or more frameworks. In some embodiments, semantic matching algorithms are used by framework integration module 150 to map data points to specific framework dimensions based on one or more of: keywords, context, and relationships.

[0131]

[0118] At step 604, upon providing the one or more frameworks to the external LLM 220 and / or the LLM module 108, processing circuitry 102, executing instructions stored in memory 104, structures a framework-based entity- specific query to the respective LLM. The framework-based entity specific query may request the respective LLM to provide an output in relation to a particular framework, or portion of a framework, such as the ‘technological’ portion of PESTEL, for example. The framework-based entity specific query may specify a specific context to which the output of the respective LLM should focus on, for example. An example frameworkbased entity-specific query may be ‘analyse eco-friendly retail trends for urban millennials under 'Technological' and 'Environmental' dimensions. Provide three actionable recommendations with confidence scores’, for example.

[0132]

[0119] The framework-based entity specific query is based on the first dataset 126, the second dataset 130, and the one or more frameworks. In some embodiments, the framework-based entity -specific query may specify a framework of the one or more frameworks for the respective LLM to generate the key data points 134 based on. Processing circuitry 102 may structure the framework-based query in relation to the specific entity for which propositions are being generated using method 600. That is, the query may be structured differently depending on the specific entity for which the method 600 is being performed.

[0120] Server 100, upon structuring the framework-based entity-specific query, communicates, or provides, the query to the one or more external LLMs 220 and / or LLM module 108. In response to receiving the framework-based entity-specific query, the respective LLM, external LLM 220 and / or LLM module 108, provides an LLM- generated response, being the one or more key data points 134 in relation to the specific entity, as previously described, and further in view of the one or more provided frameworks. The one or more key data points 134 are based, at least in part, on the first dataset 126, the second dataset 130, the structured query, and the one or more frameworks. Upon receiving the LLM-generated response, processing circuitry 102 proceeds to step 606.

[0133]

[0121] To perform steps 604, processing circuitry 102 may execute query structure module 132 stored in memory 104. Executing the query structure module 132 may cause server 100 to structure a query specific to the entity for which propositions are to be generated in view of the one or more frameworks provided to an LLM, such as external LLM 220 or LLM module 108, provide the query to an LLM, and receive from the LLM key data points 134 relating to the specific entity in view of the one or more frameworks.

[0134]

[0122] At step 606, processing circuitry 102 executes instructions, stored in memory 104, to generate one or more entity -specific propositions 138 based on the key data points 134 received at step 604, as previously described in relation to step 310 of method 300, and further based on the one or more frameworks provided to the respective LLM. In some embodiments, the key data points 134 generated at step 604 are not generated based on the one or more frameworks, and the entity- specific propositions 138 generated at step 606 are generated based on the one or more frameworks. That is, the key data points 134 are generated as described in relation to step 308 of method 300 and the entity -specific propositions 138 are generated as described in relation to step 606 of method 600. In some embodiments, both the key data points 134 generated at step 604 and the entity -specific propositions 138 generated at step 606 are generated based on the one or more frameworks.

[0123] To perform step 606, processing circuitry 102 may execute proposition generation module 136 stored in memory 104, as previously described. Executing the proposition generation module 136 may cause server 100 to structure a proposition query based on the one or more frameworks, provide the proposition query to the external LLM 220 and / or the LLM module 108, and in response to receiving the proposition query, the respective LLM provides an LLM-generated response, being one or more propositions 138 in relation to the specific entity and the one or more frameworks. In some embodiments, the proposition query is structured based on the one or more frameworks provided to the respective LLM. In some embodiments, the generated entity- specific frameworked propositions are validated and evaluated using the process described in relation to method 400.

[0135]

[0124] Figure 7 illustrates an example computer system 700 according to some embodiments. In particular embodiments, one or more computer systems 700 perform one or more steps of one or more methods described or illustrated herein. In particular embodiments, one or more computer systems 700 provide functionality described or illustrated herein. In particular embodiments, software running on one or more computer systems 700 performs one or more steps of one or more methods described or illustrated herein or provides functionality described or illustrated herein. Particular embodiments include one or more portions of one or more computer systems 700. Herein, reference to a computer system may encompass a computing device, and vice versa, where appropriate. Moreover, reference to a computer system may encompass one or more computer systems, where appropriate. The proposition generation server 100, external LLM 220, external web pages 215, client device(s) 211, client server 210, data store 105, network 200, communications module 106, LLM module 108, seed data retrieval module 120, web scrape module 124, web crawl module 128, query structure module 132, proposition generation module 136, agent construct generation module 140, validation module 142, content generation, validation, and publishing module 144, content package performance analysis module 146, retrieval-augmentation generation module 148, framework integration module 150 and / or API integration module 152 may incorporate a subset or all of the computing components described with reference to the computer system 700 to provide the functionality described in this specification.

[0125] This disclosure contemplates any suitable number of computer systems 700 to implement each of the proposition generation server 100, external LLM 220, external web pages 215, client device(s) 211, client server 210, data store 105, network 200, communications module 106, LLM module 108, seed data retrieval module 120, web scrape module 124, web crawl module 128, query structure module 132, proposition generation module 136, agent construct generation module 140, validation module 142, content generation, validation, and publishing module 144, content package performance analysis module 146, retrieval- augmentation generation module 148, framework integration module 150 and / or API integration module 152.

[0136]

[0126] Computer system 700 may be an embedded computer system, a system-on- chip (SOC), a single-board computer system (SBC) (such as, for example, a computer- on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a mobile telephone, a personal digital assistant (PDA), a server, a tablet computer system, or a combination of two or more of these. Where appropriate, computer system 700 may include one or more computer systems 700; be unitary or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud, which may include one or more cloud components in one or more networks. Where appropriate, one or more computer systems 700 may perform without substantial spatial or temporal limitation one or more steps of one or more methods described or illustrated herein. As an example and not by way of limitation, one or more computer systems 700 may perform in real-time or in batch mode one or more steps of one or more methods described or illustrated herein. One or more computer systems 700 may perform at different times or at different locations one or more steps of one or more methods described or illustrated herein, where appropriate.

[0137]

[0127] In particular embodiments, computer system 700 includes a processor 702, memory 704, storage 706, an input / output (I / O) interface 708, a communication interface 710, and a bus 712. Although this disclosure describes and illustrates a particular computer system having a particular number of particular components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.

[0138]

[0128] In particular embodiments, processor 702 includes hardware for executing instructions, such as those making up a computer program. Processor 702 may perform the functions of processing circuitry 102 as described herein, for example. As an example and not by way of limitation, to execute instructions, processor 702 may retrieve (or fetch) the instructions from an internal register, an internal cache, memory 704, or storage 706; decode and execute them; and then write one or more results to an internal register, an internal cache, memory 704, or storage 706. In particular embodiments, processor 702 may include one or more internal caches for data, instructions, or addresses. This disclosure contemplates processor 702 including any suitable number of any suitable internal caches, where appropriate. As an example and not by way of limitation, processor 702 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction caches may be copies of instructions in memory 704 or storage 706, and the instruction caches may speed up retrieval of those instructions by processor 702. Data in the data caches may be copies of data in memory 704 or storage 706 for instructions executing at processor 702 to operate on; the results of previous instructions executed at processor 702 for access by subsequent instructions executing at processor 702 or for writing to memory 704 or storage 706; or other suitable data. The data caches may speed up read or write operations by processor 702. The TLBs may speed up virtual-address translation for processor 702. In particular embodiments, processor 702 may include one or more internal registers for data, instructions, or addresses. This disclosure contemplates processor 702 including any suitable number of any suitable internal registers, where appropriate. Where appropriate, processor 702 may include one or more arithmetic logic units (ALUs); be a multi-core processor; or include one or more processors 702. Although this disclosure describes and illustrates a particular processor, this disclosure contemplates any suitable processor.

[0139]

[0129] In particular embodiments, memory 704 includes main memory for storing instructions for processor 702 to execute or data for processor 702 to operate on. As an example and not by way of limitation, computer system 700 may load instructions from storage 706 or another source (such as, for example, another computer system 700) to memory 704. Processor 702 may then load the instructions from memory 704 to an internal register or internal cache. To execute the instructions, processor 702 may retrieve the instructions from the internal register or internal cache and decode them. During or after execution of the instructions, processor 702 may write one or more results (which may be intermediate or final results) to the internal register or internal cache. Processor 702 may then write one or more of those results to memory 704. In particular embodiments, processor 702 executes only instructions in one or more internal registers or internal caches or in memory 704 (as opposed to storage 706 or elsewhere) and operates only on data in one or more internal registers or internal caches or in memory 704 (as opposed to storage 706 or elsewhere). One or more memory buses (which may each include an address bus and a data bus) may couple processor 702 to memory 704. Bus 712 may include one or more memory buses, as described below. In particular embodiments, one or more memory management units (MMUs) reside between processor 702 and memory 704 and facilitate accesses to memory 704 requested by processor 702. In particular embodiments, memory 704 includes random access memory (RAM). This RAM may be volatile memory, where appropriate. Where appropriate, this RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Moreover, where appropriate, this RAM may be single-ported or multi-ported RAM. This disclosure contemplates any suitable RAM. Memory 704 may include one or more memories 704, where appropriate. Although this disclosure describes and illustrates particular memory, this disclosure contemplates any suitable memory.

[0140]

[0130] In particular embodiments, storage 706 includes mass storage for data or instructions. As an example and not by way of limitation, storage 706 may include a hard disk drive (HDD), a floppy disk drive, flash memory, an optical disc, a magnetooptical disc, magnetic tape, or a Universal Serial Bus (USB) drive or a combination of two or more of these. Storage 706 may include removable or non-removable (or fixed) media, where appropriate. Storage 706 may be internal or external to computer system 700, where appropriate. In particular embodiments, storage 706 is non-volatile, solid- state memory. In particular embodiments, storage 706 includes read-only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically alterable ROM (EAROM), or flash memory or a combination of two or more of these. This disclosure contemplates mass storage 706 taking any suitable physical form. Storage 706 may include one or more storage control units facilitating communication between processor 702 and storage 706, where appropriate. Where appropriate, storage 706 may include one or more storages 706. Although this disclosure describes and illustrates particular storage, this disclosure contemplates any suitable storage.

[0141]

[0131] In particular embodiments, VO interface 708 includes hardware, software, or both, providing one or more interfaces for communication between computer system 700 and one or more VO devices. Computer system 700 may include one or more of these VO devices, where appropriate. One or more of these VO devices may enable communication between a person and computer system 700. As an example and not by way of limitation, an VO device may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet, touch screen, trackball, video camera, another suitable VO device or a combination of two or more of these. An VO device may include one or more sensors. This disclosure contemplates any suitable VO devices and any suitable VO interfaces 708 for them. Where appropriate, VO interface 708 may include one or more device or software drivers enabling processor 702 to drive one or more of these VO devices. VO interface 708 may include one or more VO interfaces 708, where appropriate. Although this disclosure describes and illustrates a particular VO interface, this disclosure contemplates any suitable VO interface.

[0142]

[0132] In particular embodiments, communication interface 710 includes hardware, software, or both providing one or more interfaces for communication (such as, for example, packet-based communication) between computer system 700 and one or more other computer systems 700 or one or more networks. As an example and not by way of limitation, communication interface 710 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wire-based network or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network, such as a WI-FI network. This disclosure contemplates any suitable network and any suitable communication interface 710 for it. As an example and not by way of limitation, computer system 700 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 700 may communicate with a wireless PAN (WPAN) (such as, for example, a BLUETOOTH WPAN), a WI-FI network, a WLMAX network, a cellular telephone network (such as, for example, a Global System for Mobile Communications (GSM) network), or other suitable wireless network or a combination of two or more of these. Computer system 700 may include any suitable communication interface 710 for any of these networks, where appropriate. Communication interface 710 may include one or more communication interfaces 710, where appropriate. Although this disclosure describes and illustrates a particular communication interface, this disclosure contemplates any suitable communication interface.

[0143]

[0133] In particular embodiments, bus 712 includes hardware, software, or both coupling components of computer system 700 to each other. As an example and not by way of limitation, bus 712 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a serial advanced technology attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus or a combination of two or more of these. Bus 712 may include one or more buses 712, where appropriate. Although this disclosure describes and illustrates a particular bus, this disclosure contemplates any suitable bus or interconnect.

[0134] Herein, a computer-readable non-transitory storage medium or media may include one or more semiconductor-based or other integrated circuits (ICs) (such, as for example, field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical discs, optical disc drives (ODDs), magneto-optical discs, magneto-optical drives, floppy diskettes, floppy disk drives (FDDs), magnetic tapes, solid-state drives (SSDs), RAM-drives, SECURE DIGITAL cards or drives, any other suitable computer-readable non- transitory storage media, or any suitable combination of two or more of these, where appropriate. A computer-readable non-transitory storage medium may be volatile, nonvolatile, or a combination of volatile and non-volatile, where appropriate.

[0144]

[0135] It will be appreciated by persons skilled in the art that numerous variations and / or modifications may be made to the above-described embodiments, without departing from the broad general scope of the present disclosure. The present embodiments are, therefore, to be considered in all respects as illustrative and not restrictive.

Claims

CLAIMS:

1. An automated computer- implemented method for entity-specific proposition generation using artificial intelligence, the method comprising: receiving a seed dataset from one or more entity- specific devices and extracting one or more seed URLs from the seed dataset; performing a web scrape on the one or more seed URLs to generate a first dataset; extracting one or more scraped URLs from the first dataset and performing a web crawl on the one or more scraped URLs to generate a second dataset; structuring an entity -specific query based on the first and second datasets; receiving, by providing the entity- specific query to a first large language model, one or more key data points in relation to the specific entity; and generating one or more entity- specific propositions based on the key data points.

2. The method of claim 1, further comprising: providing instructions to a second large language model to generate multiple agent constructs, wherein the multiple agent constructs are each configured with different modelled behaviour parameters; subsequently, structuring a validation query to the second large language model based on the one or more entity- specific propositions; receiving, by providing the validation query to the second large language model, proposition validation data based on the interaction of the multiple agentconstructs with the one or more entity -specific propositions indicative of the validity of each of the one or more entity -specific propositions.

3. The method of claim 2, wherein the different modelled behaviour parameters include one or more of: age, gender, ethnicity, education, income, employment status, relationship status, personality, and location.

4. The method of claim 2 or claim 3, wherein the different modelled behaviours are determined, in part, on one or more of the seed dataset, the first dataset, and the second dataset.

5. The method of any one of claims 2 to 4, further comprising: generating one or more content packages based on the generated entityspecific propositions and proposition validation data; and publishing the one or more content packages to one or more platforms.

6. The method of claim 5, wherein each of the one or more content packages are platform specific content packages.

7. The method of claim 5 and 6, wherein the content packages are generated for a target user.

8. The method of claim 7, wherein the one or more content packages are generated based on one or more of the target user’s demographics, industry, and location.

9. The method of any one of claims 5 to 8, wherein the one or more content packages include one or more of: text, image, audio, and video.

10. The method of any one of claims 5 to 9, further comprising, prior to publishing the one or more content packages:providing the one or more content packages to the second large language model; subsequently, structuring a query to the second large language model based on the one or more content packages; receiving, by providing the query to the second large language model, content package validation data based on the interaction of the multiple agent constructs with the one or more content packages indicative of the validity of each of the one or more content package; and wherein the publishing includes publishing the one or more content packages to one or more platforms based on the content package validation data.

11. The method of any one of claims 5 to 10, further comprising: retrieving performance metrics of the one or more content packages from the one or more platforms; and generating performance results in relation to the one or more entity -specific propositions based on the retrieved performance metrics.

12. The method of any one of claims 1 to 11, wherein performing the web scrape extracts markup language code from the one or more seed URLs to generate the first dataset.

13. The method of any one of claims 1 to 12, wherein performing the web crawl extracts markup language code from the one or more scraped URLs to generate the second dataset.

14. The method of any one of claims 1 to 13, further comprising, prior to structuring the entity- specific query:applying retrieval-augmentation generation by providing one or more of the seed dataset, the first dataset, and the second dataset to the first large language model.

15. The method of any one of claims 1 to 14, further comprising, prior to structuring the entity- specific query: providing one or more frameworks to the first large language model; and wherein structuring the entity -specific query is further based on the provided one or more framework.

16. The method of claim 15, wherein the one or more frameworks comprise one or more of: double diamond, PESTEL, and design thinking.

17. A system for automated entity -specific proposition generation using artificial intelligence, the system comprising a server, the server comprising: processing circuitry; memory accessible to the processing circuitry; a communications module; wherein the memory stores instructions that, when executed by the processing circuitry, cause the processing circuitry to perform the method of any one of claims 1 to 16.

18. The system of claim 17, further comprising a large language model module, wherein the entity -specific query is provided to the large language model module.

19. The system of claim 18, wherein instructions to generate multiple agent constructs are provided to the large language model module, and wherein the validation query is provided to the large language model module.

20. The system of any one of claims 17 to 19, further comprising one or more entity-specific devices, wherein the one or more entity- specific devices are in communication with the server via the communications module.

21. A non-transitory computer readable medium having one or more computer program instructions configured to perform a method, the method comprising: receiving a seed dataset from one or more entity- specific devices and extracting one or more seed URLs from the seed dataset; performing a web scrape on the one or more seed URLs to generate a first dataset; extracting one or more scraped URLs from the first dataset and performing a web crawl on the one or more scraped URLs to generate a second dataset; structuring an entity -specific query based on the first and second datasets; receiving, by providing the entity- specific query to a first large language model, one or more key data points in relation to the specific entity; and generating one or more entity- specific propositions based on the key data points.

Citation Information

Patent Citations

  • System and method for managing artificial conversational entities enhanced by social knowledge

    US10599644B2

  • Omnichannel data communications system using artificial intelligence (AI) based machine learning and predictive analysis

    US11017176B2

  • Multi-service business platform system having entity resolution systems and methods

    US11775494B2

  • Methods and systems for a content development and management platform

    US11836199B2