An intelligent multimodal knowledge synthesizer system and associated methods
The intelligent multimodal knowledge synthesizer system addresses the limitations of existing AI systems by integrating diverse data sources and employing advanced AI techniques to deliver user-specific, hyper-personalized product advisories, enhancing user engagement and sales growth.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NAGARRO SOFTWARE PTE LTD
- Filing Date
- 2026-01-20
- Publication Date
- 2026-07-30
AI Technical Summary
Existing artificial intelligence systems struggle to provide cohesive, personalized user experiences due to limitations in data integration, video analysis, and user interaction capabilities, often relying on static knowledge bases and failing to analyze data from multiple modalities, leading to generic responses that are not tailored to individual users, and are difficult to scale across different platforms.
An intelligent multimodal knowledge synthesizer system that integrates diverse data sources into a unified knowledge graph, employing advanced AI techniques like natural language processing and multi-modal content analysis to provide user-specific, hyper-personalized product and service advisories, using a multimodal content parser, knowledge mixer, knowledge graph builder, and Al playbook agent to generate personalized advice.
The system effectively provides personalized product and service advisories to multiple users simultaneously, improving user engagement and sales growth by leveraging diverse data sources and advanced AI techniques to deliver tailored responses across various platforms.
Smart Images

Figure IN2026050098_30072026_PF_FP_ABST
Abstract
Description
[0001] AN INTELLIGENT MULTIMODAL KNOWLEDGE SYNTHESIZER SYSTEM AND ASSOCIATED METHODS
[0002] TECHNICAL FIELD OF INVENTION
[0003] The present invention generally relates to an artificial intelligence system for natural language communication. More particularly, the present invention relates to a multimodal knowledge synthesizer system for producing an automated userspecific hyper-personalized product and service advisory for a user.
[0004] BACKGROUND OF THE INVENTION
[0005] With the advent of digital technologies, sales of products and services have increasingly relied on digital interfaces, be it online ecommerce websites, or in-store kiosks. Over time, the digital sales and user engagement landscape is increasingly becoming complex, with users interacting with brands across multiple channels, from eCommerce websites to social media and in-store kiosks.
[0006] Additionally, with the advent of artificial intelligence (Al) systems, a lot of sales support, marketing, and branding is being handled by artificial intelligence driven sales and engagement tools. However, prior art artificial intelligence driven sales and engagement tools struggle to deliver a cohesive, personalized experience. This is due to the limitations in data integration, video analysis, and user interaction capabilities. These Al tools often rely on static knowledge bases, leading to generic and uninspired user interactions. Additionally, these tools struggle with analyzing data from multiple modalities, such as, text, image, video, audio, loT, sensors, and are often restricted to basic features, such as, keyword extraction or simple transcription. The existing tools lack the capability to extract meaningful insights from videos without sound or across different languages, limiting their applicability in global and industry contexts. Therefore, these existing tools often provide generic responses that are not tailored to individual users. This lack of personalization reduces the effectiveness of userengagement and requires additional cost and effort to make it personalized. Further, existing tools are typically for single-channel interactions, making it difficult to scale across different platforms (e.g., eCommerce, social media, in-store kiosks). Therefore, the existing tools fail to drive significant engagement or sales growth.
[0007] Therefore, there is a continued need for an automated intelligent multimodal knowledge synthesizer that overcomes the limitations of prior methods by integrating diverse data sources into a unified knowledge graph, and employ advanced Al techniques, including natural language processing, Gen-AI, multi-modal content analysis, to provide a holistic understanding of user preferences and product usage to provide a personalized sales experience to the users. Additionally, there is a need for a system that can provide automated user-specific hyper personalized product and service advisory to several hundred or thousand users simultaneously, thereby significantly improving the quality of product and service advisory for all users, for example, customers of an ecommerce platform linked to the system.
[0008] SUMMARY OF THE INVENTION:
[0009] The following presents a simplified summary of the invention in order to provide a basic understanding of some aspects of the invention. This summary is not an extensive overview of the present invention. It is not intended to identify the key / critical elements of the invention or to delineate the scope of the invention. Its sole purpose is to present some concept of the invention in a simplified form as a prelude to a more detailed description of the invention presented later.
[0010] Aspects of the present invention relate to an intelligent multimodal knowledge synthesizer system for producing an automated user-specific, hyper personalized product and service advisory for a user. The system comprises a multimodal content parser module configured to extract and analyze data from different sources of information about a product or the user, a knowledge mixer module configured to mix the knowledge extracted by the multimodal content parser module from said different sources, a knowledge graph builder module configured to build a knowledge graphbased on the knowledge produced by the knowledge mixer. The system further comprises a grounding and fine-tuning module that grounds and fine-tunes a foundational model using one or more knowledge graphs, and an Al playbook agent module that extracts relevant knowledge from the grounded and fine-tuned foundational model in response to a user query to generate personalized advice and recommendations about the product for the user.
[0011] According to some aspects, the system comprises a content crawler that crawls content about the product or the user from online repositories, data sources, forums, brochures, catalogs and social media platforms and inputs the crawled content comprising video, image, audio or text data into the multimodal content parser.
[0012] According to some aspects, the multimodal content parser module comprises an image and video content parser, which includes a content slicer that fragments the input image or video into video fragments of different sizes, at least one sliding window analyzer that analyzes at least one video fragment at a time, wherein the input image is treated as a single video fragment for analysis by the sliding window analyzer and a content knowledge module that stores the knowledge of each video fragment.
[0013] According to some aspects, the multimodal content parser module comprises an audio content parser, which includes a content slicer that fragments the input audio into audio fragments of different sizes, at least one sliding window analyzer that analyzes at least one audio fragment at a time to analyze the audio fragment and a content knowledge module that stores the knowledge of each audio fragment.
[0014] According to some aspects, the multimodal content parser module comprises a text content parser and a content knowledge module that stores the knowledge of the text content.
[0015] According to some aspects the intelligent multimodal knowledge synthesizer system comprises a knowledge mixer module, the module comprises knowledge vectors construction module that constructs knowledge vectors from user data, tensor creation module that arranges each of the knowledge vectors into a multi-dimensional tensor, knowledge vector clustering module that groups the knowledge vectors in therelevant tensor slice into one or more clusters, context vectorization module that encodes a user’s query’s context into a query context vector, cluster selection module, that selects one or more most relevant clusters based on the query context vector, an attention weight determination module that determines attention weights for each of the knowledge vectors in the one or more most relevant clusters.
[0016] According to some aspects, the attention weights determination module, determines the attention weights by computing standard scaled dot product between the query context vector and each of the knowledge vectors of the one or more most relevant clusters, which are further normalized using softmax function to generate final weights.
[0017] According to some aspects, the cluster selection module, slices the multidimensional tensor to determine one or more relevant tensor slices based on the dimensions determined from the user’s profile or the user’s query.
[0018] According to some aspects, the aggregated knowledge vector module, calculates an aggregated knowledge vector by summing all the weighted knowledge vectors related to an entity.
[0019] According to some aspects, the knowledge vector clustering module selects a clustering algorithm, computes similarity among different knowledge vectors, applies a clustering algorithm to partition knowledge vectors into one or more clusters and characterizes and labels the one or more clusters.
[0020] According to some aspects, the knowledge graph builder module creates a detailed relational knowledge graph using the information received from the knowledge mixer module.
[0021] According to some aspects, the grounding and fine-tuning foundational model module grounds the foundational model using the data of the knowledge graph and first party files from the embedded vector database, and industry playbook templates from a reinforcement, and feedback and training application, wherein industry templates are text configuration files, which are used to refine the system context, and help in effective response formulation or creating / executing a relevant action plan.Some other aspects of the present invention relate to a method for producing an automated user-specific hyper personalized product advisory for a user. The method comprises extracting and analyzing data from different sources of information about a product or the user by the multimodal content parser module, mixing the knowledge extracted by the multimodal content parser module by the knowledge mixer module, building a knowledge graph based on the knowledge produced by the knowledge mixer, by the knowledge graph builder module; grounding and fine-tuning the foundational model using one or more knowledge graphs, by the grounding and fine-tuning module; extracting relevant knowledge from the grounded and fine-tuned foundational model in response to a user query and generating personalized response for the user, by the Al playbook agent.
[0022] According to some aspects, the method of intelligent multimodal knowledge synthesizing further comprises crawling content about the product or the user from online repositories, data sources, forums, brochures, catalogs and social media platforms and input the crawled content comprising video, image, audio or text data into the multimodal content parser.
[0023] According to some aspects, the extracting and analyzing comprises, parsing an image and video content, which includes slicing an input image or video into video fragments of different sizes, analyzing at least one video fragment at a time, wherein the input image is treated as a single video fragment for the analysis and storing the knowledge of each video fragment.
[0024] According to some aspects, the extracting and analyzing comprises parsing an audio content, which includes slicing the input audio into audio fragments of different sizes, analyzing at least one audio fragment at a time and storing the knowledge of each audio fragment.
[0025] According to some aspects, the extracting and analyzing comprises parsing a text content and storing the knowledge of the text content.
[0026] According to some aspects, method of knowledge mixing comprises constructing knowledge vectors from user data, arranging each of the knowledge vectors into amulti-dimensional tensor, clustering the knowledge vectors in the relevant tensor slice into one or more clusters, encoding a user’s query’s context into a query context vector, selecting one or more most relevant clusters based on the query context vector and determining attention weights for each of the knowledge vectors in the one or more most relevant clusters.
[0027] According to some aspects, determining of the attention weights comprises computing standard scaled dot product between the query context vector and each of the knowledge vectors of the one or more most relevant clusters and normalizing using SoftMax function to generate final weights.
[0028] According to some aspects, the method further comprises slicing the multidimensional tensor to determine one or more relevant tensor slices based on the dimensions determined from the user’s profile or the user’s query.
[0029] According to some aspects, the method further comprises calculating the aggregated knowledge vector by combining all the weighted knowledge vectors related to an entity.
[0030] According to some aspects, the clustering of knowledge vectors comprises selecting a clustering algorithm, computing similarity among different knowledge vectors, applying a clustering algorithm to partition knowledge vectors into one or more clusters, characterizing and labeling the one or more clusters.
[0031] According to some aspects, building the knowledge graph creates a detailed relational knowledge graph using the information received from the knowledge mixer module.
[0032] According to some aspects, the grounding and fine-tuning comprises using the data of the knowledge graph and first party files from the embedded vector database, and industry playbook templates from a reinforcement, and feedback and training application, wherein industry templates are text configuration files, which are used to refine the system context, and help in effective response formulation or creating / executing a relevant action plan.Other aspects, advantages, and salient features of the invention will become apparent to those skilled in the art from the following detailed description, which, taken in conjunction with the annexed drawings, discloses exemplary embodiments of the invention.
[0033] BRIEF DESCRIPTION OF ACCOMPANYING DRAWINGS:
[0034] Some of the objects of the invention have been set forth above. These and other objects, features, aspects and advantages of the present invention will become better understood with regard to the following description, appended claims and accompanying drawings where:
[0035] FIG. 1 is a schematic diagram representing various components of an intelligent multimodal knowledge synthesizer system for producing a user-specific hyper personalized product and service advisory;
[0036] FIG. 2 is a flow diagram representing the method steps followed by intelligent multimodal knowledge synthesizer system of FIG. 1 for producing a user-specific hyper personalized product and service advisory;
[0037] FIG. 3 is a schematic diagram representing a video content parser;
[0038] FIG. 4 is a flow diagram representing the method steps followed by the video content parser of FIG. 3;
[0039] FIG. 5 is a schematic diagram representing a video crawler for the video content parser;
[0040] FIG. 6 is a flow diagram representing the method steps followed by the video crawler of FIG. 5;
[0041] FIG. 7 is a schematic diagram representing an audio content parser 700;FIG. 8 is a flow diagram representing a method 800 followed by the audio content parser shown in FIG. 7;
[0042] FIG. 9 is a schematic diagram representing an audio crawler for the audio content parser;
[0043] FIG. 10 is a flow diagram representing the method steps followed by the audio crawler of FIG. 9;
[0044] FIG. 11 is a schematic diagram representing a text content parser configured to parse text content;
[0045] FIG. 12 is a flow diagram representing a method followed by the text content parser shown in FIG. 11 ;
[0046] FIG. 13 is a schematic diagram representing a text crawler for the text content parser.
[0047] FIG. 14 discloses a method of operation of the text crawler as shown in FIG. 13;
[0048] FIG. 15 is a schematic diagram representing an loT sensor data content parser configured to parse loT sensor data;
[0049] FIG. 16 is a flow diagram representing a method followed by the loT Sensor data content parser shown in FIG. 15;
[0050] FIG. 17 is a schematic diagram representing the relationship between the multimodal content, the multimodal content parser, the knowledge mixer, and the knowledge graph builder.
[0051] FIG. 17a is a schematic diagram representing knowledge vectors in multidimensional space created in a knowledge vector construction module.FIG. 17b is a schematic diagram representing knowledge vectors into a three-dimensional tensor that is created in a tensor creation module.
[0052] FIG. 17c is a schematic diagram representing a simpler depiction of knowledge vectors in a two-dimensional (DI and D2) tensor (tensor not shown).
[0053] FIG. 17d is a schematic diagram of an exemplary two-dimensional tensor that is created in a tensor creation module.
[0054] FIG 17e is a flow diagram representing the method steps followed by of the knowledge vector creation module and tensor creation module.
[0055] FIG. 17f is a schematic diagram representing clustering of the knowledge vectors into multiple clusters (Cl, C2 and C3) by a knowledge vector clustering module.
[0056] FIG. 17g is a flow diagram representing the method steps followed by knowledge vector clustering module.
[0057] FIG. 17h is a schematic diagram representing encoding of a user’s query into a query context vector by a context vectorization module.
[0058] FIG. 17i illustrates a visual representation of an exemplary user’s query into a query context vector by context vectorization module.
[0059] FIG. 17j illustrates query context vector probing the clusters to find the most relevant one or more clusters by cluster selection module.
[0060] FIG. 17k illustrates the shortlisted most relevant cluster based on the query context vector by cluster selection module.
[0061] FIG. 171 illustrates shortlisted knowledge vectors which are part of the shortlisted most relevant cluster by cluster selection module.FIG. 17m illustrates calculation of the final attention weights for one of the shortlisted knowledge vectors by attention weights determination module.
[0062] FIG. 17n is a flow diagram representing the method steps followed by of the attention weight determination module.
[0063] FIG 17o is a flow diagram representing the method steps followed by of the knowledge mixer module.
[0064] FIG. 17p, represents different modules in a knowledge mixer module.
[0065] FIG. 18 is a schematic diagram representing the functional components of the knowledge graph builder;
[0066] FIGS. 19A-19B are schematic diagrams representing creation of knowledge graphs database. FIG. 19A is a schematic representing creation of product knowledge graphs database in the system, and FIG. 19B is a schematic representing a creation of user knowledge graphs database.
[0067] FIGS. 20A-20B are schematic diagrams representing knowledge graphs. FIG.
[0068] 20A is a schematic representing a user knowledge graph, and FIG. 20B is a schematic representing a knowledge graph on a product in focus.
[0069] FIG. 21 is a schematic diagram representing the grounding and fine-tuning foundational model module.
[0070] FIG. 22 is a schematic diagram representing the interaction between the user device application, the Omni Channel Integration Sub-system, the Al Playbook Agent, and the Grounded and Fine-tuned Foundational Model; and the interaction between Al Playbook agent and the Action Framework and External Service Providers.
[0071] FIG. 23 is a flow diagram representing a method followed by the Al Playbook Agent;FIG. 24 is a schematic diagram representing an omni-channel integration subsystem;
[0072] FIG. 25 is a schematic diagram representing flow of information through the system when the user engages the system with a query in accordance with the present invention;
[0073] FIGS. 26A-26B are schematic diagrams representing conversations between a user and prior-art Al agents, where FIG. 26A (Prior Art) is a schematic diagram representing a chat between a user Jane and a non-personalized Al agent, and FIG. 26B (Prior Art) is a schematic diagram representing a chat between a user Jane and an Al agent with some personalization; and
[0074] FIGS. 27A-27B illustrate a conversation between a user and the Al playbook agent of the present invention, where FIG. 27A is a schematic diagram representing a chat between a user Jane and a highly personalized Al agent; and FIG. 27B is a schematic diagram representing various aspects of personalization in the highly personalized Al agent response of FIG. 27A.
[0075] Corresponding reference numerals indicate corresponding parts throughout the drawings.
[0076] DETAILED DESCRIPTION OF INVENTION:
[0077] The following detailed description should be read with reference to the drawings in which similar elements in different drawings are numbered the same. The drawings, which are not necessarily to scale, depict illustrative embodiments and are not intended to limit the scope of the invention. Although examples of construction, dimensions, and materials are illustrated for the various elements, those skilled in the art will recognize that many of the examples provided have suitable alternatives that may be utilized.
[0078] Definitions and AbbreviationsApplication programming interface (API): API is a connection between computers or between computer programs. It is a type of software interface, offering a service to other pieces of software.
[0079] Large Language Model (LLM): A large language model is a large-scale machine learning model trained on a broad set of natural language data that can be adapted and fine-tuned for a wide variety of applications and downstream tasks. The term LLM is used at several instances in the description text, and it can also refer to the foundational model (defined below), or inputs thereto or output thereof, i.e., LLM input, LLM Output.
[0080] Foundational Model: A foundational model is a large-scale machine learning model trained on a broad data set that can be adapted and fine-tuned for a wide variety of applications and downstream tasks. A foundational model can be an LLM or can include an LLM within itself in addition to other models such as, image, audio, or video processing models.
[0081] Grounding: In generative Al, grounding is the ability to connect model output to verifiable sources of information. When models are provided with access to specific data sources, then grounding tethers their output to these data and reduces the chances of inventing or hallucinating content. This is particularly important in situations where accuracy and reliability are important.
[0082] Fine-tuning: Fine-tuning is the process of taking a pretrained machine learning model and further training it on a smaller, targeted data set. The aim of fine-tuning is to maintain the original capabilities of a pretrained model while adapting it to suit more specialized use cases.
[0083] Industry Playbook Templates: Industry playbook templates are text configuration files, which are used to refine the system context, and help in effective response formulation or creating / executing a relevant action plan.Al Avatar: Al avatar refers to a digital representation or embodiment of an individual that is created and controlled using artificial intelligence techniques. It is an interactive virtual character that can simulate human-like behaviors, emotions, and interactions.
[0084] Knowledge Schema: A knowledge schema is an organization of knowledge in a data structure having nodes and their respective relations compatible with a corresponding knowledge graph.
[0085] Module: A module, as discussed in several instances in the specification below, is a combination of hardware and / or software components that work together to perform various functions of the system as defined and elaborated in the specification.
[0086] Cloud: In the context of the present invention, the term cloud means a collection of one or more servers comprising data storage mediums, such as hard drive, solid state drives, etc., that are located remotely from one another or a data processing system, but are connected to each other and the data processing system via a network, such as the internet or an intranet, that allows for storage and retrieval of data for processing at the data processing system, to and from this collection of remotely located servers.
[0087] Detailed Description
[0088] FIG. 1 illustrates a schematic diagram representing various components of an intelligent multimodal knowledge synthesizer system 100 for producing a user-specific hyper personalized product and service advisory. FIG. 1 depicts the system 100 in a networked environment comprising a user device 10, a network 12 (such as internet or intranet), a computer server 14, and a data storage medium 16, such as a networked data storage or hard disk drive, or solid-state drive linked to the server 14. The system 100 is configured on the server 14 and the data storage medium 16, and is configured to interact with multiple different types of user devices 10, including but not limited to, mobile phones, smart displays, tablets, laptops, desktops, in-store kiosks, etc. The system 100 interacts with the user devices 10 via a user device application 18, such asa window, a chat interface, a map, a user interface (UI) widget, etc. and user data sources 20, such as keyboard, touchscreen, mouse, camera, microphone, loT sensors, such as such as heart rate sensor, camera, ambient light sensor, temperature sensor, microphone, etc. In some embodiments, data from the user data sources 20 can be in the form of typed (on keyboard or touchscreen) or spoken (on microphone) user inputs 20a, such as, typed or spoken prompt strings, e.g. “Is this shampoo good for my type of hair?”, “how to clean makeup?”, “how to enhance the product application and recommended associated beauty styles”, or “Is the product waterproof and what is the feedback on its dryness?”, etc., to the user device application 18 or data from loT sensors 20b, such as measurements from a heart rate sensor, images from a camera, intensity of daylight from an ambient light sensor, measurements from a temperature sensor, sound captured by a microphone, etc.
[0089] In some embodiments, the system 100 can also interact with other sources or repositories of multimodal data 22 on the network 12, such as Youtube®, Instagram®, Spotify®, Soundcloud®, Reddit®, Wikipedia®, websites, blogs, weather information sensors, product catalogs, and industry data.
[0090] In some embodiments, the system 100 can also interact with external service providers 24, such as salons, therapy centers, etc. on the network 12. In some embodiments, the system 100 can interact with said external service providers 24 via application programming interfaces (APIs).
[0091] The system 100 comprises a number of modules configured on the server 14 to allow for functioning of the system 100. These modules, as described below, are a combination of hardware and / or software components that work together to perform various functions of the system 100. In some embodiments, the system 100 includes specialized hardware components, operably connected to the server 14, designed for artificial intelligence and neural network-related tasks, such as graphics processing units (GPUs), field-programmable gate arrays (FPGAs), and application-specific integrated circuits (ASICs), Neural Processing Units (NPUs), etc. In some embodiments, the system 100 includes specialized software components designed torun on aforementioned specialized hardware components for processing artificial intelligence and neural network-related tasks, such as video and image processing programs, speech-to-text conversion programs, language translation programs, text summarization programs, key words and key concepts identification programs, audio signal processing programs, programs to process loT sensors data, programs to build knowledge schemas, programs to build knowledge graphs, etc. In some other embodiments, these modules can be software programs installed directly on the server 14.
[0092] In some embodiments, the system 100 includes a multimodal content parser module 102 configured to analyze multimodal data 22 received from the network 12. In some embodiments, the multimodal data 22 comprises an image or video content 22a, such as a Youtube® video or an Instagram® post or a reel, audio content 22b, such as, an audio podcast on Spotify® or Soundcloud®, text content 22c, such as product reviews, discussions, comments, brochures, manuals, etc., and loT (Internet of Things) sensor data 22d, such as, health data from a smart watch, weather information, etc., from different sources of information about a user or a product, such as, but not limited to, a personal care product, a beauty product, an electronic device, a household appliance, and an automotive component or accessory.
[0093] In some embodiments, the system 100 includes a knowledge mixer module 104 configured to mix the knowledge extracted by the multimodal content parser 102 module from said different sources of multimodal data. The knowledge mixer module 104 uses different contextual cues to mix data, such as, recency of the information extracted from the multimodal data, and popularity of the multimodal data, for example, views of a Youtube® video or likes of an Instagram® Post or Reel, etc.
[0094] In some embodiments, the system 100 includes a knowledge graph builder module 106 configured to build a knowledge graph based on the knowledge produced by the knowledge mixer module 104. The knowledge graph builder module 106 creates a detailed relational knowledge graph 107 using the information received from the knowledge mixer module 104.In some embodiments, the system 100 includes a grounding and fine-tuning module 108 that grounds and fine-tunes a foundational model 109 using the information in the knowledge graph. Some well-known foundational models that can be used with the system 100 are: OpenAI® Generative Pre-trained Transformer 4 (GPT-4®), Meta® Llama, Google® Gemini®, AWS® Amazon Titan®, Anthropic Claude®, Cohere Command®, Databricks DBRX®, IBM Granite®, Microsoft Phi®, Mistral Al®, and Nvidia Nemotron®.
[0095] In some embodiments, the system 100 includes an Al playbook agent module 110 that extracts relevant knowledge from the grounded and fine-tuned foundational model 109 in response to a user query to generate personalized advice and recommendations about a product for a user.
[0096] In some embodiments, the system 100 comprises an omni-channel integration sub-system 112 that ensures integration of the system 100 across different platforms and channels, including a website, a mobile application, a desktop application, and an in-store device. In some embodiments, the omni-channel integration sub-system 112 comprises a number of application programming interfaces (APIs) discussed in detail in sections below.
[0097] In some embodiments, the system 100 comprises an action framework 114 that communicates with external service providers 24 for booking services external to the system 100. In some embodiments, the action framework 114 comprises external action APIs that communicate with external service providers 24 via their websites or other internet-based interfaces.
[0098] FIG. 2 discloses an overview of the method 200 for producing a user-specific hyper personalized product advisory using the intelligent multimodal knowledge synthesizer system 100.
[0099] In some embodiments, the method comprises a step 202 of retrieving multimodal content about a product or a user from the network 12. This multimodal content 22 can be stored temporarily on memory of the server 14 or on the data storage medium 16 or on a cloud storage platform.In some embodiments, the method comprises a step 204 of passing the multimodal content through the multimodal content parser module 102 for extracting and analyzing data from the multimodal content to produce a knowledge schema for each piece of multimodal content.
[0100] In some embodiments, the method comprises a step 206 of mixing the knowledge schemas of each multimodal content passed through the multimodal content parser 102 in the knowledge mixer module 104 to create a mixed knowledge schema.
[0101] In some embodiments, the method comprises a step 208 of building a knowledge graph 107 based on the mixed knowledge schema by the knowledge graph builder module 106.
[0102] In some embodiments, the method comprises a step 210 of grounding and fine-tuning the foundational model 109 using the knowledge graph 107, by the grounding and fine-tuning module 108.
[0103] In some embodiments, the method comprises a step 212 of querying the grounded and fine-tuned foundational model 109 in response to a user query and generating a personalized advice and recommendation about the product for the user, by the Al playbook agent module 110.
[0104] The method 200 steps disclosed above are only an overview to provide a bird’s eye view of the working of the system 100. Each step of the method 200 further includes several sub-steps which will be revealed in the discussions below.
[0105] Additionally, although, the steps in the above embodiment of the method 200 are shown to be performed sequentially. In some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any other synchronous or asynchronous format that would yield the end results as the embodiment of the method 200 discussed above.
[0106] The following sections describe each system component and associated method steps in detail.Multimodal content Parser 102
[0107] As described above, the multimodal content parser 102 includes an image and video content parser 300, an audio content parser 700, a text content parser 1100, and an loT content parser 1500. The multimodal content parser 102 directs the content input to it in one of said content parsers based on the type of the content. Each of aforesaid content parsers are discussed in detail in sections below.
[0108] Image and Video Content Parser
[0109] FIG. 3 is a schematic diagram representing an image and video content parser 300 configured to parse video content, for example, an input image or video 22a about the product in focus to extract meaningful information from said input image or video 22a. The image and video content parser 300 includes a content slicer 300a that fragments an input video into video fragments of different sizes. The image and video content parser 300 includes at least one sliding window analyzer 300b that analyzes at least one video fragment at a time to generate content knowledge about the video fragment. In some embodiments, the sliding window analyzer 300b runs machine learning based programs such as image processing, object identification, text identification, audio signal processing, speech recognition, speech to text conversion, language translation, text summarization, key words and key concepts identification for each frame of a video fragment to extract relevant knowledge from the video fragment. The image and video content parser 300 includes a content knowledge module 300c that stores the knowledge extracted from each video fragment and creates a knowledge schema. The image and video content parser 300 includes an Industry Template Record Store 300d that includes preexisting industry templates and knowledge schemas that can enhance the knowledge extracted from the video fragments. The knowledge schema created from the knowledge extracted from all the video fragments is passed to the Knowledge Mixer 104 for mixing with knowledge schemas created from other multimodal content.In some embodiments, the image and video content parser 300 parses static images as well, by treating each image as a single video fragment and repeating the process discussed above. In some embodiments, if the image size exceeds the size limit of a single video fragment, the image can be fragmented into multiple fragments and parsed by the image and video content parser 300 similar to the parsing of a video file as discussed above.
[0110] FIG. 4 is a flow diagram representing a method 400 followed by the image and video content parser 300 shown in FIG. 3.
[0111] In some embodiments, the method comprises a step 402 of receiving an input video file from the memory of the server 14 or the data storage medium 16 or any cloud data storage on the network 12.
[0112] In some embodiments, the method comprises a step 404 of creating video fragments of the video file 22a by the content slicer 300a. In some embodiments, the content slicer 300a can fragment the video file 22a into video fragments based on the time length of the fragment, for example, 15 seconds, 30 seconds, 60 seconds, etc. In some other embodiments, the content slicer 300a can fragment the video based on the size of the video clips, for example, 16 Mb, 32 Mb, 64 Mb, etc.
[0113] In some embodiments, the method comprises a step 406 of creating content knowledge about each video fragment by the sliding window analyzer 300b using machine learning based programs as discussed above. In some other embodiments, the sliding window analyzer 300b inputs the video fragment to a video to text conversion LLM (large language model), and receives content knowledge, such as multilingual video transcription, and description of video scenes as textual output from the LLM, which are then further processed to form the content knowledge about the video fragment. In some embodiments, in step 406, pre-stored knowledge schemas from the Industry Template Record Store 300d are merged with the content knowledge extracted by the sliding window analyzer 300b from each video fragment. In some embodiments, in step 406, the content knowledge extracted from previous video fragments is fed intothe sliding window analyzer 3OOb for the current video fragment to maintain contextual consistency in the content knowledge extracted from the video fragments.
[0114] In some embodiments, the method comprises a step 408, in which the content knowledge created from feeding each video fragment by the sliding window analyzer 300b into the content knowledge module 300c that combines the content knowledge of all the video fragments of the input video file 22a and creates a knowledge schema in the knowledge graph data structure and stores it in memory of the server 14 or on the data storage medium 16 or on the cloud. In some embodiments, in step 408, a social media engagement matrix comprising data points such as likes / dislikes, comments, view count, number of followers of the creator, etc. are also fed into the content knowledge module 300c to create the knowledge schema.
[0115] In some embodiments, the method comprises a step 410, where the knowledge schema created by the content knowledge module 300c is passed to the knowledge mixer 104, where knowledge schemas from multiple different pieces of multimodal content are merged to form a single large knowledge schema about the product or user.
[0116] Although, the steps in the above embodiment of the method 400 are shown to be performed sequentially. In some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any other synchronous or asynchronous format that would yield the end results as the embodiment of the method 400 discussed above.
[0117] FIG. 5 is a schematic diagram representing an image and video crawler 500 for the image and video content parser 300. The image and video crawler 500 crawls image and video platforms or repositories, such as Youtube® or Instagram® or others on the internet to download and analyze images and videos on said repositories relating to the product in focus to feed into the image and video content parser 300.In some embodiments, the image and video crawler 500 includes a platform selection module 502 that selects the image and / or video content platform or repository to download and analyze videos from, such as Youtube® or Instagram® or others.
[0118] In some embodiments, the image and video crawler 500 includes a search module 504 that generates specific searches related to the product in focus to identify relevant images and videos to download.
[0119] In some embodiments, the image and video crawler 500 includes a download module 506 that downloads the image and / or video file that is identified to be related to the product or the user.
[0120] In some embodiments, the image and video crawler 500 includes a Social Media Engagement Matrix Extraction Module 508 that extracts the social media engagement matrix of the image or the video.
[0121] In some embodiments, the downloaded image and / or video file from the download module and its social media engagement matrix from the social media engagement matrix extraction module 508 are fed into the image and video content parser 300.
[0122] In some embodiments, image and video crawler 500 comprises an image and video deletion module 510 that deletes the image and / or video and its associated social media engagement matrix after feeding them into the video content parser 300.
[0123] FIG. 6 discloses a method 600 of operation of the image and video crawler 500 as shown in FIG. 5.
[0124] In some embodiments, the method 600 comprises a step 602 of receiving an input information about a user or a product. In some embodiments, the input information can be user or product information retrieved from a user database or a productcatalog / database. In some other embodiments, the input information can be provided by the user or by an administrator of the server 14 or the input information may automatically be provided via the user application 18 that connects the system 100 to a product web page on an online marketplace or digital store linked to the system 100.
[0125] In some embodiments, the method 600 comprises a step 604 of selecting a platform to search for the images and / or videos about the user or the product via the platform selection module 502. In some embodiments, the platform can be selected sequentially from a pre-stored list stored in the memory of the server 14. In some other embodiments, the platform selection module 502 may use an internet search utility to dynamically identify video content platforms on the internet 12.
[0126] In some embodiments, the method 600 comprises a step 606 of searching for images and / or videos about or related to the user or product, via the search module 504. In some embodiments, the search module 504 comprises a keyword and / or key string generation application that generates keywords and key strings for searching images and videos on the image and video content platforms.
[0127] In some embodiments, the method comprises a step 608 of downloading the image and / or video file identified as an image and / or video related to the user or product via the download module 506. In some embodiments, the download module 506 downloads the image and / or video file to the memory of the server 14. In some embodiments, the download module 506 has pre-saved permissions for downloading image and video content from the image and video content platform owners.
[0128] In some embodiments, the method 600 comprises a step 610 of extracting the social media engagement matrix of the image and / or video being downloaded simultaneously as the video is being downloaded. The social media engagement matrix comprises data points such as likes / dislikes, comments, view count, number offollowers of the creator, name and other relevant details of the creator, title and description of the video file, etc.
[0129] In some embodiments, the method 600 comprises a step 612 of feeding the downloaded image and / or video file and corresponding social media engagement matrix as an input to the image and video content parser 300, as described above.
[0130] In some embodiments, the method 600 comprises a step 614 of deleting the image and / or video file and its corresponding social media engagement matrix from the memory of the server, after it is processed by the Video Content Parser 300.
[0131] Although, the steps in the above embodiment of the method 600 are shown to be performed sequentially. In some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any other synchronous or asynchronous format that would yield the end results as the embodiment of the method 600 discussed above.
[0132] Audio Content Parser
[0133] FIG. 7 is a schematic diagram representing an audio content parser 700 configured to parse audio content, for example, an input audio 22b about a product or a user to extract meaningful information from said input audio 22b.
[0134] In some embodiments, the audio content parser 700 includes a content slicer 700a that fragments the input audio into audio fragments of different sizes.
[0135] In some embodiments, the audio content parser 700 includes at least one sliding window analyzer 700b that analyzes at least one audio fragment at a time to generate content knowledge about the audio fragment.
[0136] In some embodiments, the sliding window analyzer 700b runs machine learning based programs such as audio signal processing, speech recognition, speech to textconversion, language translation, text summarization programs, keywords and key concepts identification on each audio fragment to extract relevant knowledge from the audio fragment.
[0137] The audio content parser 700 includes a content knowledge module 700c that stores the knowledge extracted from each audio fragment and creates a knowledge schema. The audio content parser 700 includes an Industry Template Record Store 700d that includes preexisting industry templates and knowledge schemas that can enhance the knowledge extracted from the audio fragments. The knowledge schema created from the knowledge extracted from all the audio fragments is passed to the Knowledge Mixer 104 for mixing with knowledge schemas created from other multimodal content.
[0138] FIG. 8 is a flow diagram representing a method 800 followed by the audio content parser 700 shown in FIG. 7.
[0139] In some embodiments, the method comprises a step 802 of receiving an input audio file from the memory of the server 14 or the data storage medium 16 or the cloud.
[0140] In some embodiments, the method comprises a step 804 of creating audio fragments of the audio file 22b by the content slicer 700a. In some embodiments, the content slicer 700a can fragment the audio file 22b into audio fragments based on the time length of the fragment, for example, 15 seconds, 30 seconds, 60 seconds, etc. In some other embodiments, the content slicer 700a can fragment the audio based on the size of the audio clips, for example, 16 Mb, 32 Mb, 64 Mb, etc.
[0141] In some embodiments, the method comprises a step 806 of creating content knowledge about each audio fragment by the sliding window analyzer 700b using machine learning based programs as discussed above.
[0142] In some other embodiments, the sliding window analyzer 700b inputs the audio fragment to an audio to text conversion LLM (large language model), and receives content knowledge, such as multilingual audio transcription as textual output from theLLM, which are then further processed to form the content knowledge about the audio fragment.
[0143] In some embodiments, in step 806, pre-stored knowledge schemas from the Industry Template Record Store 700d are merged with the content knowledge extracted by the sliding window analyzer 700b from each audio fragment.
[0144] In some embodiments, in step 806, the content knowledge extracted from previous audio fragments is fed into the sliding window analyzer 700b for the current audio fragment to maintain contextual consistency in the content knowledge extracted from the audio fragments.
[0145] In some embodiments, the method comprises a step 808, where the content knowledge created from each audio fragment by the sliding window analyzer 700b is fed into the content knowledge module 700c that combines the content knowledge of all the audio fragments of the input audio file 22b and creates a knowledge schema in the knowledge graph data structure and stores it in memory of the server 14 or the data storage medium 16 or the cloud.
[0146] In some embodiments, in step 808, a social media engagement matrix comprising data points such as likes / dislikes, comments, view count, number of followers of a creator, etc. are also fed into the content knowledge module 700c to create the knowledge schema.
[0147] In some embodiments, the method comprises a step 810, where the knowledge schema created by the content knowledge module 700c is passed to the knowledge mixer 104, where knowledge schemas from multiple different pieces of multimodal content are merged to form a single large knowledge schema about the product in focus.
[0148] Although, the steps in the above embodiment of the method 800 are shown to be performed sequentially; in some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any othersynchronous or asynchronous format that would yield the end results as the embodiment of the method 800 discussed above.
[0149] FIG. 9 is a schematic diagram representing an audio crawler 900 for the audio content parser 700. The audio crawler 900 crawls audio platforms or repositories, such as Spotify® or Soundcloud® or others on the internet to download and analyze audio on said repositories relating to the product or the user to feed into the audio content parser 700.
[0150] In some embodiments, the audio crawler 900 includes a platform selection module 902 that selects the audio content platform or repository to download and analyze audio files from, such as Spotify® or Soundcloud® or others.
[0151] In some embodiments, the audio crawler 900 includes a search module 904 that generates specific searches related to the product or the user to identify relevant audio to download.
[0152] In some embodiments, the audio crawler 900 includes a download module 906 that downloads the audio file that is identified to be related to the product or the user.
[0153] In some embodiments, the audio crawler 900 includes a Social Media Engagement Matrix Extraction Module 908 that extracts the social media engagement matrix of the audio.
[0154] In some embodiments, the downloaded audio file from the download module and its social media engagement matrix from the social media engagement matrix extraction module are fed into the audio content parser 700.
[0155] In some embodiments, audio crawler 900 comprises an audio deletion module 910 that deletes the audio and its associated social media engagement matrix after feeding them into the audio content parser 700.FIG. 10 discloses a method 1000 of operation of the audio crawler 900 as shown in FIG. 9.
[0156] In some embodiments, the method 1000 comprises a step 1002 of receiving an input about the product or the user. In some embodiments, the input can be provided by the user or by an administrator of the server 14 or can be stored on the data storage medium 16 or the input may automatically be provided via the user application that connects the system 100 to a product web page on an online marketplace or digital store linked to the system 100.
[0157] In some embodiments, the method 1000 comprises a step 1004 of selecting a platform to search for the audios about the product or the user via the platform selection module 902. In some embodiments, the platform can be selected sequentially from a pre-stored list stored in the memory of the server 14. In some other embodiments, the platform selection module 902 may use an internet search utility to dynamically identify audio content platforms on the internet 12.
[0158] In some embodiments, the method 1000 comprises a step 1006 of searching for audio about or related to the product or the user, via the search module 904. In some embodiments, the search module 904 comprises a keyword and / key string generation application that generates keywords and key strings for searching audio content on the audio content platforms.
[0159] In some embodiments, the method comprises a step 1008 of downloading the audio file identified as an audio related to the product or the user via the download module 906. In some embodiments, the download module 906 downloads the audio file to the memory of the server 14. In some embodiments, the download module 906 has pre-saved permissions for downloading audio content from the audio content
[0160]
[0161] In some embodiments, the method 1000 comprises a step 1010 of extracting the social media engagement matrix of the audio being downloaded simultaneously as the audio is being downloaded. The social media engagement matrix comprises data points such as likes / dislikes, comments, view count, number of followers of the creator, name and other relevant details of the creator, title and description of the audio file, etc.
[0162] In some embodiments, the method 1000 comprises a step 1012 of feeding the downloaded audio file and corresponding social media engagement matrix as an input to the audio content parser 700, as described above.
[0163] In some embodiments, the method 1000 comprises a step 1014 of deleting the audio file and its corresponding social media engagement matrix from the memory of the server 14 or the data storage medium 16, after it is processed by the Audio Content Parser 700.
[0164] Although, the steps in the above embodiment of the method 1000 are shown to be performed sequentially. In some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any other synchronous or asynchronous format that would yield the end results as the embodiment of the method 1000 discussed above.
[0165] Text Content Parser
[0166] FIG. 11 is a schematic diagram representing a text content parser 1100 configured to parse text content, for example, an input text 22c about the product in focus to extract meaningful information from said input text 22c. In some embodiments, the text content parser 1100 includes a text summarizer module 1102 that summarizes the input text 22c. In some embodiments, the text content parser 1100 includes a key topics identifier module 1104 that identifies key topics and themes discussed in the input text 22c. In some embodiments, the text content parser 1100 includes a relational graph creator module 1106 that creates a relational graph betweenthe words of the input text 22c. In some embodiments, the text content parser 1100 includes a social media engagement matrix processing module 1108 that processes a social media engagement information matrix, such as name of creator, connections or subscribers of the creator, title of the text, likes or dislikes of the text, number of reshares of the text, etc. In some embodiments, the output from the text summarizer module 1102, key topics identifier module 1104, relational graph creator module 1106, and the social media engagement matrix processing module 1108 is combined with pre-recorded industry templates from an industry template record store 1110 in a content knowledge generator module 1112 to create a knowledge schema of the input text 22c. The created knowledge schema is passed to the knowledge mixer 104 for mixing with knowledge schemas created from other multimodal content.
[0167] FIG. 12 is a flow diagram representing a method 1200 followed by the text content parser 1100 shown in FIG. 11.
[0168] In some embodiments, the method comprises a step 1202 of receiving an input text file 22c from the memory of the server 14 or from the data storage medium 16, or the cloud.
[0169] In some embodiments, the method comprises a step 1204 of creating a summary of the input text in the text summarizer module 1102. In some embodiments, the text summarizer module 1102 uses an LLM to summarize the input text 22c.
[0170] In some embodiments, the method comprises a step 1206 of identifying key topics and themes in the input text using the key topics identifier module 1104. In some embodiments, the key topics identifier module 1104 uses an LLM to identify key topics and themes in the input text 22c.
[0171] In some embodiments, the method comprises a step 1208 of creating a relational graph of the input text using the relational graph creator module 1106. In some embodiments, the relational graph creator module 1106 uses an LLM to create a relational graph of the input text 22c.
[0172] In some embodiments, the method comprises a step 1210 of processing a social media engagement matrix comprising data points such as likes / dislikes, comments,view count, number of followers of the creator, etc., using the social media engagement matrix processing module 1108.
[0173] In some embodiments, the method comprises a step 1212 of combining the outputs of text summarizer module 1102, key topics identifier module 1104, relational graph creator module 1106, social media engagement matrix processing module 1108, with pre-recorded industry templates from an industry template record store 1110 in content knowledge generator module 1112 to create a knowledge schema in the knowledge graph data structure and store it in memory of the server 14 or the data storage medium 16 or the cloud.
[0174] In some embodiments, the method comprises a step 1214, where the knowledge schema created by the content knowledge generator module 1112 is passed to the knowledge mixer 104, where knowledge schemas from multiple different pieces of multimodal content are merged to form a single large knowledge schema about the product or the user.
[0175] Although, the steps in the above embodiment of the method 1200 are shown to be performed sequentially; in some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any other synchronous or asynchronous format that would yield the end results as the embodiment of the method 1200 discussed above.
[0176] FIG. 13 is a schematic diagram representing a text crawler 1300 for the text content parser 1100. The text crawler 1300 crawls social media and content platforms or repositories, such as Facebook®, X®, Reddit®, Wikipedia®, blogs, forums, websites, etc., on the internet to download and analyze text from said repositories relating to the product or the user to feed into the text content parser 1300.
[0177] In some embodiments, the text crawler 1300 includes a platform selection module 1302 that selects the social media or content platform or repository to downloadand analyze text from, such as, Facebook®, X®, Reddit®, Wikipedia®, blogs, forums, websites, or others.
[0178] In some embodiments, the text crawler 1300 includes a search module 1304 that generates specific searches related to the product in focus to identify relevant text to download.
[0179] In some embodiments, the text crawler 1300 includes a download module 1306 that downloads the text that is identified to be related to the product or the user.
[0180] In some embodiments, the text crawler 1300 includes a Social Media Engagement Matrix Extraction Module 1308 that extracts the social media engagement matrix of the Text, for example, likes and reshares of an article on Facebook® or a post on X®.
[0181] In some embodiments, the downloaded text from the download module 1306 and its social media engagement matrix from the social media engagement matrix extraction module 1308 are fed into the text content parser 1100.
[0182] In some embodiments, text crawler 1300 comprises a text deletion module 1310 that deletes the text and its associated social media engagement matrix after feeding them into the text content parser 1100.
[0183] FIG. 14 discloses a method 1400 of operation of the text crawler 1300 as shown in FIG. 13.
[0184] In some embodiments, the method 1400 comprises a step 1402 of receiving an input about the product or the user. In some embodiments, the input can be provided by the user or by an administrator of the server or the input may automatically be provided via the user application that connects the system 100 to a product web page on an online marketplace or digital store linked to the system 100.In some embodiments, the method 1400 comprises a step 1404 of selecting a text content platform to search for the text about the product or the user via the platform selection module 1302. In some embodiments, the text content platform can be selected sequentially from a pre-stored list stored in the memory of the server 14. In some other embodiments, the platform selection module 1302 may use an internet search utility to dynamically identify social media and content platforms on the internet 12.
[0185] In some embodiments, the text content platform can include any one of online databases, online forums, product brochures, product catalogs. In some embodiments, the text content platform can include any one of social media platforms, such as Facebook®, X®, Reddit®. In some embodiments, the text content platform can include past sales data by sales channels, including store, eCommerce, call-center, third-party marketplaces, product inventory snapshot, product catalogue with various product attributes, and product pricing data by sales channel. In some embodiments, the text content platform can include a database comprising data relating to marketing campaign historical data, views, and number of conversions, etc. In some embodiments, the text content platform can include a database comprising various customer attributes from one or more customer record systems.
[0186] In some embodiments, the method 1400 comprises a step 1406 of searching for text about or related to the product or the user, via the search module 1304. In some embodiments, the search module 1304 comprises a keyword and / or key string generation application that generates keywords and key strings for searching text on the text content platforms.
[0187] In some embodiments, the method comprises a step 1408 of downloading the text identified as text related to the product or the user via the download module 1306. In some embodiments, the download module 1306 downloads the text to the memory of the server 14. In some embodiments, the download module 1306 has pre-saved permissions for downloading text content from the text content platform owners.In some embodiments, the method 1400 comprises a step 1410 of extracting the social media engagement matrix of the text being downloaded simultaneously as the text is being downloaded. The social media engagement matrix comprises data points such as likes / dislikes, comments, view count, number of followers of the creator, name and other relevant details of the creator, title and description of the text, for example, an article, note, discussion forum, etc.
[0188] In some embodiments, the method 1400 comprises a step 1412 of feeding the downloaded text and corresponding social media engagement matrix as an input to the text content parser 1100, as described above.
[0189] In some embodiments, the method 1400 comprises a step 1414 of deleting the text and its corresponding social media engagement matrix from the memory of the server 14, after it is processed by the text content parser 1100.
[0190] Although, the steps in the above embodiment of the method 1400 are shown to be performed sequentially. In some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any other synchronous or asynchronous format that would yield the end results as the embodiment of the method 1400 discussed above.
[0191] loT Sensor Data Content Parser
[0192] FIG. 15 is a schematic diagram representing an loT sensor data content parser 1500 configured to parse loT sensor data, for example, an loT sensor data 22d related to the user and / or the product to extract meaningful information from said IOT sensor data 22d. In some embodiments, the loT sensor data content parser 1500 includes a weather sensor data processing module 1502 that processes data received from weather sensors. In some embodiments, the loT sensor data content parser 1500 includes a user location sensor data processing module 1504 that processes sensor data relating to the location of the user. In some embodiments, the loT sensor data content parser 1500includes a user biosensor data processing module 1506 that processes user’s biological data, such as, pulse, heartbeat, body temperature, E.C.G (Electrocardiogram), or other biological data received from one or more user devices associated with the user. In some embodiments, the loT sensor data content parser 1500 includes a supply-side Sensors Data Processing Module 1508 that processes sensor data relating to the seller or logistics service provider with regards to the product, for example, data pertaining to the location of the warehouse in which the product is stored, tracking data relating to a product ordered by the user, quantity of the product available with the seller, weather conditions at the warehouse. In some embodiments, the output from the weather sensor data processing module 1502, user time and location sensor data processing module 1504, user biosensor data processing module 1506, and supply-side Sensors Data Processing Module 1508 are combined with pre-recorded industry templates from an industry template record store 1510 and user demographic and biographic data from a user information record store 1512 in an loT sensor data knowledge generator module 1514 to create a knowledge schema of loT sensor data. The created knowledge schema is passed to the knowledge mixer 104 for mixing with knowledge schemas created from other multimodal content.
[0193] FIG. 16 is a flow diagram representing a method 1600 followed by the loT Sensor data content parser 1500 shown in FIG. 15.
[0194] In some embodiments, the method comprises a step 1602 of receiving an input loT sensor data 22d from the memory of the server 14 or the data storage medium 16 or the cloud. In some embodiments, unlike the use of crawlers used for capturing video, image, audio, or text content, the loT sensor data is received in the memory of the server 14 from trusted and verified data sources, pre-registered with the server 14. In some embodiments, the weather sensor data is received from trusted weather monitoring websites or weather stations. In some embodiments, the user location data is received from the user device 10 (as shown) or any other related user device registered with the user. In some embodiments, the user biosensor data is received from the user device 10 or any other user device registered with the user, for example, anywearable device, such as a smartwatch, or a smart ring. In some embodiments, the supply-side sensors data can be received from several sources, such as seller, warehouse, in-transit transport vehicle, and product packaging. In some embodiments, the supply-side sensors data includes, but not limited to, location or weather conditions of the seller, warehouse, in-transit transport vehicle, product packaging. In some embodiments, the supply-side sensors data includes, but not limited to, the quantity of products available in the warehouse or available with the seller, or available in the closest in-transit transport vehicle.
[0195] In some embodiments, the method comprises a step 1604 of processing weather sensor data in the weather sensor module 1502. In some embodiments, the weather sensor data processing module 1502 determines the current weather and climatic conditions in the region. In some embodiments, in this step 1604, the weather sensor data processing module 1604 encodes this weather information for further processing by the knowledge generator module 1514.
[0196] In some embodiments, the method comprises a step 1606 of processing data relating to the user location in the user location sensor data processing module 1504. In some embodiments, the user location sensor data may include data, such as, GPS sensor data received from the user device 10. In some embodiments, in this step 1606, the user location sensor data processing module 1504 encodes this user location information for further processing by the knowledge generator module 1514.
[0197] In some embodiments, the method comprises a step 1608 of processing user biosensor data received from one or more user devices registered with the user in the user biosensor data processing module 1506. In some embodiments, in this step 1608, the user biosensor data processing module 1506 encodes this user bio-sensor information for further processing by the knowledge generator module 1514.
[0198] In some embodiments, the method comprises a step 1610 of processing the supply side sensor data in the supply-side sensors data processing module 1508. In some embodiments, in this step 1610, the supply-side sensors data processing module1508 encodes this supply side sensor information for further processing by the knowledge generator module 1514.
[0199] In some embodiments, the method comprises a step 1612 of combining data from the weather sensor data processing module 1502, user location sensors data processing module 1504, user biosensors data processing module 1506, and supply-side sensors data processing module 1508 with pre-recorded industry templates from an industry template record store 1510 and user demographic and biographic data from a user information record store 1512 in an loT sensor data knowledge generator module 1514 to create a knowledge schema of loT sensor data in the knowledge graph data structure and stores it in memory of the server 14 or the data storage medium 16 or the cloud.
[0200] In some embodiments, the method comprises a step 1614, where the knowledge schema created by the loT sensor data knowledge generator module 1514 is passed to the knowledge mixer 104, where knowledge schemas from multiple different pieces of multimodal content are merged to form a single large knowledge schema about the product or the user.
[0201] Although, the steps in the above embodiment of the method 1600 are shown to be performed sequentially; in some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any other synchronous or asynchronous format that would yield the end results as the embodiment of the method 1600 discussed above.
[0202] Knowledge Mixer
[0203] FIG. 17 is a schematic diagram representing the relationship between the multimodal content, the multimodal content parser, the knowledge mixer, and the knowledge graph builder.
[0204] In some embodiments, as shown, the multimodal content is fed into the multimodal content parser 102, which then converts the multimodal content intodiscrete knowledge schemas for each piece of video, audio, text, and loT content, as discussed above.
[0205] In some embodiments, the knowledge schemas created by the multimodal content parser 102 are fed into the knowledge mixer 104. The knowledge mixer 104 is a processing module that sorts, rearranges, and merges knowledge schemas as per a recency and popularity factors model. In some embodiments, for example, a first knowledge schema of a piece of content is more recent than a second knowledge schema about another piece of content, both relating to the same topic or theme, then the first knowledge schema will get a higher precedence over the second knowledge schema while mixing the knowledge schemas. Similarly, if a third knowledge schema about a piece of content is more popular (for example, has more views, likes, votes, or a combination thereof) than a fourth piece of content, both relating to the same topic or theme, then the third knowledge schema will get a higher precedence over the fourth knowledge schema. In some embodiments, the knowledge mixer sorts the information provided in all the knowledge schemas it receives from the multimodal content parser, removes redundancies, and creates a single knowledge schema about the product or the user which is fed into the knowledge graph builder 106.
[0206] In some embodiments, the knowledge mixer 104 sorts, rearranges, and merges knowledge schemas as per a dynamic multi-dimensional knowledge structure that comprises a multi-dimensional knowledge vector (K) to represent complex product knowledge for each product, or SKU (Stock Keeping Unit), where the knowledge vector is represented as:
[0207]
[0208] Where‘s’ represents an SKU,
[0209] ‘t’ represents a time window,
[0210] ‘Sbase(s,t)’ is a base score that combines Recency ‘R(s,t)’ and Popularity ‘P(s,t), and is calculated by the formula:
[0211]
[0212] where wrand wpare weight coefficients,
[0213] ' V(s, t)' is Sales Revenue for SKU 's' for the given time window 't' ,
[0214] 'M(s, t)' is Product Margin for SKU ' s' for the given time window 't' ,
[0215] 'C(s, t)' is Marketing Sponsorship / Spend for SKU 's' for the given time window 't' ,
[0216] ‘Vnorm(s,t)’ is normalized sales revenue within the range of [0,1], calculated using the formula:
[0217] >
[0218]
[0219] Where Vmin(t) and Vmax(t) are the minimum and maximum revenues across all relevant SKUs in the time window ‘t’, similarly,
[0220] ‘MnOmi(s,t)’ is normalized product margin within the range of [0,1], and
[0221] ‘Cnorm(s,t)’ is normalized marketing / sponsorship spend within the range of [0,1], Table 1 depicts an example of a metric being normalized. A metric value from 1 to 5 that has been normalized to a value between 0 to 1 using the above formula.
[0222] For example, rating 2 would become:
[0223] (2-l) / (5-l) = % = 0.25Table 1
[0224]
[0225] Similarly, for a metric having value from 0-100, a metric value of 65 would become:
[0226] (65-0) / (100-0) = 65 / 100 = 0.65
[0227] In some embodiments, as part of normalisation, and to remove statistical error introduced by outliers, a ‘Robust Scaling’ technique is used, which is robust to outliers and scales data using statistics that are themselves resistant to outliers, namely the median and the interquartile range (IQR), where:
[0228] xi' = (xi-median(x)) / IQR(x)
[0229] where the IQR is the range between the 1st quartile (25th percentile) and the 3rd quartile (75th percentile). In such embodiments, because extreme values do not influence the median and IQR in the tails of the distribution, the ‘Robust Scaling’ technique is more stable when the data contains outliers. In some embodiments, for metrics like metrics (V, M, C) in the Knowledge Vector, Robust Scaling is a highly recommended approach as it will ensure that extreme outliers.
[0230] In some embodiments, the multi-attribute knowledge vector (K) provides a comprehensive and commercially relevant understanding of a product's value within
[0231]
[0232] In some embodiments, the Knowledge Vector (K), by integrating a base engagement score (S_base) with crucial business metrics like sales revenue (V), product margin (M), and marketing spend (C), aligns with a necessary evolution in sales channel intelligence, which demands a "thorough perspective" that balances both customer needs and inventory authority, including financial performance and stock availability. The Knowledge Vector (K) allows segmenting products not merely by popularity but by their strategic role as "best-sellers, profitable items, and VIP items," a classification that requires the dimensions revenue (V), product margin (M), and marketing spend (C).
[0233] A knowledge vector can also be represented in a multidimensional space. Fig.
[0234] 17a depicts the spatial representation of two-dimensional knowledge vectors created by knowledge vectors construction module. Three vectors KV1, KV2 and KV3 are shown having two dimensions each Vnorm and Mnorm. The three vectors are spread in the two-dimensional space depending on their values of Vnorm and Mnorm. A person skilled in the art would appreciate that the knowledge vectors can have any number of dimensions (i.e., Vnorm, Mnorm, Cnorm, Sbase, etc.).
[0235] In some embodiments, the knowledge mixer (104) aggregates SKU-level Knowledge Vectors (V) into dynamic higher-level Aggregation Tensor, T(E,t).
[0236] Aggregation of SKU level knowledge vectors into an Aggregation Tensor helps in answering higher-level, context-dependent queries. This approach allows for a shift from the traditional static, rule-based weighting system to a dynamic, learned model that can capture the complex nuances of user intent and adapt in real-time. To create an Aggregation Tensor we arrange the constructed knowledge vectors into multidimensional tensor space which naturally represents the multi-dimensional relationships between various Knowledge Vectors, allowing for a more nuanced and flexible aggregation process.Fig. 17b depicts arrangement of knowledge vectors KV1-KV9 into a three-dimensional tensor with generic dimensions DI, D2 and D3 created by tensor creation module.
[0237] Fig. 17c depicts a simpler arrangement of knowledge vectors KV1-KV9 into a two-dimensional tensor with only two dimensions (DI and D2) and without tensor gird for ease of representation.
[0238] As shown in Fig. 17b, the knowledge vectors can be arranged into aggregation tensor consisting of higher dimensions (i.e., DI: Category, D2: Geography, D3: Demographics). These tensors can be sliced by different dimensions such as DI: i.e., product category, D2: i.e., geography, and D3: i.e., demographics, which enables the knowledge mixer to do a context-aware mixing or knowledge packaging. This structure enables an Al Playbook agent 110 to move beyond generic recommendations and answer specific queries, such as "What are the trending running shoes for women in the Northeast?" by directly accessing the relevant tensor slice. This approach is particularly valuable in domains such as e-commerce, where situational factors such as time, location, and social context heavily influence user preferences.
[0239] In some embodiments, the knowledge mixer aggregates SKU-level Knowledge Vectors (V) into dynamic higher-level Aggregation Tensor, T (E, t). The Tensor, T (E, t) can be sliced by dimensions, such as, product category (J), geography (G) and enables the knowledge mixer to do a context-aware mixing or knowledge packaging.
[0240] Fig. 17d depicts a Tensor (E, t) created by tensor creation module, aggregating a plurality of knowledge vectors (total 16) with four metrics each S (s, t), V (s, t), M (s, t) and C (s, t). The tensor can be sliced among multiple dimensions. For example, we may slice it based on product category (J) and geography (G). Therefore, for a query "Show me all the running shoes", the Tensor can be sliced according to the product category (J), where there are four possible categories:
[0241] JI: Formal ShoesJ2: Sneakers
[0242] J3: Hiking Shoes
[0243] J4: Running Shoes
[0244] And the outcome would be a whole slice J4 containing three knowledge vectors related to three SKUs or products.
[0245] Further, if there are three possible geographies:
[0246] Gl: Hills
[0247] G2: Dessert
[0248] G3: Plains
[0249] Then for a more specific query like “Show me the running shoes best for hilly terrains”, then the Tensor can be sliced by tensor slicing module against the dimensions category (J) and geography (G). The output Tensor slice would be J4+G1 containing one knowledge vector.
[0250] Fig. 17d only depicts two possible dimensions for slicing the Tensor. However, there can be many more dimensions into a Tensor. As the number of dimensions grow the Tensor can be used to address more and more specific queries. For example, for a query, “What are the trending running shoes for women ideal for hilly terrain during winters”. There is a total number of four dimensions in this query: shoes, gender, terrain and weather. Therefore, a Tensor with four possible dimensions can be created to answer such a query.
[0251] Fig. 17e depicts the method (17100) followed by the knowledge vectors construction module and the tensor creation module, comprising receiving Raw Data Metrics in step 17102 like R(s,t), P(s,t), V(s,t), M(s,t), C(s,t) related to an SKU by the knowledge mixer. In step 17104, the knowledge mixer computes base engagement score, which is calculated based on the recency and popularity of the product. In step 17106 the Raw Data Metrics like R(s,t), P(s,t), V(s,t), M(s,t), C(s,t) are normalized tothe value [0,1], In step 17108, the robust scaling is done to the normalized values of the Raw data Metrics. In step 17110, a multi-dimensional knowledge vector is constructed using the normalized Raw Data Metrics and the base engagement score. In Step 17112 the constructed knowledge vectors are arranged in a Multi-dimensional Tensor.
[0252] In some embodiments, the SKU level Knowledge Vectors are_clustered under multiple clusters. Fig. 17f depicts clustering of knowledge vectors KV1-KV9 into three clusters C1-C3 performed by knowledge vector clustering module.
[0253] Fig. 17g depicts a flowchart describing the method of clustering of SKU level Knowledge Vectors followed by knowledge vector clustering module. The method begins by ingesting (17202) the full set of SKU-level Knowledge Vectors K(s, t), followed by selection of the clustering algorithm (17204). SKU-level Knowledge Vectors K(s, t), are then processed by a high-dimensional vector indexing algorithm (17208) that partitions the vector space and groups nearby vectors into clusters or related structures based on similarity (17206) computation (cosine similarity and / or Euclidean distance). This indexing / clustering can be performed using various algorithms (17204), such as graph-based approaches, i.e., HNSW which construct a navigable graph connecting geometrically close vectors, tree-based techniques like Annoy which recursively divide the space with hyperplanes and quantization -based strategies that first apply k-means clustering and store results in an inverted file index. Other clustering algorithms like DBSCAN and Hierarchical can also be used. The output is a persistent, queryable knowledge cluster / index or vector database that enables sub-linear lookup performance (e.g., O(log n)) rather than requiring a full linear scan.
[0254] In some embodiments as shown in step 17210, clusters can be characterized and labeled. In some embodiments as shown in step 17212, strategic insights can be extracted from the clusters. In some embodiments as shown in step 17214, behaviorally refined product segments can be defined based on the created clusters.In some embodiments, as shown in Fig. 17h, a user’s query context is extracted from a user’s query by context vectorization module.
[0255] In this phase, the context of the user query is extracted, for example, a keyword “budget” used in the user’s query can be translated to low price. Therefore, only the SKU / products with low price will be selected. However, in case of apparels, budget may mean less than Rs. 1000, however in case of smartphone, budget may mean less than Rs. 10000 and similarly for cars, budget may mean less than Rs 5 Lakhs. Similarly, for apparels, “trending” may mean high sales. However, in case of smartphones high revenue is a much better proxy for "trending", since a premium product like an Apple® iPhone® might not sell even half the number of a cheaper android phone. However, the overall revenue of iPhone® might be 4 to 5 times than that of its cheaper counter parts. The right metric to gauge the trending or bestseller behaviour in case of electronics would be revenue. Therefore, considering the word “premium” from the query can be encoded to “High Revenue” and “High Margin”. Therefore, depending on the category of the SKU / product, a single keyword can be translated to different contexts. This translation is performed by the query context encoding module.
[0256] Table 1 shows few real-life examples of context-based encoding of query keywords.
[0257] Table 1
[0258] &
[0259]
[0260] &
[0261] &
[0262]
[0263] In some embodiments, as shown in Fig. 17i, the extracted query context from the user’s query is encoded into a query context vector, that can be mapped into a multi-dimensional vector space in which the clusters of knowledge vectors are present.
[0264] In some embodiments, the encoding may result in knowledge vector weights. As explained above, the word “premium” from the query can be encoded to “High Revenue” and “High Margin”. Therefore, as shown in Fig. 17i based on the keyword and corresponding encoding, the attention may get focused on Revenue and Margin. The corresponding attention weights would become:
[0265] S(s,t) = 0 (Base Score)
[0266] V(s,t) = 1 (Revenue)
[0267] M(s,t) = 1 (Margin)
[0268] C(s,t) = 0 (Marketing Spend).
[0269] In some embodiments, the most relevant cluster is selected based on the encoded context. The above context encoding basically means that the query wants to find out high revenue and high margin running shoes with no consideration to base score andmarketing spend. The output would be knowledge vector weights according to the encoded query containing attention weights.
[0270] In Fig. 17i if the dimensions Hl, H2 and H3 represent:
[0271] Hl as comfortable then all attributes would have average score
[0272] H2 as bestseller then S(s,t) signifying Base Score would have higher score and all other attributes average
[0273] H3 as premium then R(s,t) signifying Revenue and M(s,t) signifying Margin both would be higher score while other attributes would have average score.
[0274] Therefore, according to Fig. 17i, when the knowledge vectors of slice J4 are weighted for “premium” keyword, the knowledge vector in slice J4+H3 (i.e., the most relevant cluster) would have the highest weight since it has a high score against Revenue and Margin and the query also focuses on these two attributes only.
[0275] The above processing enables real-time (sub-second) query responses over product catalogues with millions of SKUs. The retrieval time scales sub-linearly O(logn) with the number of SKUs. This processing avoids computationally wasteful attention calculations on irrelevant items, focusing resources on a small set of promising candidates, combining the speed of an approximate k-NN search (Retrieval) with the precision of a complete neural attention model (Ranking), achieving the best of both.
[0276] Fig. 17j gives a generic view of the query context vector acting as a probe for an approximate k-nearest-neighbour search over the Knowledge Clusters / Index, performed by cluster selection module, allowing the system to efficiently identify a small, highly relevant set of candidates (Cl cluster consisting of KV1-KV4 as shown in Fig. 17j and 17k) without evaluating millions of irrelevant SKUs / KVs.
[0277] In some embodiments as shown in Fig. 171, these candidates (the identified most relevant cluster Cl members KV1-KV4) are then passed to the next stage, where thefull atention mechanism such as the neural network or scaled dot-product attention is applied only to this focused subset.
[0278] In some embodiments as shown in Fig. 17m, a query embedding / context vector, along with each SKU's Knowledge Vector K(s,t), is fed into a small neural network (also referenced as: atention network) by atention weight determination module. In some embodiments, the attention network outputs a vector of atention weights, specific to that query and that SKU. In some embodiments, the aggregated knowledge vector is a weighted sum of dynamically computed atention weights. In some embodiments, this method allows for a sophisticated and adaptive understanding of user preferences. For example, for a user query like "profitable and trending," the knowledge mixer might learn to pay high atention to the Margin (M norm) and Base Score (S base) dimensions, as the static model would. However, it might also learn that for the "Electronics" category, Revenue (V norm) is a much better proxy for "trending" than the base engagement score, while for "Fashion," the Marketing Spend (C norm) is more indicative. This ability to "adaptively prioritise features based on the context" is provided by the above-described attention mechanism.
[0279] In some embodiments, a standard scaled dot-product atention score, ‘e’ is calculated as:
[0280]
[0281] Where qctxis a query embedding vector that encodes user's query context (e.g., "profitable and trending"), and
[0282] Where d_k is the dimensionality of the vectors, which serves as a scaling factor to prevent the dot product from growing too large.In some embodiments, the relevance scores are then normalised across all SKUs s' within the target entity E (e.g., a product category) using the SoftMax function. This produces the final weights, a, where each weight represents the relative importance of an SKU for the given query:
[0283]
[0284] In some embodiments, the Aggregated Knowledge Vector, A(E,t), for the entity is then computed as the weighted sum of the individual Knowledge Vectors of its constituent SKUs.
[0285]
[0286] In some embodiments, the knowledge mixer 104 applies clustering algorithms to the SKU-level knowledge vectors (K), which allows the knowledge mixer 104 to move beyond simple product grouping to uncover latent, behaviourally defined segments within the product catalogue.
[0287] Fig. 17n is a flow diagram representing a method followed by context vectorization module, cluster selection module and attention weight determination module. Once the clusters are created, the next process creates attention weights (17300), where first a user’s query is received in step 17302, which is encoded into a query context vector in step 17304. In step 17306 at least one cluster is identified / shortlisted using the query context vector as a probe in the multi-dimensional space. In step 17308 a standard scaled dot product is computed between query embedding / context vector and each knowledge vector from the shortlisted cluster to generate attention weight vectors. In step 17310 the attention weight vectors are normalized using softmax function to create final attention weights.As described above, a dynamic Query Context Vector “q ctx" is generated from the user’s query, which is used to probe the clusters through an approximate k-NN search (17306), retrieving only the most relevant candidates (members of the most relevant one or more clusters) while avoiding computation over millions of SKUs. These candidates are then passed through a precise attention mechanism (17308 and 17310), allowing the system to apply its full neural attention model only where it matters, yielding final, context-dependent attention weights. The output either an Aggregated Knowledge Vector or a ranked list of items constitutes the real-time “knowledge manifestation.” This embodiment offers major advantages: sub -second performance at million-SKU scale, sub-linear O(log^9jn)retrieval complexity, efficient allocation of computational resources, and high accuracy through the combination of approximate retrieval and precise neural ranking.
[0288] Finally, the Fig. 17o, depicts that the high-level method flow followed by the Knowledge Mixer (17400). The Knowledge Mixer (17400) initially accepts the input 17402 (Base Score and Business KPI Data) from the Multimodal Content Parser. In the next step 17404 a knowledge vectors construction module in Knowledge Mixer constructs a Knowledge Vector for each SKU, based on the input data received for each and every SKU. The constructed Knowledge Vectors are further used by the tensor creation module in knowledge mixer to create Aggregated Knowledge Computation Tensor (17406). Next, the Knowledge Clustering is performed by a knowledge vector clustering module at step 17408 where different knowledge vectors are clustered into one or more clusters and after extraction of strategic insights behaviorally defined segments are created. These clusters can be further used to respond to the user queries, with additional steps of query context vectorization, most relevant cluster selection.
[0289] Fig. 17p, represents flow of data between different modules in a knowledge mixer module, which have been explained above.Referring to Fig. 17 again, in some embodiments, the knowledge mixer 104 also takes input Industry Playbook Knowledge Templates from an Industry Playbook Knowledge Templates Database 104a. In some embodiments, the Industry Playbook Knowledge Templates have data schema structure / attributes configurations that conform with product catalogue systems used in eCommerce. In some embodiments, Industry Playbook Knowledge Templates Database 104a is stored in the memory of the server 14 or on the data storage medium 16 or on the cloud. The data schema attributes when input to the foundational model (LLM) module 108 forces its LLM outputs to conform the format of said product catalogue systems, which aids in codification of said outputs in subsequent processing of said outputs in omni-channel integration sub-system 112.
[0290] FIG. 18 is a flow diagram representing a method 1800 followed by the knowledge graph builder 106.
[0291] In some embodiments, the method comprises a step 1802 of fetching mixed knowledge data schema from the knowledge mixer 104. In some embodiments, this fetching can be triggered by the knowledge mixer 104 when a mixed knowledge data schema has been processed by the knowledge mixer 104.
[0292] In some embodiments, the method comprises a step 1804 of fetching relevant industry playbook templates from the Industry Playbook Template Database 104a. This fetching can be triggered by the fetching of the mixed knowledge data schema in step 1802.
[0293] In some embodiments, the method comprises a step 1806 of extracting information from the fetched mixed knowledge data schema about the product or the user, including sentiment analysis, catalog details, and concise text summaries for improved data insights and user understanding about the product.In some embodiments, the method comprises a step 1808 of implementing postprocessing and API calls to format data, including sentiment analysis, summaries, and catalog details, ensuring compatibility with the fetched industry playbook templates for Knowledge Graph integration and improved data usability.
[0294] In some embodiments, the method comprises a step 1810 of uploading data to the Knowledge Graph in nodes and edges format using API for structured representation and efficient querying.
[0295] In some embodiments, the method comprises a step 1812 of implement an embedded Al and data pipeline to store Knowledge graph data in vector format within a vector database for efficient retrieval and similarity search.
[0296] Although, the steps in the above embodiment of the method 1800 are shown to be performed sequentially. In some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any other synchronous or asynchronous format that would yield the end results as the embodiment of the method 1800 discussed above.
[0297] Knowledge Graphs Database Creation
[0298] FIGS. 19A-19B are schematic diagrams representing creation of knowledge graphs database. FIG. 19A is a schematic representing creation of user knowledge graphs database in the system 100, and FIG. 19B is a schematic representing a creation of product knowledge graphs database in the system 100.
[0299] Referring to FIG. 19A, as the process of individual knowledge graph creation is disclosed above, this section provides description of the process 1900a of creating a user knowledge graphs database in data storage 16 or the cloud.In some embodiments, the server 14 may store in its local memory or the data storage medium 16 or the cloud, a user database 1902a comprising information about all registered users of the system 100.
[0300] In some embodiments, a user information crawler 1904a crawls the user database 1902a and one by one feeds the user information of each user from the user database 1902a to the multimodal content parser 102.
[0301] In some embodiments, the multimodal content parser 102 then collects all data available about each user on multimodal data sources 22 and sends the parsed information to the knowledge mixer 104.
[0302] In some embodiments, the knowledge mixer 104 mixes the multimodal knowledge gathered on each user and feeds it to the knowledge graph builder 106.
[0303] In some embodiments, the knowledge graph builder 106 builds a knowledge graph 107 for each user and then stores said knowledge graph on the data storage medium 16.
[0304] In some embodiments, this process 1900a is repeated to create a knowledge graph for each registered user on the user database 1902a and stores it on the data storage medium 16 or the cloud. If a knowledge graph 107 for a user already exists on the data storage medium 16, then the knowledge graph builder 106 updates the existing knowledge graph 107 of said user with new information.
[0305] In some embodiments, the process 1900a is continuously repeated after a fixed interval of time, such as, 1 day, 1 week, 2 weeks, 1 month, 6 months, etc. to update the knowledge graphs for each registered user previously recorded and create new knowledge graphs for newly registered users on the user database 1902a during that period. In some other embodiments, the process 1900a is repeated for a specific user, when said user actively engages the system 100. Several other methods and timelinesof updating the user knowledge graphs can be implemented. In some other embodiments, the process 1900a creates a single large knowledge graph for all registered users on the user database 1902a.
[0306] Referring to FIG. 19B, this section provides description of the process 1900b of creating a product knowledge graphs database in data storage medium 16 or the cloud.
[0307] In some embodiments, the server 14 may store in its local memory or the data storage medium 16 a product catalog or database 1902b comprising information about all products available for sale via the system 100.
[0308] In some embodiments, a user information crawler 1904a crawls the product database 1902b and one by one feeds the product information of each product from the product database 1902b to the multimodal content parser 102.
[0309] In some embodiments, the multimodal content parser 102 then collects all data available about each product on multimodal data sources 22 and sends the parsed information to the knowledge mixer 104.
[0310] In some embodiments, the knowledge mixer 104 mixes the multimodal knowledge gathered on each product and feeds it to the knowledge graph builder 106.
[0311] In some embodiments, the knowledge graph builder 106 builds a knowledge graph 107 for each product and then stores said knowledge graph on the data storage medium 16.
[0312] In some embodiments, this process 1900b is repeated to create a knowledge graph for each available product on the product database 1902b and stores it on the data storage medium 16. If a knowledge graph 107 for a product already exists on the data storage medium 16, then the knowledge graph builder 106 updates the existing knowledge graph 107 of said product with new information.In some embodiments, the process 1900b is continuously repeated after a fixed interval of time, such as, 1 day, 1 week, 2 weeks, 1 month, 6 months, etc. to update the knowledge graphs for each product previously recorded and create new knowledge graphs for newly added products to the product database 1902b during that period. In some other embodiments, the process 1900b is repeated for a specific product, when said product is actively searched or queried by a user on the system 100. Several other methods and timelines of updating the product knowledge graphs can be implemented. In some other embodiments, the process 1900b creates a single large knowledge graph for all available products on the product database 1902b.
[0313] In some embodiments, the processes 1900a and 1900b create a single large knowledge graph for all registered user on the user database 1902a and all available products on the product database 1902b.
[0314] Knowledge Graphs
[0315] FIGS. 20A-20B are schematic diagrams representing the knowledge graphs 107. FIG. 20A depicts a user knowledge graph 107-1 and FIG. 20B depicts a product knowledge graph 107-2. Both the user knowledge graph 107-1 and the product knowledge graph 107-2 can coexist on the data storage medium 16 or the cloud. Additionally, several other user or product knowledge graphs can coexist on the data storage medium 16 for different users or products respectively. Referring to FIGS.
[0316] 20A-20B, as shown, a node or vertex 107-la or 107-2a is a discrete word or a concept in the data. Also, as shown, an edge 107-lb, 107-2b is a logical link joining two nodes 107-la or 107-2a, respectively. In some embodiments, a single node 107-la, 107-2a can be connected to multiple other nodes by edges 107-lb, 107-2b.
[0317] Referring to FIG. 20 A, the user knowledge graph 107-1 allows the system 100 to customize responses to user queries in accordance with the unique features, lifestyle, requirements of the user. For example, the user follows the influencers, named: Pollyand Molly, or the user visits a particular store frequently, or the user has a heart condition, evident from his / her smart watch data, or the user lives in a region of cold climate. In that, the user knowledge graph 107-1 provides the system 100 with the context of the user.
[0318] Referring to FIG. 20B, the product knowledge graph 107-2 allows the system 100 to know the attributes and social acceptance of a product in the market. For example, the product is popular with a certain class of influencers, the product has some recommended usage guidelines, the product has some cautionary information, the product has some side effects for a certain class of people, etc. Additionally, the tracking information of the product in the warehouse or in transit is also important for the system 100 to know to provide good product advisory. In that, the product knowledge graph 107-2 provides the system 100 with context of a best-case scenario in which the product should be used, and which users would benefit the most from said product, etc.
[0319] The knowledge graphs 107-1 and 107-2 allow the system 100 to respond to user queries with customizations tailored to both the user and the product most suitable for the user.
[0320] In some embodiments, the user knowledge graphs 107-1 and product knowledge graphs 107-2 are all combined into one single large knowledge graph.
[0321] Grounding and Fine-tuning Module
[0322] FIG. 21 is a schematic diagram representing the grounding and fine-tuning foundational model module 108.
[0323] In some embodiments, the grounding and fine-tuning foundational model module 108 receives as input, the knowledge graphs 107 (user knowledge graph 107-1 and product knowledge graph 107-2) and first-party vector database, which comprise files provided by the manufacturer or seller about the product, such as brochures, manuals,specifications, etc, or details of associated service providers registered with the system 100, from the data storage medium 16, and feeds the same into a foundational model 109.
[0324] In some embodiments, the grounding and fine-tuning foundational model module 108 receives as input, industry templates, which are text configuration files used to refine the system context, and help in effective response formulation or creating / executing a relevant action plan, from a reinforcement, feedback, and training application 108aand feeds into the foundational model 109.
[0325] In some embodiments, the foundational model 109 learns about the specifics of the user and the product using these inputs, and is converted into a grounded and finetuned foundational model 109, which can interact with the Al playbook agent module 110.
[0326] In some embodiments, there can be a grounded and fine-tuned foundational model 109 each for several products and user profiles saved on the server 14, data storage medium 16, or the cloud. In some alternate embodiments, a single foundational model can be grounded and fine-tuned for each product and user profile for a user query or a user communication session, on the fly, when the user selects said product on the user device application 18.
[0327] Al playbook agent module
[0328] FIG. 22 is a schematic diagram representing the interaction between the user device application 18, the Omni Channel Integration Sub-system 112, the Al Playbook Agent 110, and the Grounded and Fine-tuned Foundational Model 109; and the interaction between Al Playbook agent 110 and the Action Framework 114 and External Service Providers 24.
[0329] In some embodiments, the user device application 18, which can be any of a web browser, a chatbot, a mobile application, or others, allows a user to input a query about the product in the user device application 18. In some embodiments, the user query canbe in any form, such as, a text, audio format, audio-video format, or any combination thereof.
[0330] In some embodiments, the user query is communicated to the Al playbook agent 110 via the omni-channel integration sub-system 112.
[0331] In some embodiments, the Al playbook agent 110 uses an ontology service to select the appropriate product knowledge graph 107-2 based on the user query.
[0332] In some embodiments, the Al playbook agent 110 sends the user query to the foundational model 109 grounded and fine-tuned, by the grounding and fine-tuning foundational model module 108, with said appropriate product knowledge graph 107-2 and the user knowledge graph 107-1 of the user making the query. In some embodiments, the Al playbook agent 110 sends a reference of the appropriate user knowledge graph 107-1 and product knowledge graph 107-2 to the grounding and fine-tuning foundational model module 108 based on the user query or user communication session at the user device application 18, for grounding and fine-tuning the foundational model 109.
[0333] In some embodiments, the grounded and fine-tuned foundational model 109 provides to the Al playbook agent 110, a hyper personalized response to the modified user query.
[0334] In some embodiments, the Al playbook agent 110 sends the hyper personalized response to an action framework 114. In some embodiments, the Al playbook agent 110 also sends relevant data from 1stparty vector database of external service providers 24 and product customizations stored in the data storage medium 16 to the Action Framework 114. In some embodiments, the action framework 114 uses an ontology service to determine and provide details of relevant external services and product customizations to the Al playbook agent 110 based on the output of the grounded and fine-tuned foundational model 109 and data from the 1stParty Vector database.
[0335] In some embodiments, the Al playbook agent 110 updates the hyper personalized response with suggestions and links about product customizations and external services suitable for the user.In some embodiments, the Al playbook agent 110, sends this hyper personalized response to the user device application 18, via the omni-channel integration sub-system 112.
[0336] In some embodiments, the user device application 18 communicates the hyper personalized response to the user. In some embodiments, the user device application 18 can communicate the response via a text message, a multimedia message, a voice message, or an audio-visual message created using an Al Avatar, such as Al Avatars provided by HeyGen®, or Veed Studios®.
[0337] FIG. 23 is a flow diagram representing a method 2300 followed by the Al Playbook Agent 110;
[0338] In some embodiments, in a step 2302, the Al Playbook Agent 110 receives a user query from the user device application 18 via the omni channel integration sub-system 112.
[0339] In some embodiments, in a step 2304, the Al Playbook Agent 110, passes the user query through an ontology selection service to determine a relevant product knowledge graph 107-2.
[0340] In some embodiments, in a step 2306, the Al Playbook Agent 110 calls the foundational model 109 grounded and fine-tuned with the relevant product knowledge graph 107-2, the user knowledge graph 107-1, and the 1stparty vector database, and sends the user query to said foundational model 109.
[0341] In some embodiments, in a step 2308, the Al Playbook Agent 110 receives the hyper personalized response from the grounded and fine-tuned foundational model 109.
[0342] In some embodiments, in step 2310, the Al Playbook Agent 110 sends the hyper personalized response to an action framework 114, where the action framework 114 uses an ontology service to determine and provide details of relevant external services and product customizations to the Al Playbook Agent 110 based on the output of the grounded and fine-tuned foundational model 109.In some embodiments, in step 2312, the Al Playbook Agent 110 updates the hyper personalized response with details of relevant product customizations and external services. In some embodiments, the details of the external services may also include links to book said external services on the websites or platforms of the external service providers.
[0343] In some embodiments, in step 2314, the Al Playbook Agent 110 send the updated hyper personalized response to the Omni Channel Integration Sub-system 112.
[0344] Although, the steps in the above embodiment of the method 2300 are shown to be performed sequentially. In some other embodiments, the steps of the method can be performed in any sequence, in parts, parallelly, in a batch process, or in any other synchronous or asynchronous format that would yield the end results as the embodiment of the method 2300 discussed above.
[0345] Omni Channel Integration Sub-system
[0346] FIG. 24 is a schematic diagram representing an omni-channel integration subsystem 112. In some embodiments, the omni-channel integration sub-system 112 ensures integration of the system across different platforms and channels, including a website, a mobile application, a desktop application, and an in-store device.
[0347] In some embodiments, the omni-channel integration sub-system 112 comprises a LLM Response Analyzer module 112a. The LLM Response Analyzer Module 112a uses semantic application programming interfaces (APIs) for semantic and ontologybased translation or codification of the LLM response received from the Al Playbook Agent 110. In said embodiments, the LLM Response Analyzer Module 112a converts the text response received from the foundational model 109 into text to code objects and API endpoints.
[0348] In some embodiments, the omni-channel integration sub-system 112 comprises a post response codifier module 112b. The post response codifier module 112b receives the text to code objects and API endpoints outputted by the LLM Response Analyzermodule 112a. The post response codifier module 112b executes APIs that can be either external or internal to the system 100. In some embodiments, the post response codifier module 112b runs the service booking workflow in the ecosystem partner’s infrastructure via the action framework 114 (refer to FIG. 1).
[0349] In some embodiments, the omni-channel integration sub-system 112 comprises a consolidate and summarize module 112c. The consolidate and summarize module 112c consolidates and summarizes the LLM response received from the Al Playbook Agent 110 using an Al language API.
[0350] In some embodiments, the omni-channel integration sub-system 112 comprises a render module 112d. The render module 112d receives the consolidated and summarized LLM response from the consolidate and summarize module 112c. The render module 112d renders the response using User Interface (UI) widgets or user device application(s) 18 on the user device 10. In some embodiments, the UI widgets or user device application 18 includes a chat widget to show text data to the user. In some embodiments, the UI widgets or user device application 18 includes a map widget or user device applications 18 to show geolocation data to the user. In some embodiments, the UI widgets or user device application 18 includes a chart widget or user device applications 18 to show numerical insight data to the user.
[0351] In some embodiments, the omni-channel integration sub-system 112 comprises a regime manager agent module 112e. The regime manager agent module 112e manages the interaction and communication of data between the LLM Response Analyzer module 112a, Post response codifier module 112b, Consolidate and summarize module 112c, and Render Module 112d.
[0352] Action Framework
[0353] In addition to working with the Al playbook agent 110 (FIG. 23), as noted above, the post response codifier module 112b runs the service booking workflow in the ecosystem partner’s infrastructure via the action framework 114. In some embodiments, the action framework 114 comprises external action APIs that allow theomni-channel integration sub-system 112 to communicate with external service providers 24 such as Salons, Therapy Centers, etc.
[0354] Semantic Service
[0355] In some embodiments, the system 100 is configured with a semantic service to offer product personalization or related external services, comprising: advisory to a sales agent on what all products to sell next to a given user (buyer), advisory to a marketing agent on what marketing campaign to run for specific user types (usersegments), advisory to an Al agent preparing a regime plan across a day several days, including consumption of product or combination of several related products and services, advisory regarding product, such as, a real-time human language-based audio, video or text advisory on relevance of a product in accordance with a user need, advisory on how to use the product, product purchase reminders, and automated purchase assistance with ecommerce and retail stores in the partner network, advisory on relevant services, including services from business partner automatically scheduling and reserving a consultation or series of consultations in a logical sequence, running enterprise workflows to enable certain actions, and automated product, service reminders for the bookings or product consumption as per the product consumption plan.
[0356] User Interaction Example
[0357] FIG. 25 is a schematic diagram representing flow of information through the system 100 when the user engages the system 100 with a query in accordance with the present invention.
[0358] In some embodiments, a registered user, “User 1”, inputs information 2502, in the user device application 18. In some embodiments, the information 2502 can be a chat prompt, with a question “Which hair colour would be best for my hair?”
[0359] In some embodiments, the information 2502 is passed to the Al Playbook Agent 110 via the Omni Channel Integration Sub-system 112 from the User Device Application 18. In some embodiments, the Al Playbook Agent 110 processes theinformation 2502. In some embodiments, the Al Playbook Agent 110 deconstructs the information 2502 to identify key concepts in the prompt: “Userl, Hair, Colour”.
[0360] In some embodiments, the Al Playbook Agent 110 uses the deconstructed information 2502 to identify the ontology of products and sub category for the ontology of products which is most closely related to the user input. For example, in this case, the ontology is selected as ‘Beauty’ and Sub-Category as ‘Hair Colours’.
[0361] In some embodiments, the Al playbook agent 110 also identifies the relevant user knowledge graph 107-1 and product knowledge graphs 107-2 that need to be processed by the grounding and fine-tuning foundation model module 108 and sends a reference to said user knowledge graph 107-1 and product knowledge graphs 107-2 as information 2504 to the grounding and fine-tuning foundation model module 108. In some embodiments, the information 2504 can include
[0362] “Reference to:
[0363] User 1 Knowledge Graph
[0364] Product 1 Knowledge Graph
[0365] Product 2 Knowledge Graph ”
[0366] In some embodiments, the grounding and fine-tuning foundation model module 108, uses the information 2504 to retrieve the relevant user knowledge graph 107-1 and product knowledge graphs 107-2 from the data storage medium 16 and sends information 2506 from said user knowledge graph 107-1 and product knowledge graphs 107-2 to the foundation model 109 for grounding and fine-tuning the foundation model 109. In some embodiments, the information 2506 may include:
[0367] “Note the following Information:
[0368] Userl ’s Name: Neha
[0369] Gender: Female
[0370] Hair Type: Dry
[0371] Location: South Delhi
[0372] Fav. Influencer: MeghaProduct 1: ABC Colour
[0373] Best For Hair Type: Oily
[0374] Discount Available: 10%
[0375] Megha ’s Video Clips on Product 1
[0376] Product 2: XYZ Colour
[0377] Best For Hair Type: Dry
[0378] Discount Available: 5%>
[0379] Megha ’s Video Clips on Product 2 ”
[0380] In some embodiments, the Al playbook agent 110, after grounding and fine-tuning the foundation model 109, sends information 2508 comprising the user query to the grounded and fine-tuned the foundation model 109. In some embodiments, the information 2508 includes:
[0381] “Respond to the following query of User 1 based on the noted information: Which hair colour will be best for my hair? ”
[0382] In some embodiments, the foundational model processes the information 2508 and produces information 2510, which is sent to the action framework 140 via the Al playbook agent 110. In some embodiments, the information 2510 includes:
[0383] “XYZ Colour will work best for your Dry Hair, Neha. It is available at a discount of 5%> on our ecommerce platform.
[0384] Also, here is a list of video clip of your favorite influencer, Megha explaining the usage and benefits of XYZ Colour:
[0385] [Video Clip 1], [Video Clip 2], [Video Clip 3], and [video clip 4]”
[0386] In some embodiments, the Al playbook agent 110 also sends a reference to relevant 1stParty Vector Database(s) to the action framework 114, via information 2512. The data storage medium 16 may contain many 1stParty Vector Databases, that may relate to many different types of curated information provided by themanufacturer / seller of the products on the ecommerce platform linked to the system 100. For example, a 1stParty Vector Database may include, links to product page of each product on the ecommerce platform, another 1stParty Vector Database may include information of 3rdParty Service Partners registered with the ecommerce platform. In some embodiments, the information 2512 includes:
[0387] “Reference to:
[0388] 1st Party Vector Database(s) ”
[0389] In some embodiments, the action framework 114 processes the information 2510 and information 2512 and adds relevant links to product pages and information of 3rdParty Service Partners to the information 2510 from the 1stParty Vector Database(s), and sends it as information 2514 via the omni-channel integration sub-system 112 to the user device application 18 as a response to the user query (information 2502). In some embodiments, information 2514 includes:
[0390] “XYZ Colour will work best for your Dry Hair, Neha. It is available ata discount of 5° / o on our ecommerce platform: Link to Buy: XYZ Hair Colour.
[0391] You can get this colour applied on your hair by professionals in our partner saloon near you in South Delhi: Link to Book Appointment
[0392] Also, here is a list of video clips of your favorite influencer, Megha explaining the usage and benefits of XYZ Colour
[0393] [Video Clip 1], [Video Clip 2], [Video Clip 3], and [video clip 4]”
[0394] This information flow 2500 repeats if the user sends a subsequent question / query via the user device application 18.
[0395] In some embodiments, the system 100 is capable of handling a large number of such user queries simultaneously. In some embodiments, the system 100 can provide the above-described automated user-specific hyper personalized product and service advisory to several hundred or thousand users of the ecommerce platform linked to the system 100 simultaneously, thereby significantly improving the quality of product and service advisory for all customers of the ecommerce platform.Comparative analysis of Example Responses
[0396] The following sections provide exemplary use cases for the multi-modal knowledge synthesizer system 100. For a comparative analysis of the improvements provided by the system 100 over the prior-art, FIGS. 26A-26B show a conversation between a user Jane and prior art Al systems in a chat box, i.e., user device application 18, and FIGS. 27A-27B show the conversation between the user Jane, and the present multi-modal knowledge synthesizer system 100 in the chat box, i.e., user device application 18. We note that the chat box is just one exemplary format of conversation system that the present multi-modal knowledge synthesizer system 100 can be adapted for. The present multi-modal knowledge synthesizer system 100 can be adapted for several other multimodal communication systems, including text, voice, audio-video, etc.
[0397] Referring to FIG. 26A (Prior- Art), as can be seen, the User queries “Suggest me a good skin product? ”, in response to which a Level 1 Prior Art Al agent responds “Hello, Jane, I suggest you these discounted conditioners and moisturizers. ” This response is very generic and lacks any personalization at all.
[0398] Referring to FIG. 26B (Prior- Art), as can be seen, in response to the same user query, a Level 2 Prior Art Al Agent responds “Hello, Jane, given your purchase patterns, I understand you are looking for skin moisturizer for daily use. ” This response is largely generic, with some level of customization based on the user’s purchase patterns is made.
[0399] Referring to FIG. 27 A, the multi-modal knowledge synthesizer system 100 of the present application, responds to the same query as “Hello, Jane, I understand that you are looking for skin moisturizer that can help you improve your extremely dry skin condition, and we are offering you a 5% discount on it. We understand that you have limited time during the day, and you have tropical summers coming soon! Interestingly, you are close to one of our highly rated salons, we recommend you take a skin cleansing therapy at least once in coming weeks.Below are few videos to explain product usage: video slice a, video slice x, video slice y, video slice z ” As can be seen this response is hyper personalized and more adept to the needs and requirements of the user, Jane. Additionally, the response recommends some videos that explains the usage of the product.
[0400] Referring to FIG. 27B, the hyper personalization in the response of the multimodal knowledge synthesizer system 100 is explained. For example, identifying that the user is looking for a “skin moisturizer” requires a deep product affinity model, and time series data indexing. Identifying that the user has “extremely dry skin”, requires knowledge of the bio-signals received from loT sensors from user’s devices, such as Apple Watch. Providing a dynamic discount offer, such as the “5% discount” requires integration of promotions and offer mix planning Identifying that the weather is changing, for example “tropical summers” are coming, requires data from weather sensors or weather information and location sensors of the user devices, it also requires information about advanced feature engineering, seasonality, and SKU propensity. Recommending “highly rated salons” requires ecosystem integration, expanded catalogs i.e., both First Party and Partner catalogue of services or products. Providing advisory on recommended therapy, such as “skin cleansing therapy” requires machine learning based NBA advisory connected to extended catalogue. In addition, the video slice recommendation provides relevant video slices from the analyzed video data by the video content parser 300. Further, generating such hyper personalized response requires channel specific templates, for example, the response in a chat box may vary from an email response, or an audio response, or an Al Avatar response, etc. Additionally, the time of day, date, and location of the user also play a crucial role in creating the response. Further, it also requires the LLM and Gen Al to adapt to the brand tone.
[0401] Practical Applications and Advantages
[0402] The present invention introduces an intelligent multimodal knowledge synthesizer that overcomes the limitations of prior methods by integrating diverse data sources into a unified knowledge graph. The system of the present invention employsadvanced Al techniques, including natural language processing, Generative Al, multimodel video analysis, to provide a holistic understanding of user preferences and product usage.
[0403] The invention disclosed in the present application has an ability to fuse multimodal data — text, video, loT, and user feedback — into a dynamic, interactive experience. Also, in the present application, a knowledge synthesizer uses a dynamic advanced statistical technique during knowledge mixing stage. Further, in the present application, an Al Playbook (AIP) agent drives personalized, human-like conversations with users. A high level of personalization is achieved by leveraging the rich, contextual knowledge stored in the product knowledge graph, which is continuously updated with new data. Furthermore, the present application discloses a video content parser module that can extract valuable insights from videos without relying on sound or specific languages, making it a useful tool for global applications, i.e., in regions with different languages and cultures. Moreover, in the present application, an omnichannel integration sub-system ensures that users have a consistent experience across platforms, such as mobiles, laptops, desktops, in-store kiosks.
[0404] The invention disclosed in the present application transforms the way customers engage with businesses, offering a scalable, Al-driven solution that enhances sales and deepens user engagement across multiple channels. Some key features of the invention disclosed in the present application includes features, such as, robust multimodal data fusion that effectively combines diverse data sources for a holistic understanding of products and user preferences, intelligent multimedia content analysis that extracts valuable insights from videos, audio files, including videos with no sound or content in different languages, personalized user interaction that creates engaging and informative conversations tailored to individual user needs and preferences, and enhanced sales and user engagement that boosts product sales and strengthens customer relationships through interactive and informative experiences. Also, the system of the present invention can provide automated user-specific hyper personalized product and service advisory to several hundred or thousands of users of the ecommerce platformlinked to the system simultaneously, thereby significantly improving the quality of product and service advisory for all customers of the ecommerce platform linked to the system.
[0405] The invention disclosed in the present application represents a significant advancement in Al-driven sales and marketing, offering businesses a powerful tool to improve customer satisfaction and drive revenue growth.
[0406] The term “about” is intended to include the degree of error associated with measurement of the particular quantity and / or manufacturing tolerances based upon the equipment available at the time of filing the application.
[0407] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the present disclosure. As used herein, the singular forms “a”, “an” and “the” are intended to include the plural forms as well, unless the context clearly indicates otherwise. It will be further understood that the terms “comprises” and / or “comprising,” when used in this specification, specify the presence of stated features, integers, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, element, components, and / or groups thereof.
[0408] Those of skill in the art will appreciate that various example embodiments are shown and described herein, each having certain features in the particular embodiments, but the present disclosure is not thus limited. Rather, the present disclosure can be modified to incorporate any number of variations, alterations, substitutions, combinations, sub-combinations, or equivalent arrangements not heretofore described, but which are commensurate with the scope of the present disclosure. Additionally, while various embodiments of the present disclosure have been described, it is to be understood that aspects of the present disclosure may include only some of the described embodiments. Accordingly, the present disclosure is not to be seen as limited by the foregoing description, but is only limited by the scope of the appended claims.
Claims
We Claim:
1. An intelligent multimodal knowledge synthesizer system for producing an automated user specific hyper personalized product and service advisory for a user, the system comprising:a multimodal content parser module configured to extract and analyze data from different sources of information about a product or the user;a knowledge mixer module configured to mix the knowledge extracted by the multimodal content parser module from said different sources;a knowledge graph builder module configured to build a knowledge graph based on the knowledge produced by the knowledge mixer;a grounding and finetuning module that grounds and finetunes a foundational model using one or more knowledge graphs; andan Al playbook agent module that extracts relevant knowledge from the grounded and finetuned foundational model in response to a user query to generate personalized response for the user.
2. The intelligent multimodal knowledge synthesizer system as claimed in claim 1, wherein the system comprises a content crawler that crawls content about the product or the user from online repositories, data sources, forums, brochures, catalogs and social media platforms and inputs the crawled content comprising at least of video, image, audio or text data into the multimodal content parser.
3. The intelligent multimodal knowledge synthesizer system as claimed in claim 1, wherein the multimodal content parser module comprisesan image and video content parser, which includes:a content slicer that fragments the input image or video into video fragments of different sizes,at least one sliding window analyzer that analyzes at least one video fragment at a time, wherein the input image is treated as a single video fragment for analysis by the sliding window analyzer; anda content knowledge module that stores the knowledge of each video fragment.
4. The intelligent multimodal knowledge synthesizer system as claimed in claim 1, wherein the multimodal content parser module comprises:an audio content parser, which includes:a content slicer that fragments the input audio into audio fragments of different sizes,at least one sliding window analyzer that analyzes at least one audio fragment at a time to analyze the audio fragment; anda content knowledge module that stores the knowledge of each audio fragment.
5. The intelligent multimodal knowledge synthesizer system as claimed in claim 1, wherein the multimodal content parser module comprises:a text content parser; anda content knowledge module that stores the knowledge of the text content.
6. The intelligent multimodal knowledge synthesizer system as claimed in claim 1, wherein the knowledge mixer module comprises:a knowledge vectors construction module that constructs knowledge vectors from user data;a tensor creation module that arranges each of the knowledge vectors into a multi-dimensional tensor;a knowledge vector clustering module that groups the knowledge vectors into one or more clusters;a context vectorization module that encodes a user’s query’s context into a query context vector;a cluster selection module that selects one or more most relevant clusters based on the query context vector; andan attention weight determination module, that determines attention weights for each of the knowledge vectors in the one or more most relevant clusters.
7. The intelligent multimodal knowledge synthesizer system as claimed in claim 6, wherein the attention weights determination module determines the attention weights by computing a standard scaled dot product between the query context vector and each of the knowledge vectors of the one or more most relevant clusters, which are further normalized using SoftMax function to generate final weights.
8. The intelligent multimodal knowledge synthesizer system as claimed in claim 6, wherein:the cluster selection module slices the multi-dimensional tensor to determine one or more relevant tensor slices based on the dimensions determined from the user’s profile or the user’s query.
9. The intelligent multimodal knowledge synthesizer system as claimed in claim 6, wherein an aggregated knowledge vector module calculates an aggregated knowledge vector by summing all the weighted knowledge vectors related to an entity.
10. The intelligent multimodal knowledge synthesizer system as claimed in claim 6, wherein the knowledge vector clustering module:selects a clustering algorithm;computes similarity between different knowledge vectors;applies a clustering algorithm to partition knowledge vectors into one or more clusters;characterizes and labels the one or more clusters.
11. The intelligent multimodal knowledge synthesizer system as claimed in claim 1, wherein the knowledge graph builder module creates a detailed relationalknowledge graph using the information received from the knowledge mixer module.
12. The intelligent multimodal knowledge synthesizer system as claimed in claim 1, wherein the grounding and finetuning foundational model module grounds the foundational model using the data of the knowledge graph and first party files from the embedded vector database, and industry playbook templates from a reinforcement, and feedback and training application, wherein industry templates are text configuration files, which are used to refine the system context, and help in effective response formulation or creating / executing a relevant action plan.
13. A method for producing an automated user specific hyper personalized product advisory for a user, the method comprising:extracting and analyzing data from different sources of information about a product or the user by the multimodal content parser module;mixing the knowledge extracted by the multimodal content parser module by the knowledge mixer module;building a knowledge graph based on the knowledge produced by the knowledge mixer, by the knowledge graph builder module;grounding and finetuning the foundational model using one or more knowledge graphs, by the grounding and finetuning module; andextracting relevant knowledge from the grounded and finetuned foundational model in response to a user query and generating a personalized response for the user, by the Al playbook agent.
14. The method as claimed in claim 13, further comprises:crawling content about the product or the user from online repositories, data sources, forums, brochures, catalogs and social media platforms and inputting the crawled content comprising video, image, audio or text data into the multimodal content parser.
15. The method as claimed in claim 13, wherein the extracting and analyzing comprises:parsing an image and video content which includes:slicing an input image or video into video fragments of different sizes; analyzing at least one video fragment at a time, wherein the input image is treated as a single video fragment for the analysis; andstoring knowledge of each video fragment.
16. The method as claimed in claim 13, wherein the extracting and analyzing comprises:parsing an audio content, which includes:slicing the input audio into audio fragments of different sizes; analyzing at least one audio fragment at a time; andstoring the knowledge of each audio fragment.
17. The method as claimed in claim 13, wherein extracting and analyzing comprises:parsing a text content; andstoring the knowledge of the text content.
18. The method as claimed in claim 13, wherein knowledge mixing comprises:constructing knowledge vectors from user data;arranging each of the knowledge vectors into a multi-dimensional tensor; clustering the knowledge vectors in the relevant tensor slice into one or more clusters;encoding a user’s query’s context into a query context vector;selecting one or more most relevant clusters based on the query context vector; anddetermining attention weights for each of the knowledge vectors in the one or more most relevant clusters.
19. The method as claimed in claim 18, wherein the determining of the attention weights comprises:computing standard scaled dot product between the query context vector and each of the knowledge vectors of the one or more most relevant clusters; and normalizing using SoftMax function to generate final weights.
20. The method as claimed in claim 18, wherein the method further comprises slicing the multi-dimensional tensor to determine one or more relevant tensor slices based on the dimensions determined from the user’s profile or the user’s query.
21. The method as claimed in claim 18, wherein the method further comprises:calculating the aggregated knowledge vector by combining all the weighted knowledge vectors related to an entity.
22. The method as claimed in claim 18, wherein the clustering of knowledge vectors comprises:selecting a clustering algorithm;computing similarity among different knowledge vectors;applying a clustering algorithm to partition knowledge vectors into one or more clusters;characterizing and labeling the one or more clusters.
23. The method as claimed in claim 13, wherein building the knowledge graph creates a detailed relational knowledge graph using the information received from the knowledge mixer module.
24. The method as claimed in claim 13, wherein the grounding and finetuning comprises:using the data of the knowledge graph and first party files from the embedded vector database, and industry playbook templates from a reinforcement, andfeedback and training application, wherein industry templates are text configuration files, which are used to refine the system context, and help in effective response formulation or creating / executing a relevant action plan.