System and method for answering a natural language user query in an industrial environment
The system addresses the limitations of LLMs in industrial environments by using a domain-specific first LLM and a knowledge base, combined with a public LLM, to provide accurate and efficient responses to natural language queries while maintaining data confidentiality.
Patent Information
- Application Number
- PCT/EP2024/074558
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-25
- Filing Date
- 2024-09-03
- Publication Date
- 2025-10-30
AI Technical Summary
Existing large language models (LLMs) face challenges in industrial environments due to limited access to high-quality data, computational resource constraints, and data privacy concerns, necessitating domain-specific fine-tuning, which is time-consuming and resource-intensive, and lack domain-specific knowledge, leading to inaccurate responses to queries outside their training dataset.
A system utilizing a first LLM trained on industrial domain data and a knowledge base, combined with a second public LLM, processes natural language queries by generating domain-specific prompts, ensuring accurate responses while maintaining data confidentiality, and iteratively refining the first LLM's parameters for efficient domain-specific training.
Enables accurate and efficient answering of natural language queries in industrial environments by integrating domain-specific knowledge with open-world knowledge, ensuring privacy and reducing the need for extensive data and computational resources.
Smart Images

Figure EP2024074558_30102025_PF_FP_ABST
Abstract
Description
[0001] SYSTEM AND METHOD FOR ANSWERING A NATURAL LANGUAGE USER QUERY IN AN
[0002] INDUSTRIAL ENVIRONMENT
[0003] CROSS-REFERENCE TO RELATED PATENT APPLICATIONS
[0004] This application claims priority to and the benefit of patent application number 202441032988 titled “SYSTEM AND METHOD FOR ANSWERING A NATURAL LANGUAGE USER QUERY IN AN INDUSTRIAL ENVIRONMENT”, filed in the Indian Patent Office on April 25, 2024. The specification of the above referenced patent application is incorporated herein by reference in its entirety.
[0005] The present invention relates to a field of data retrieval systems and more particularly relates to a system and method for answering a natural language user query in an industrial environment.
[0006] In an industrial environment, numerous assets are employed for performing different operations. The assets may include mechanical systems, electromechanical systems, electronic systems, building facility and so on. Several predictive analytics and data analytics techniques are used to monitor condition, behavior and performance of assets in the industrial environment. However, the interpretation of data output from such systems is a challenge. Even with level of automation and intelligent systems that are present, the industry still relies heavily on the operator’s domain know how to manipulate and understand the process. Interestingly, even when all the data that is required is available in the systems already employed, it is still a challenge to retrieve specific information desired for a particular event at a given time with current approaches. For example, consider from the point of view of the plant operators, maintenance engineers, supervisors, or factory personnel who would like to interrogate the system from time to time to get insight on the state of the system, get specific answers regarding the problems being faced, specific KPI values at that time, causal analysis of the anomalies already happened, potential solutions to the problems already faced. In addition, it is also seen that operators / factory managers / building managers would like to know additional information on the equipment deployed ranging from model make, type, age, ideal operating conditions to complete process in which the equipment are deployed, operated in, potential weaknesses and inefficiencies in the process. This can further be extended to examining the inventories, supply chain of raw materials, storage, logistics and others. The lack of systems that can retrieve and provide requested information in an industrial environment brings several challenges to the plant operators, such as uninformed or incorrect decision-making, relying heavily on domain knowledge, incorrect interpretation of industrial system results, etc. In a nutshell, data retrieval from industrial systems in a manner that is understandable to the user is a challenge. Data retrieval refers to the process of querying one or more data sources for desired information in some manner. In conventional data retrieval based on natural language processing (Natural Language Processing, NLP) technology, queries are typically expressed by natural language questions of users, algorithm models determine intent, direction and goal of the retrieval by performing semantic analysis, such as regular expressions, keyword extraction, named entity recognition, etc., on the natural language questions, then obtain the required information from data sources including databases, file systems, web pages, emails, etc., through various computer program interfaces, and finally package the results back to the users. With the development of Al technology, a large language model (Large Language Model, LLM) technology is raised, and by constructing a neural network model with extremely large parameters (GPT 3 about 1750 hundred million), unsupervised pre-training is performed on a huge corpus (GPT 3 pre-training corpus about 45 TB), and the nature is that human knowledge and modes expressed in text, codes and the like are parameterized into the neural network model, so that the large language model is naturally provided with understanding of general human knowledge and can be directly used for general knowledge retrieval and generalization.
[0007] LLMs are typically trained on a large language model data set, hence requiring huge computing power and availability of data. However, these availability of both of these crucial elements is uncertain due to multiple reasons like limited access to high-quality data, computational resource constraints, and potential data privacy concerns. These factors can limit the ability to effectively train LLMs and hence require domain-specific fine-tuning. Additionally, even with access to sufficient data and computing resources, training LLMs can still be a time-consuming and resource-intensive process, requiring significant expertise in natural language processing and deep learning techniques. In certain cases, data itself cannot be shared outside the organization, to use external LLM API’s due to privacy constraints.
[0008] Creating an in-house Al model that can perform a specific task with limited data while also exhibiting the behavior and style of a large language model (LLM) is a significant technical challenge for organizations. Another major challenge is that LLMs are typically trained on massive amounts of data, which allows them to learn the intricacies of natural language and generate human-like responses.
[0009] Furthermore, other challenges with existing public LLMs are that it is not available as APIs, absence of domain adaption technologies along with data confidentiality, and no run time production set-up while managing data traffic and load. Moreover, current LLMs are not capable of generating responses to data that is not available in the training dataset. In other words, current LLMs are unable to generate responses to prompts or situations that it has not seen before.
[0010] Organizations face a significant technical challenge in creating an in-house Al model that can exhibit the behavior and style of a large language model (LLM) while also performing a specific task with limited data. LLMs typically require huge computing power and access to a large language model data set, but these resources are often uncertain due to various factors such as limited access to high-quality data, computational resource constraints, and data privacy concerns. As a result, domain-specific fine-tuning may be necessary to effectively train LLMs.
[0011] Even with sufficient access to data and computing resources, training LLMs can be a timeconsuming and resource-intensive process, requiring significant expertise in natural language processing and deep learning techniques. In some cases, privacy constraints prevent the sharing of data outside the organization, which limits the ability to use external LLM APIs.
[0012] Furthermore, most LLMs out there have been trained on open world knowledge and has been proven to be reasonably good at answering generic questions. But these LLMs lack domain specific knowledge that are accumulated in industry specific knowledge silos because of their confidential nature. Therefore, there is a need to combine the open world knowledge along with domain specific know how to answer several of the queries, a plant operator might have.
[0013] In light of the above, there exists a need for answering a natural language user query in an industrial environment in an accurate manner combining confidential industrial asset information with domain knowledge and open-world knowledge as disclosed.
[0014] Therefore, it is an object of the present invention to provide a system, apparatus and method for answering a natural language user query in an industrial environment.
[0015] The object of the present invention is achieved by a method for answering a natural language user query in an industrial environment. The method comprises receiving the natural language query from the user. The natural language query pertaining to a particular domain of the industrial environment.
[0016] Throughout the present disclosure, the term “industrial environment” as used herein refers to a technical set-up with a plurality of assets such as a power plant, wind farm, power grid, manufacturing facility, process plants, buildings (residential or non-residential areas) and so on. Examples of an industrial environment may include a complex industrial set-up such as a manufacturing facility, process plants, storage facility, transportation. It will be appreciated that the industrial environment may refer to any vertical and / or domain in business. For example, different verticals treated as industrial environment for the purpose of this disclosure, may include but not limited to automobiles, textiles, every distribution, energy production, buildings, factories, and medical equipment. For the sake of simplicity and brevity of the invention, the invention is explained with respect to industrial environment. A person skilled in the art would understand that the concepts disclosed herein can be applied across multiple domains such as finance, marketing, legal, medicine, etc. Therefore, the claims appended herein shall be construed limiting to the industrial environment only.
[0017] Throughout the present disclosure, the term “assets” refer to any device, system, instrument or machinery manufactured or used in an industry that may be employed for performing an operation. In some cases, assets may also include any devices or instruments deployed or functioning in a non-industrial environment such as buildings. Example of assets include any machinery in a technical system or technical instal lation / facility such as motors, gears, bearings, shafts, switchgears, rotors, circuit breakers, protection devices, remote terminal units, transformers, reactors, disconnectors, gear-drive, gradient coils, magnet, chillers, radio frequency coils, appliances, electronic devices, chillers, pumps, heat exchangers, cooling towers, air compressors, boilers, fluid bed driers, coating machines, carbonation towers etc.
[0018] Throughout the present disclosure, the term “one or more sources” as used herein refers to databases from where information pertaining to the particular domain is extracted. In an example, the one or more sources are product catalogs, product manuals, industry standards, regulatory documents, patents, technical specifications, and reference materials. The sources can also be knowledge graphs for the domain, ontologies, historical databases, etc. The “one or more sources” also comprise sensors such as position sensors, rotary encoders, dynamometers, proximity sensors, current sensors, accelerometers, temperature sensors, acoustic sensors, voltage sensors associated with the assets in the industrial environment that provide data related to the one or more assets. The one or more sources also comprise Al model reports, RCA reports, anomaly reports, alarm reports pertaining to models deployed in the industrial environment.
[0019] Throughout the present disclosure, the term “natural language query” as user herein refers to a combination of linguistic, semantic, and pragmatic elements that together convey the user's information needs in a contextually rich manner. The natural language query can be understood as a type of search or input query that is expressed in ordinary human language, rather than using formalized syntax or specific query languages. It allows users to communicate with computers or systems using natural language, just like they would with other humans. In an example, the natural language query is input via user through a user interface. It is to be noted that the term “natural language query” and “query” have been interchangeably used hereinafter.
[0020] Throughout the present disclosure, the term “user” as used herein refers to a human interacting with the system for responses. The user may also refer to a virtual assistant or co-pilot capable of querying the system with natural language query.
[0021] Through the present disclosure, the term “prompts” refer to textual input provided to the model to generate responses or perform specific tasks, prompts consist of one or more sentences or passages of text that serve as input to the LLM. This input can include questions, statements, commands, or any other form of textual communication. The prompts convey the intent of the natural language query to the language model. In general, prompts may include contextual information relevant to the desired task or response. This could include background information, constraints, preferences, or other contextual cues that help guide the response of the LLM. In some examples, the prompts follow a specific format or structure depending on the requirements of the task or application. This could involve using certain keywords, phrases, or formatting conventions to convey instructions or constraints to the LLM. In some cases, the prompts may include examples or samples of the desired output to provide additional context or guidance to the model. This aids to clarify the user's expectations and improve the quality of the generated responses. In the context of LLMs, the prompts refer to as cues or instructions for the model to produce desired outputs based on the input provided, for example, the natural language query. Advantageously, the prompts are generated for accurately querying the first LLM and / or second LLM. It should be understood that the claims refer to a “first set of prompts”, a “second set of prompts”, a “third set of prompts”, a fourth set of prompts” which are used throughout the description at multiple occasions, the above definition applies to all. Also, the term “prompts” has been used hereinafter to refer collectively to “first set of prompts”, “second set of prompts”, “third set of prompts”, fourth set of prompts”. For brevity, the first set of prompts”, “second set of prompts”, “third set of prompts”, fourth set of prompts” are also referred to as “prompts” individually as per the context.
[0022] Throughout the present disclosure, the term “domain taxonomy” as used herein refers to a hierarchical classification system that organizes concepts, entities, or topics within a specific domain or field of knowledge. The domain taxonomy provides a structured framework for categorizing and organizing information based on relationships, similarities, and hierarchies. The domain taxonomy is organized in a hierarchical manner, with broader categories or concepts at the top level and increasingly specific subcategories or subtopics at lower levels. This hierarchical structure enables the system to navigate from general to more detailed levels of information. Furthermore, the domain taxonomy defines a classification scheme that categorizes entities or concepts into distinct classes or categories based on shared characteristics or attributes. Furthermore, the domain taxonomy capture relationships and dependencies between different classes or categories within the domain. Moreover, the domain taxonomy uses standardized terminology and naming conventions to ensure consistency and clarity in the classification of concepts. The domain taxonomy also comprises specific field of knowledge, industry, or subject area, reflecting the unique characteristics, concepts, and relationships within that domain. It may incorporate domain-specific terminology, concepts, and principles relevant to the subject area. It is to be noted that the domain taxonomy evolves over time to accommodate changes in the domain, new discoveries, or emerging trends with respect to the chosen domain. In specificity, the domain taxonomy is a hierarchical structure comprising keywords in the domain mapped to corresponding description pertaining to the domain. It is to be understood that the domain taxonomy is generated based on a requirement of the user or a use case for which the system is deployed for.
[0023] Through the present disclosure, the term “Large Language Model (LLM)” refers to a type of artificial intelligence model designed to understand and generate human-like text based on vast amounts of natural language data. These models are built using deep learning architectures, particularly transformer-based architectures. LLMs are trained on massive datasets containing billions or even trillions of words from diverse sources such as books, articles, websites, and other textual content. The large-scale training enables the model to learn complex patterns, structures, and nuances of human language. LLMs are typically built using deep learning architectures, particularly transformer architectures. Transformers employ self-attention mechanisms to process input sequences and capture long-range dependencies, enabling the model to generate coherent and contextually relevant text. LLMs undergo two main stages of training: pre-training and fine- tuning. During pre-training, the model is trained on a large corpus of text data using unsupervised learning techniques to learn general language patterns and semantics. In fine-tuning, the pretrained model is further optimized on domain-specific or task-specific datasets to adapt its knowledge and capabilities to specific applications. LLMs have the ability to generate human-like text based on given prompts or inputs. They can produce coherent paragraphs, articles, stories, code, or responses to questions by predicting the next words or tokens in the sequence based on the context provided. LLMs exhibit a strong understanding of context and semantics in natural language. They can infer meaning, resolve ambiguity, and generate text that is contextually relevant and coherent with the given input.
[0024] Throughout the present disclosure, the term “second LLM” as used herein refers to a public domain LLM. Some examples of the second LLM include ChatGPT, T5 (Text-to-Text-Transfer- Transformer), GPT-Neo, GPT-J, and GPT-NeoX, XLNet, Roberta - Robustly Optimized BERT Approach, DeBERT,, DistilBERT, GPT 3.5, GPT 4, GPT-vision, LLAMA-3, Mistral 7b instruct, Phi3, dolphin-mixtral, Llava, Gemini pro, gemini 1.5 and other state of the art LLMs.
[0025] Throughout the present disclosure, the term “first LLM” as used herein refers to an artificial intelligence model designed to understand and generate human-like text based on industrial domain data and one or more Al model deployed in the industrial environment. These models are built using deep learning architectures, particularly transformer-based architectures. The first LLM is trained on massive datasets acquired from the one or more data sources such as sensors, RCA reports, prediction models, one or more Al models, database including asset specification, predetermined knowledge graphs, images, videos, audios, or a combination thereof. The large- scale training enables the model to learn complex patterns, structures, and nuances of human language. The first LLM is built using deep learning architectures, particularly transformer architectures. Transformers employ self-attention mechanisms to process input sequences and capture long-range dependencies, enabling the model to generate coherent and contextually relevant text. The first LLM is an in-house Al model that can exhibit the behavior and style of a large language model (LLM) while also performing a specific task with limited data.
[0026] Throughout the present disclosure, the term “knowledge base” as used herein refers to a heterogeneous database comprising information pertaining to the domain and the industrial environment. The knowledge base is a centralized repository or database that contains domainspecific information, asset specific information, process- specific information, sensor-specific information, etc. The knowledge base also contains information relevant to the specific industry, sector, or domain in which the organization operates. This may include technical specifications, manufacturing processes, equipment manuals, safety procedures, regulatory requirements, and industry standards. The knowledge base also comprises information retrieved from one or more Al models deployed in the industrial environment. The knowledge base also comprises information extracted from documentation of performance parameters, efficiency parameters, anomalies, root cause analysis, prediction and resolution anomalies, etc. The knowledge base also comprises information extracted from knowledge graphs comprising information of the plant and its assets in a hierarchical manner. The knowledge base also comprises images, videos, audios etc. having information of the industrial environment.
[0027] The method comprises parsing the natural language query to determine one or more entities and an intent of the natural language query. In an embodiment, the method of determining an intent of the natural language query comprises parsing the natural language query to extract one or more entities and determine semantic and syntactic relationships therebetween. Further, the method comprises determining an intent of the natural language query based on a machine learning model trained on a historical database of one or more entities annotated with their corresponding intent. Advantageously, the objective of the correct intent identification of the query is to route the user query to the appropriate data sources so as to provide an accurate answer.
[0028] The method comprises generating a first set of prompts from the natural language query based on a defined taxonomy pertaining to the particular domain of the industrial environment and the determined intent of the user. The domain taxonomy is a hierarchical structure comprising keywords in the domain mapped to corresponding description pertaining to the domain.
[0029] In an embodiment, the method of generating prompts from the natural language query for querying the first LLM comprises retrieving information that is similar to the keywords in the natural language query from domain taxonomy. Further, the method comprises generating prompts for querying the first LLM based on the descriptions extracted from the domain taxonomy. Advantageously, the prompts are generated for accurately querying the first LLM and / or second LLM. Furthermore, the correct prompt generation eliminated the need for sending huge prompts to the first LLM and improves the efficiency of the response.
[0030] In an embodiment, the method of generating domain taxonomy comprises identifying a plurality of keywords from one or more sources pertaining to a particular domain. The method comprises extracting a textual description of each of the keywords from the one or more sources. The method comprises defining a relationship between the keywords and the textual description by annotating the keywords. The method comprises generating the domain taxonomy with identified keywords and corresponding textual description based on a defined hierarchy. Advantageously, the domain taxonomy will be utilized to get more open world information which will enable the system to make more richer inferences. Furthermore, the descriptions in the domain taxonomy also ensure accurate generations of prompts.
[0031] The method comprises determining if an answer to the natural language query is known to a first LLM based on the determined intent of the natural language query. The first LLM is specifically trained on the particular domain of the industrial environment and enriched with a knowledge base. The knowledge base comprises information from one or more data sources combined with domain knowledge from a second LLM. Advantageously, the system is capable of intelligently deciding the which workflow to be followed for accurately answering the natural language query.
[0032] In an embodiment, the method of generating the knowledge base pertaining to the particular domain comprises acquiring information from one or more data sources in the industrial environment. The one or more data sources comprises at least one of: sensors, RCA reports, prediction models, one or more Al models, database including asset specification, predetermined knowledge graphs, images, videos, audios, or a combination thereof. The method comprises generating domain taxonomy based on the extracted information. The domain taxonomy comprises a set of keywords and corresponding descriptions extracted from the acquired information. The method comprises generating a set of questions using the domain taxonomy based on a predefined template. The method comprises identifying a confidentiality level of the generated set of questions. The method comprises querying the second LLM using the generated set of questions that do not contain confidential entities. The method comprises triggering Al workflows in the one or more Al models for the set of questions that contain confidential entities. The method comprises generating the knowledge base by combining output from Al workflows and second LLM. Advantageously, the set of questions generated vary significantly and aims to include all potential queries such that the knowledge base is accurately enriched with correct information. Advantageously, the information in the knowledge base is stored in in a structured manner to facilitate easy navigation, retrieval, and access to information. Information is typically categorized into topics, sections, or modules based on the subject matter, making it easier for users to locate relevant content.
[0033] In an embodiment, the method of generating the first LLM comprises generating a set of questions from the knowledge base using the domain taxonomy. The method comprises querying the second LLM with the generated set of questions iteratively with different variations of the generated set of questions. The method comprises querying the first LLM with the generated set of questions iteratively with different variations of the generated set of questions. The method comprises iteratively calculating a distillation loss between the responses of the first LLM and the second LLM . The method comprises tuning one or more parameters of the first LLM such that the distillation loss decreases after each iteration. The method comprises generating the first LLM when the calculated distillation loss for each of the set of questions in below a predefined threshold. Advantageously, the technique for fine tuning ensures that the fine tuning of the first LLM is done on-premises and data need not be out to external sources (like ChatGPT API’s) mitigating privacy issues. Advantageously, the technique to generate the first LLM enables focused domain-specific training in an efficient manner and eliminating the need to use large corpus of data.
[0034] The method comprises querying the natural language query to the first LLM using the generated prompts, if it is determined that the answer is known to the first LLM. The method comprises generating a response to the natural language query based on an output of the first LLM.
[0035] In an embodiment, the response to the natural language query is output in at least one of the formats: textual, images, videos, reports, and a combination thereof. Advantageously, the system intelligently understands the intent of the user and structures a response most suitable for the user to interpret the response.
[0036] In an embodiment, the method further comprises comprising identifying a confidentiality level of the natural language query of the user based on the determined entities of the natural language query identified based on the comparison with a database comprising a plurality of keywords that are annotated as confidential.
[0037] In an embodiment, the method of answering the received natural language query comprises querying to the first LLM using the generated prompts, if it is identified that the natural language query comprises one or more entities marked as confidential in the database. Advantageously, in order to maintain confidentiality of customer data or plant data, the natural language query comprising confidential entities are only directed to the first LLM i.e. the local LLM enriched with knowledge base.
[0038] In an embodiment, the method of answering the received natural language query comprises executing one or more Al workflows in the industrial environment based on the intent of the natural language query, if it is determined that the answer to the natural language query is not known to the first LLM, and the natural language query comprises one or more entities marked as confidential in the database. The method comprises updating the knowledge base with output of the one or more Al workflows triggered. The method comprises generating a second set of prompts based on the domain taxonomy. The method comprises querying the natural language query to the first LLM enriched with the updated knowledge base using the second set of prompts. The method comprises generating a response to the natural language query based on an output of the first LLM enriched with the updated knowledge base. Advantageously, the system is capable of answering queries that are not present in the knowledge base or in the training database of the first LLM, thereby making sure that new queries are also answered accurately.
[0039] In an embodiment, the method of answering the received natural language query comprises querying to the second LLM using the generated prompts, if it is identified that the natural language query does not comprise one or more entities marked as confidential in the database, and the answer to the natural language query is not known to the second LLM. Advantageously, the confidentiality check ensures that the queries having non-confidential entities are only forwarded to the second LLM or the public LLM.
[0040] In an embodiment, the method of answering the received natural language query identified as comprising entities marked as both confidential and non-confidential comprises identifying one or more confidential entities in the natural language query. The method comprises generating a third set of prompts by replacing the confidential entities with general descriptions using the domain taxonomy, and a fourth set of prompts by retaining the confidential entities. The method comprises querying the second LLM with the third set of prompts and the first LLM with the fourth set of prompts. The method comprises structuring, by the processing unit, a response from the first LLM and the second LLM in order to generate a natural language response to the natural language query.
[0041] The object of the present invention is also achieved by an apparatus for answering a natural language user query in an industrial environment. The apparatus comprising one or more processing units and a memory unit communicatively coupled to the one or more processing units. The memory unit comprises a module stored in the form of machine-readable instructions executable by the one or more processing units, wherein the module is configured to perform method steps as mentioned above.
[0042] The object of the present invention is also achieved by a system for answering a natural language user query in an industrial environment. The system comprising one or more client devices configured for inputting the natural language query and outputting the answer to the natural language query, a first LLM trained specifically on a particular domain and enriched with a knowledge base, wherein the first LLM is a local LLM, the first LLM being communicatively coupled to the one or more client devices, a second LLM, wherein the second LLM is a public LLM, one or more data sources for acquiring information pertaining to one or more assets installed in the industrial environment, and an apparatus communicatively coupled to the one or more client devices. The apparatus is configured for answering a natural language user query in an industrial environment.
[0043] The object of the present invention is also achieved by a system for answering a natural language user query in an industrial environment. The system further comprises a plurality of first LLMs trained specifically on corresponding plurality of industrial domains, a plurality of knowledge bases generated based on specific domain knowledge, Al models, manuals, text documents for the plurality of industrial domain, the plurality of knowledge bases communicatively coupled to corresponding plurality of first LLMs, and a domain selection module configured for identifying a domain of the received natural language query, and selecting a suitable first LLM and corresponding knowledge base based on the identified domain of the received natural language query.
[0044] The object of the present invention is also achieved by a computer-program product having machine-readable instructions stored therein, which when executed by one or more processing units, cause the processing units to perform the method as abovementioned.
[0045] The object of the present invention is also achieved by computer-readable storage medium comprising instructions which, when executed by one or more processing units cause the one or more processing units to perform the method as abovementioned.
[0046] The above-mentioned attributes, features, and advantages of this invention and the manner of achieving them, will become more apparent and understandable (clear) with the following description of embodiments of the invention in conjunction with the corresponding drawings. The illustrated embodiments are intended to illustrate, but not limit the invention.
[0047] FIG 1 illustrates a block-diagram of a system for answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention;
[0048] FIG 2 illustrates a block-diagram of a system for answering a natural language user query in an industrial environment, in accordance with another embodiment of the present invention;
[0049] FIG 3 illustrates an apparatus answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention;
[0050] FIG 4 illustrates a block diagram of an architecture of the system of FIG 1 , in accordance with an embodiment of the present invention;
[0051] FIG 5 illustrates a block diagram of an architecture of the system of FIG 1 , in accordance with another embodiment of the present invention;
[0052] FIG 6 depicts a flowchart of a method for answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention;
[0053] FIG 7 illustrates a block diagram of a knowledge base with one or more data sources, in accordance with an embodiment of the present invention; FIG 8 depicts a flowchart of a method for generating the knowledge base pertaining to the particular domain, in accordance with an embodiment of the present invention;
[0054] FIG 9 illustrates a block diagram of an architecture for generating the knowledge base pertaining to the particular domain, in accordance with an embodiment of the present invention;
[0055] FIG 10 depicts a flowchart of a method for generating the first LLM, in accordance with an embodiment of the present invention;
[0056] FIG 11 illustrates a block diagram of an architecture for generating the first LLM, in accordance with an embodiment of the present invention;
[0057] FIG 12 illustrates a block diagram of another architecture for generating the first LLM, in accordance with another embodiment of the present invention;
[0058] FIG 13 illustrates a block diagram representing a method for answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention;
[0059] FIG 14 illustrates an exemplary method flow for answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention;
[0060] FIG 15 illustrates an exemplary method flow for answering a natural language user query in an industrial environment, in accordance with another embodiment of the present invention;
[0061] FIG 16 illustrates an exemplary method flow for answering a natural language user query in an industrial environment, in accordance with another embodiment of the present invention; and
[0062] FIG 17 illustrates a Graphical User Interface (GUI) of the user device for answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention.
[0063] Hereinafter, embodiments for carrying out the present invention are described in detail. The various embodiments are described with reference to the drawings, wherein like reference numerals are used to refer to like elements throughout. In the following description, for purpose of explanation, numerous specific details are set forth in order to provide a thorough understanding of one or more embodiments. It may be evident that such embodiments may be practiced without these specific details.
[0064] Description of embodiments
[0065] Referring to FIG. 1 , illustrated is a block-diagram of a system 100 for answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention. The system comprises a first large language model (LLM) 102, a second large language model (LLM) 104, an apparatus 106, a client device 110, a knowledge base 112, and one or more sources 114-1 to 114-N. The first LLM 102, the second LLM 104, the apparatus 106, user device 110, the knowledge base 112, and the one or more sources 114-1 to 114-N are communicatively coupled to each other via a communication network 108.
[0066] The industrial environment is a technical set-up with a plurality of assets such as a power plant, wind farm, power grid, manufacturing facility, process plants, buildings (residential or non- residential areas) and so on. Examples of an industrial environment may include a complex industrial set-up such as a manufacturing facility, process plants, storage facility, transportation. It will be appreciated that the industrial environment may refer to any vertical and / or domain in business. For example, different verticals treated as industrial environment for the purpose of this disclosure, may include but not limited to automobiles, textiles, every distribution, energy production, buildings, factories, and medical equipment.
[0067] The assets (not shown) are any device, system, instrument or machinery manufactured or used in an industry that may be employed for performing an operation. In some cases, assets may also include any devices or instruments deployed or functioning in a non-industrial environment such as buildings. Example of assets include any machinery in a technical system or technical installation / facility such as motors, gears, bearings, shafts, switchgears, rotors, circuit breakers, protection devices, remote terminal units, transformers, reactors, disconnectors, gear-drive, gradient coils, magnet, chillers, radio frequency coils, appliances, electronic devices, chillers, pumps, heat exchangers, cooling towers, air compressors, boilers, fluid bed driers, coating machines, carbonation towers etc.
[0068] The term “Large Language Model (LLM)” 102 and 104 refers to a type of artificial intelligence model designed to understand and generate human-like text based on vast amounts of natural language data. These models are built using deep learning architectures, particularly transformerbased architectures. LLMs are trained on massive datasets containing billions or even trillions of words from diverse sources such as books, articles, websites, and other textual content. The large-scale training enables the model to learn complex patterns, structures, and nuances of human language. LLMs are typically built using deep learning architectures, particularly transformer architectures. Transformers employ self-attention mechanisms to process input sequences and capture long-range dependencies, enabling the model to generate coherent and contextually relevant text. LLMs undergo two main stages of training: pre-training and fine-tuning. During pre- training, the model is trained on a large corpus of text data using unsupervised learning techniques to learn general language patterns and semantics. In fine-tuning, the pre-trained model is further optimized on domain-specific or task-specific datasets to adapt its knowledge and capabilities to specific applications. LLMs have the ability to generate human-like text based on given prompts or inputs. They can produce coherent paragraphs, articles, stories, code, or responses to questions by predicting the next words or tokens in the sequence based on the context provided. LLMs exhibit a strong understanding of context and semantics in natural language. They can infer meaning, resolve ambiguity, and generate text that is contextually relevant and coherent with the given input.
[0069] The second LLM 104 is a public domain LLM. Some examples of the second LLM include ChatGPT, T5 (Text-to-Text-Transfer-Transformer), GPT-Neo, GPT-J, and GPT-NeoX, XLNet, Roberta - Robustly Optimized BERT Approach, DeBERT,, DistilBERT, etc.
[0070] The first LLM 102 is an artificial intelligence model designed to understand and generate humanlike text based on industrial domain data and one or more Al model deployed in the industrial environment. These models are built using deep learning architectures, particularly transformerbased architectures. The first LLM 102 is trained on massive datasets acquired from the one or more data sources such as sensors, RCA reports, prediction models, one or more Al models, database including asset specification, predetermined knowledge graphs, images, videos, audios, or a combination thereof. The large-scale training enables the model to learn complex patterns, structures, and nuances of human language. The first LLM 102 is built using deep learning architectures, particularly transformer architectures.. The first LLM 102 is an in-house Al model that can exhibit the behavior and style of a large language model (LLM) while also performing a specific task with limited data.
[0071] The knowledge base 112 is a heterogeneous database comprising information pertaining to the domain and the industrial environment. The knowledge base 112 is a centralized repository or database that contains domain-specific information, asset specific information, process- specific information, sensor-specific information, etc. The knowledge base 112 also contains information relevant to the specific industry, sector, or domain in which the organization operates. This may include technical specifications, manufacturing processes, equipment manuals, safety procedures, regulatory requirements, and industry standards. The knowledge base 112 also comprises information retrieved from one or more Al models deployed in the industrial environment. The knowledge base 112 also comprises information extracted from documentation of performance parameters, efficiency parameters, anomalies, root cause analysis, prediction and resolution anomalies, etc. The knowledge base 112 also comprises information extracted from knowledge graphs comprising information of the plant and its assets in a hierarchical manner. The knowledge base 112 also comprises images, videos, audios etc. having information of the industrial environment.
[0072] The term “one or more sources” 114-1 to 114-N as used herein refers to databases from where information pertaining to the particular domain is extracted. In an example, the one or more sources are product catalogs, product manuals, industry standards, regulatory documents, patents, technical specifications, and reference materials. The sources 114-1 to 114-N can also be knowledge graphs for the domain, ontologies, historical databases, etc. The “one or more sources” 114-1 to 114-N also comprise sensors such as position sensors, rotary encoders, dynamometers, proximity sensors, current sensors, accelerometers, temperature sensors, acoustic sensors, voltage sensors associated with the assets in the industrial environment that provide data related to the one or more assets. The one or more sources 114-1 to 114-N also comprise Al model reports, RCA reports, anomaly reports, alarm reports pertaining to models deployed in the industrial environment.
[0073] The client device 110 provides a user interface for the user to interact with the system. Nonlimiting examples of client devices 110 include, personal computers, workstations, personal digital assistants, human machine interfaces. The client device 110 may enable the user to input one or more queries through a web-based interface.
[0074] The present invention provides the system 100 capable of answering natural language query for an industrial environment in an efficient and accurate manner. The system 100 can be understood as an industrial intelligence layer (HL), a chatbot or a query engine capable of assisting the users / plant operators with a structured response to the queries. The system 100 can also be deployed as a co-pilot assisting the plant operators with proactive insights on the plant operations by automatically querying the system 100.
[0075] It should be understood that the components disclosed herein are only for the purpose of illustration and should not be construed limiting to the claims appended herein.
[0076] Referring to FIG 2, illustrated is a block-diagram of a system 200 for answering a natural language user query in an industrial environment, in accordance with another embodiment of the present invention. The system 200 comprises one or more first LLMs 202-1 to 202-N and corresponding knowledge bases 204-1 to 204-N, the second large language model (LLM) 104, an apparatus 106, a client device 110, and one or more sources 206-1 to 206-N. The comprises one or more first LLMs 202-1 to 202-N and corresponding knowledge bases 204-1 to 204-N, the second large language model (LLM) 104, an apparatus 106, a client device 110, and one or more sources 206- 1 to 206-N are communicatively coupled to each other via a communication network 108.
[0077] It is to be appreciated that the system of claim 1 is scaled up to train multiple first LLMs 202-1 to 202-N using knowledge bases 204-1 to 204-N such that the system 200 is capable of answering queries pertaining to varied domains in an accurate manner. The one or more source 206-1 to 206-N are intelligently selected to generate knowledge bases 204-1 to 204-N for corresponding first LLMs 202-1 to 202-N. Each of the first LLMs 202-1 to 202-N is trained specifically on a particular domain. The different domain can be manufacturing, automotive, chemical, pharmaceutical, healthcare, power generation, food & beverage, process industry, energy management etc. In an example, the different domains can also be asset / task specific such as maintenance of chillers, remaining useful life of motors, supply chain management in a plant, energy efficiency in a plant, etc.
[0078] The apparatus 106 comprises a domain selection module (not shown) configured for identifying a domain of the received natural language query. Once the domain of the query is identified, a suitable first LLM and corresponding knowledge base is selected based on the identified domain of the received natural language query. Such a system is a domain agnostic system capable of accurately answering queries from varied domains.
[0079] Referring to FIG 3, illustrated is an apparatus 106 answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention. The apparatus 106 is configured for answering the natural language query in the industrial environment. In the present embodiment, the apparatus 110 is deployed in a cloud computing environment. As used herein, “cloud computing environment” refers to a processing environment comprising configurable computing physical and logical resources, for example, networks, servers, storage, applications, services, etc., and data distributed over the network 108, for example, the internet. The cloud computing environment provides on-demand network access to a shared pool of the configurable computing physical and logical resources. The apparatus 106 may include a module configured for answering a natural language user query in an industrial environment using a combination of knowledge from a locally trained domain specific LLM and a public LLM. Additionally, the apparatus 106 may include a network interface (not shown) for communicating with first LLM 102, second LLM 104, client devices 110, and one or more sources 114-1 to 114-N. In another embodiment, the apparatus 106 can be an edge computing device. As used herein “edge computing” refers to computing environment that is capable of being performed on an edge device (e.g., connected to one or more sensing units in an industrial setup and one end and to a remote server(s) such as for computing server(s) or cloud computing server(s) on other end), which may be a compact computing device that has a small form factor and resource constraints in terms of computing power. A network of the edge computing devices can also be used to implement the apparatus 106. Such a network of edge computing devices is referred to as a fog network.
[0080] As shown in FIG 3, the apparatus 106 comprises a processing unit 135, a memory unit 140, a storage unit 145, an input unit 155, an output unit 160 and a standard interface or bus 199. The apparatus 106 can be a computer, a workstation, a virtual machine running on host hardware, a microcontroller, or an integrated circuit. As an alternative, the apparatus 106 can be a real or a virtual group of computers (the technical term for a real group of computers is “cluster”, the technical term for a virtual group of computers is “cloud”).
[0081] The term “processing unit” 135, as used herein, means any type of computational circuit, such as, but not limited to, a microprocessor, microcontroller, complex instruction set computing microprocessor, reduced instruction set computing microprocessor, very long instruction word microprocessor, explicitly parallel instruction computing microprocessor, graphics processor, digital signal processor, or any other type of processing circuit. The processing unit 135 may also include embedded controllers, such as generic or programmable logic devices or arrays, application specific integrated circuits, single-chip computers, and the like. In general, a processing unit 135 may comprise hardware elements and software elements. The processing unit 135 can be configured for multi-threading, i.e. the processing unit 135 may host different calculation processes at the same time, executing the either in parallel or switching between active and passive calculation processes.
[0082] The processing unit 135 is configured to receive the natural language query from the user, the natural language query pertaining to a particular domain of the industrial environment. Further, the processing unit 135 is configured to parse the natural language query to determine one or more entities and an intent of the natural language query. Further, the processing unit 135 is configured to generate a first set of prompts from the natural language query based on a defined taxonomy pertaining to the particular domain of the industrial environment and the determined intent of the user wherein the domain taxonomy is a hierarchical structure comprising keywords in the domain mapped to corresponding description pertaining to the domain. Further, the processing unit 135 is configured to determine if an answer to the natural language query is known to a first LLM based on the determined intent of the natural language query, wherein the first LLM is specifically trained on the particular domain of the industrial environment and enriched with a knowledge base, wherein the knowledge base comprises information from one or more data sources combined with domain knowledge from a second LLM. Further, the processing unit 135 is configured to query the natural language query to the first LLM using the generate prompts if it is determined that the answer is known to the first LLM. Further, the processing unit 135 is configured to generate a response to the natural language query based on an output of the first LLM.
[0083] The memory unit 140 may be volatile memory and non-volatile memory. The memory unit 140 may be coupled for communication with the processing unit 135. The processing unit 135 may execute instructions and / or code stored in the memory unit 140. A variety of computer-readable storage media may be stored in and accessed from the memory unit 140. The memory unit 140 may include any suitable elements for storing data and machine-readable instructions, such as read only memory, random access memory, erasable programmable read only memory, electrically erasable programmable read only memory, a hard drive, a removable media drive for handling compact disks, digital video disks, diskettes, magnetic tape cartridges, memory cards, and the like.
[0084] The memory unit 140 comprises the industrial intelligence module 165 in the form of machine- readable instructions on any of the above-mentioned storage media and may be in communication to and executed by the processing unit 135. The module 165 comprises a domain knowledge acquisition module 170, an asset data acquisition module 172, a domain taxonomy generation module 175, a knowledge base creation module 180, a first LLM generation module 182, a natural language query intent extraction module 185, a confidentiality module 190, a prompt generation module 192, an Al workflow trigger module 195, and a response generation module 198.
[0085] The domain knowledge acquisition module 170 is configured for acquiring data from the one or more data sources having information related to a particular domain as per requirement of a user, for example, the requirement can be for varied domains like, power generation, automation factory, chemical plant, pharmaceutical plant, automation plant, etc. The domain knowledge acquisition module 170 is configured for determining relevant one or more data sources and then selecting the same for acquiring the information. The information acquired from the one or more selected data sources can be domain related information such as a manufacturing process information, product related information, assembly line related information, material related information, and the like. The one or more data sources can be product information databases, product manuals, or any other public databases from where domain related information can be acquired. The asset data acquisition module 172 is configured for acquiring data from the one or more data sources having information related to one or more assets relevant to the targeted domain. The asset data acquisition module 172 selects one or more assets based on the requirements of the user and collects information from the one or more data sources. The asset data is acquired for one or more data sources like product manuals, sensors, Al models, RCA reports, database including asset specification, predetermined knowledge graphs, images, videos, audios, or a combination thereof.
[0086] The domain taxonomy generation module 175 is configured for identifying a plurality of keywords from one or more sources pertaining to a particular domain. The domain taxonomy generation module 175 is configured for a textual description of each of the keywords from the one or more sources comprising both domain related information and asset related information. The domain taxonomy generation module 175 is configured for defining a relationship between the keywords and the textual description by annotating the keywords. The domain taxonomy generation module 175 is configured for generating the domain taxonomy with identified keywords and corresponding textual description based on a defined hierarchy.
[0087] The knowledge base creation module 180 is configured for generating a set of questions using the domain taxonomy based on a predefined template. The knowledge base creation module 180 is configured for determining a confidentiality level of the questions generated and then query the second LLM or the public LLM for the questions that do not comprise confidential entities. Further, the knowledge base creation module 180 is configured for triggering Al workflows in the one or more Al models for the set of questions that comprise confidential entities. The knowledge base creation module 180 is configured for generating the knowledge base by combining the output from Al workflows the second LLM and store the answers in an indexed manner.
[0088] The first LLM generation module 182 is configured for generating a set of questions from the knowledge base using the domain taxonomy. The first LLM generation module 182 is configured for querying the second LLM with the generated set of questions iteratively with different variations of the generated set of questions. The first LLM generation module 182 is configured for querying the first LLM with the generated set of questions iteratively with different variations of the generated set of questions. Further, the first LLM generation module 182 is configured for iteratively calculating a distillation loss between the responses of the first LLM and the second LLM. The first LLM generation module 182 is configured for tuning one or more parameters of the first LLM such that the distillation loss decreases after each iteration. The first LLM generation module 182 is configured for generating the first LLM when the calculated distillation loss for each of the set of questions in below a predefined threshold. The natural language query intent extraction module 185 is configured for receiving the natural language query from the user and parsing the natural language query to determine one or more entities and an intent of the natural language query.
[0089] The confidentiality module 190 is configured for classifying the query as confidential and non- confidential for further processing. The confidentiality module 190 is configured for classifying the query based on the determined entities of the natural language query identified based on the comparison with a database comprising a plurality of keywords that are annotated as confidential.
[0090] The prompt generation module 192 is configured for generating a first set of prompts from the natural language query based on a defined taxonomy pertaining to the particular domain of the industrial environment and the determined intent of the user when an answer is known to the first LLM. Further, the prompt generation module 192 is configured for generating a second set of prompts from the natural language query based on a defined taxonomy when an answer is not known to the first LLM. Further, the prompt generation module 192 is configured for generating a third set of prompts by replacing the confidential entities with general descriptions using the domain taxonomy in the natural language query, and a fourth set of prompts by retaining the confidential entities in the natural language query when the natural language query comprises entities marked as both confidential and non-confidential.
[0091] The Al workflow trigger module 195 is configured for triggering Al workflows based on the intent of the natural language query, when an answer is not known to the first LLM. The Al workflow trigger module 195 is configured for identifying the correct workflows to be triggered based on the intent of the query. Further, Al workflow trigger module 195 is configured for updating the knowledge base with the outputs of Al models.
[0092] The response generation module 198 is configured for generating a response to the natural language query based on an output of the first LLM enrich with knowledge base. Further, response generation module 198 is configured for generating a response to the natural language query based on an output of the first LLM enriched with updated knowledge base. The response generation module 198 is configured for generating a response from the query by structuring an answer retrieved from both the first LLM and the second LLM.
[0093] The storage unit 145 comprises a non-volatile memory which stores the knowledge base, reports generated from Al workflows, domain taxonomy, historical set of prompts etc. The storage unit 145 includes the database 150 that comprises the knowledge base, reports generated from Al workflows, domain taxonomy, historical set of prompts etc. The bus 199 acts as interconnect between the processing unit 135, the memory unit 140, the storage unit 145, the input unit 155 and the output unit 160. The input unit 155 enables the user to input the natural language query, and / or one or more prompts. The output unit 160 is configured for presenting the response to the natural language query.
[0094] Those of ordinary skilled in the art will appreciate that the hardware depicted in FIG 1 , FIG 2 or FIG 3 may vary for different implementations. For example, other peripheral devices such as an optical disk drive and the like, Local Area Network (LAN) / Wide Area Network (WAN) / Wireless (e.g., Wi-Fi) adapter, graphics adapter, disk controller, input / output (I / O) adapter, network connectivity devices also may be used in addition or in place of the hardware depicted. The depicted example is provided for the purpose of explanation only and is not meant to imply architectural limitations with respect to the present disclosure.
[0095] Referring to FIG 4, illustrated is a block diagram 400 of an architecture of the system of FIG 1 , in accordance with an embodiment of the present invention. The block diagram 400 illustrated a layered software architecture view of the system 100. As shown, the architecture of the system 100 comprises an artificial intelligence (Al) engine user interface layer 402, Al engine backend layer 404, data storage layer 406, data processing layer 408, communication layer 409, and data collection layer 410. Herein, the system 100 also comprises a second LLM or public LLM 411 and a first LLM or local LLM 412. It should be understood that the layers and components disclosed herein are only for the purpose of illustration and should not be construed limiting to the layers and components as illustrated herein.
[0096] The Al engine user interface layer 402 comprises an Al model library module 414, Al user manager module 416, query module 418, Al workflow manager module 420, system administration module 422. The Al engine user interface layer 402 is configured to select analytics task set or workflow and / or a particular technical installation (or assets / equipment) to be executed based on the natural language query. In an example, the selected workflow maybe process industry, building technology, discrete factory, brownfield, greenfield, or assets like heat exchangers, chillers or pumps. Such a decision can be made by user selection or automatically via the Al engine user interface layer 402 based on the set of requirements by the user. The Al engine user interface layer 402 is configured to determine required performance indicators (KPIs), fault descriptions for different assets as a function of available sensors, and integration of domainbased know-how in the front end.
[0097] The Al engine user interface layer 402 is also configured to provide web-based access, select cyber check and different user profiles. Furthermore, Al engine user interface layer 402 is configured to transmit the domain information to the Al engine backend layer 404 for further processing. Furthermore, the Al engine user interface layer 402 is configured to trigger Al model workflows, for example, selection of different Al model workflows that maybe triggered for different assets like pumps or heat exchangers. For example, there maybe an exhaustive list of fault monitoring Al workflows provided for pumps. In such a case, the Al engine user interface layer 402 is configured to select the appropriate Al workflow based on the available data sources and based on the intent of the natural language query. Furthermore, the Al engine user interface layer 402 is configured to initiate real time data ingestion services, configuration of data sources to access relevant data, selection of paths to data sources, for example, file server, OPC-LIA server, historian and the like. Furthermore, the Al engine user interface layer 402 is configured to initiate Al model training, Al model re-training, Al model testing or Al model deployment workflow based on the intent of the natural language query. The Al engine user interface layer 402 is further configured to debug the system.
[0098] The Al engine backend layer 404 comprises an Al model library module 424, Al model trainer module 426, query handler module 428, Al workflow manager module 430, system administration and security module 432. The Al engine backend layer 404 is configured to interact with the Al engine user interface layer 402. The Al engine backend layer 404 is configured to validate users for example, using LDAP, SSO and other techniques. Further, the Al engine backend layer 404 is configured to configure the cyber check on the user profiles. Further, the Al engine backend layer 404 is configured to configure the Al workflows to be execute, for example, process industry, building technology, discrete factory, or for particular assets based on the natural language query or prompts during tuning of first LLM. In an example, the Al engine backend layer 404 is also configured to allow user to select a workflow with multiple assets tied together for configuration such as overall optimizations, virtual sensing, which may require information from different assets. The Al engine backend layer 404 is configured for analyzing the intent of the query and determine which Al workflows are to be executed. Furthermore, the Al engine backend layer 404 is configured to train the selected Al models using both offline and online training, autonomous data pre-processing, optimization workflows and so forth.
[0099] Furthermore, the Al engine backend layer 404 is configured to deploy the Al models, manage the Al model library, interact with other visualization or model building tools, provide debugging data in the form of workflow logs and so forth. Furthermore, the Al engine backend layer 404 also provide a templating feature for the new assets or new asset categories in the industrial environment. The Al engine backend layer 404 provide the ability to seamlessly add any new asset into the system by adding all related metadata, diagrams, sensors, KPIs and any other available information. Furthermore, the Al engine backend layer 404 is configured to generate synthetic data using various DOE (Design Of Experiments) for the purpose of proof of concept for several workflows. Furthermore, the simulation models provide ability to incorporate new assets into the Al model library. The predictive analytics engine backend layer 404 is configured to train offline Al model and simulation using synthetic data / real data. Similarly, the predictive analytics engine backend layer 404 is configured to train online Al Model and use real-time data acquired from assets and data sources. Furthermore, the Al engine backend layer 404 is configured to test the accuracy of the Al Models using test data, ability to do blind tests and provide an insight to the quality of outputs (multivariate input data range checks).
[0100] The data storage layer 406 comprises knowledge base 434, MongoDB 436, and time series database 438. The data storage layer 406 is configured to store data sensor data, knowledge domain information, Al models and so forth as received from the data collection layer 410 through the data communication layer 409.
[0101] The data processing layer 408 comprises a KPI calculator module 440, a data cleaning module 442, a data structuring module 444, and a data normalization module 446. The data processing layer 408 is configured for pre-processing the data as received from the data storage layer 406.
[0102] The communication layer 409 comprises OPC UA module 448, TCP / IP module 450, and MQTT module 452. The communication layer 409 acts an interface between the data collection layer 410 and the data processing layer 408 and is configured for transmitting the collected data from the data collection layer 410 to the data processing layer 408.
[0103] The data collection layer 410 comprises one or more assets including chillers 454, pumps 456, HVAC 458, PLC 460, exhaust fans 462. It should be understood that these are mere examples and not limited to the particular assets mentioned here. The data collection layer 410 is configured for directly interacting with the data sources associated with the corresponding assets to acquire data from the assets in real-time.
[0104] The second LLM 411 is a public domain LLM and the first LLM 412 is a domain specific local LLM.
[0105] Referring to FIG 5, illustrated is a block diagram 500 of an architecture of the system of FIG 1 , in accordance with another embodiment of the present invention. As maybe seen in the block diagram 500, the architecture comprises a second LLM 502 (a public LLM such as Open Al ChatGPT, Gemini, Llama, Bard etc.) communicatively coupled with the Al engine interface layer 504, and Al engine backend layer 506. Further, the Al engine backend layer 506 is interacting with modules of plurality of assets 508 such as a chillers 538, coffee roasters 540, pumps 542, cement kiln 544, heat exchanger 546, motor 548, boiler 550, and HVAC 552.
[0106] The Al engine Ul layer 504 further comprises an application framework 510. The application framework 510 comprises a query Ul 512, a query handler 514, a query workflow initiator 516, asset template 518, and a first LLM agent 520. The Al engine backend layer 506 comprises a Al model training module 522, Al model testing module 524, KPI calculator module 526, workflow manager module 528, model manager 530, knowledge base 532, and a navigation module 534 comprising a Kafka module 536.
[0107] It should be understood that the layers and components disclosed herein are only for the purpose of illustration and should not be construed limiting to the layers and components as illustrated herein.
[0108] Referring to FIG 6, depicted is a flowchart of a method 600 for answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention.
[0109] At step 602, the natural language query is received from the user. The natural language query pertaining to a particular domain of the industrial environment. The natural language query is a type of search or input query that is expressed in ordinary human language, rather than using formalized syntax or specific query languages. It allows users to communicate with computers or systems using natural language, just like they would with other humans. In an example, the natural language query is input via user through a user interface. The user inputs the user query into the system via a user interface. Below are a few examples of user queries in the industrial environment:
[0110] 1. What is an estimated time for getting a burner cut?
[0111] 2. When the burner cut happens what is the root cause behind it?
[0112] 3. Generate the RCA report for the burner cut
[0113] 4. What are the alarms that occurred leading to the burner cut?
[0114] 5. When should I schedule the maintenance of my coffee roaster?
[0115] 6. How old is my green coffee batch?
[0116] At step 604, the natural language query is parsed to determine one or more entities and an intent of the natural language query. The one or more entities in the natural language query are words or phrases that express the main concepts in the query. In this context such keywords or phrases maybe related to a particular domain, asset, sensor, date and / or time, location, a process / workflow etc. In an example, the natural language query is a combination of linguistic, semantic, and pragmatic elements that together convey the user's information needs in a contextually rich manner. In order to accurately understand the intent of the user, the natural language query is interpreted and analyzed using advanced natural language processing (NLP) techniques and machine learning algorithms capable of understanding and processing the natural language query. It is to be understood that the NLP techniques are known in the art and are not explained further for the brevity of the description of the invention.
[0117] In an embodiment, the method of determining the intent of the natural language query comprises parsing the natural language query to extract one or more entities and determine semantic and syntactic relationships therebetween. The natural language query is broken down into one or more entities or phrases and then meaning of each and a relationship between each word is determined. The one or more entities maybe related to a particular domain, asset, sensor, date and / or time, location, a process / workflow etc. The method further comprises determining an intent of the natural language query based on a machine learning model trained on a historical database of one or more entities annotated with their corresponding intent. The historical database comprises one or more entities that are annotated with a corresponding intent of the entities and their relationships. The machine learning model is training on identifying the correct intent of the natural language query based on the annotated entities with their intent.
[0118] In an exemplary implementation, there are two approaches to query intent extraction. If ChatGPT or larger foundation models are available, detailed definition of various intents are given as a context to the LLM model and is prompted to classify the user query in one of the defined intents. If there is no access to foundation models, a classification model, which is trained on historical examples of user conversations annotated by intent and then the trained model is used for intent classification.
[0119] Advantageously, the objective of the correct intent identification of the query is to route the user query to the appropriate data sources so as to provide an accurate answer. For example, imagine the knowledge base in the apparatus currently has information regarding a cement kiln and the events associated with it. If the user queries a question regarding the manufacturing process of paints, then the knowledge base won't have the answer for it. So, the intent identification model will identify that question as something that is out of the scope of the local knowledge base and will redirect the query to give an answer from open-source knowledge.
[0120] The intent identification models are trained with labeled sample questions to distinguish between questions related to cement kiln and generic questions. It will also be trained to identify domain specific key words (“KN3TI3456A”) in the question. Based on the nature of the question, the intent identification model can trigger suitable actions such as querying the knowledge base, the first LLM, the second LLM, or trigger questions to the user if the original query is missing some information. At step 606, a first set of prompts are generated from the natural language query based on a defined taxonomy pertaining to the particular domain of the industrial environment and the determined intent of the user query. The prompts are textual input provided to the model to generate responses or perform specific tasks, prompts consist of one or more sentences or passages of text that serve as input to the LLM. This input can include questions, statements, commands, or any other form of textual communication. The prompts convey the intent of the natural language query to the language model. In general, prompts may include contextual information relevant to the desired task or response. This could include background information, constraints, preferences, or other contextual cues that help guide the response of the LLM. In some examples, the prompts follow a specific format or structure depending on the requirements of the task or application. This could involve using certain keywords, phrases, or formatting conventions to convey instructions or constraints to the LLM. In some cases, the prompts may include examples or samples of the desired output to provide additional context or guidance to the model. This aids to clarify the user's expectations and improve the quality of the generated responses. In the context of LLMs, the prompts refer to as cues or instructions for the model to produce desired outputs based on the input provided, for example, the natural language query. Advantageously, the prompts are generated for accurately querying the first LLM and / or second LLM.
[0121] The first set of prompts are generated from the natural language query based on a defined taxonomy. The term “domain taxonomy” refers to a hierarchical classification system that organizes concepts, entities, or topics within a specific domain or field of knowledge. The domain taxonomy provides a structured framework for categorizing and organizing information based on relationships, similarities, and hierarchies. The domain taxonomy is organized in a hierarchical manner, with broader categories or concepts at the top level and increasingly specific subcategories or subtopics at lower levels. This hierarchical structure enables the system to navigate from general to more detailed levels of information. Furthermore, the domain taxonomy defines a classification scheme that categorizes entities or concepts into distinct classes or categories based on shared characteristics or attributes. Furthermore, the domain taxonomy capture relationships and dependencies between different classes or categories within the domain. Moreover, the domain taxonomy uses standardized terminology and naming conventions to ensure consistency and clarity in the classification of concepts. The domain taxonomy also comprises specific field of knowledge, industry, or subject area, reflecting the unique characteristics, concepts, and relationships within that domain. It may incorporate domain-specific terminology, concepts, and principles relevant to the subject area. It is to be noted that the domain taxonomy evolves over time to accommodate changes in the domain, new discoveries, or emerging trends with respect to the chosen domain. In specificity, the domain taxonomy is a hierarchical structure comprising keywords in the domain mapped to corresponding description pertaining to the domain. It is to be understood that the domain taxonomy is generated based on a requirement of the user or a use case for which the system is deployed for. In an example, the domain is an industrial domain such as maintenance of chillers. The domain taxonomy comprises all information pertaining to chillers, different types of chillers, condensers, evaporators, pumps, valves, control systems, and all sensor parameters, electrical parameters, sensing units related to the different assets in the chiller prediction system and their corresponding descriptions. The domain taxonomy will also contain all information pertaining to operation of chillers, operation parameters such as temperature of fluid entering evaporator coil, typically chilled water or a refrigerant, temperature of fluid leaving the evaporator coil after absorbing heat from the chilled space, pressure inside evaporator coil, pressure inside condenser coil, pressure of refrigerant, flow rate of chiller water, compressor power, pump power, refrigerant temperature, energy efficiency ratio and other relevant parameters and the corresponding descriptions related to the same. The domain taxonomy comprises information related to predictive system of the chiller, sensor information, etc. and their corresponding descriptions.
[0122] In an embodiment, the method of generating domain taxonomy comprises identifying a plurality of keywords from one or more sources pertaining to a particular domain. The sources are databases from where information pertaining to the particular domain is extracted. In an example, the one or more sources are product catalogs, product manuals, industry standards, regulatory documents, patents, technical specifications, and reference materials. The sources can also be knowledge graphs for the domain, ontologies, historical databases, etc. Further, the method comprises extracting a textual description of each of the keywords from the one or more sources. The method comprises defining a relationship between the keywords and the textual description by annotating the keywords. Furthermore, the method comprises generating the domain taxonomy with identified keywords and corresponding textual description based on a defined hierarchy. In an example, a hierarchy is selected or defined by the system along with human in the loop, and then information extracted from the one or more sources is arranged accordingly to generate the domain taxonomy.
[0123] In an exemplary implementation, the domain taxonomy comprises the keywords from the domain which are generated by a semi-automated process of getting entities involved in the data (from different sources and format). In the process, a domain user can provide more generic description of the keywords. The domain taxonomy will have a set of keywords along with their descriptions (both domain and general world) extracted from the data or provided by human in the loop annotation. For instance, for a predictive maintenance domain of industrial coffee roaster, the time series tag variables itself become the keyword for which we can have their descriptions and general world description, like “CO920” with the description “Carbon monoxide level in industrial coffee roaster”. Advantageously, the domain taxonomy will be utilized then to get more open world information which will enable the system to make more richer inferences.
[0124] In an embodiment, the method of generating prompts from the natural language query for querying the first LLM comprises retrieving information that is similar to the keywords in the natural language query from domain taxonomy. The method further comprises generating prompts for querying the first LLM based on the descriptions extracted from the domain taxonomy.
[0125] In an exemplary implementation, the prompts are generated in varied manner. For example, the natural language query as received form the user can serve as prompt to the first LLM for generating the response. The user query is directly used to retrieve information from vector databases which is used as a context to the LLM (RAG + QnA). If second LLMs such as ChatGPT or larger foundation models are available, they are first used to elaborate on the user query, i.e. , answer in brief the user query in domain rich format. The output of the second LLM is then used to query vector databases and find similar information from the knowledge base. The information retrieved from the vector databases is used as context to answer user query from the first LLM. In another example, the received input is used to query the historical events like anomalies, root causes. The underlying Al system and time series verbalization of the Al models output provides a domain rich textual representation of time series data and the events occurred in it. This domain rich text is then used to query the vector databases to identify causes and solutions for the event. The first set of prompts generated to be provided to the first LLM are then adjusted according to the event. The information retrieved from the database is then given to the user. The information retrieved from the vector database is structured according to the event using the available second LLM or public LLM and then the first set of prompts are generated. The first set of prompts are then used to query the first LLM.
[0126] At step 608, it is determined if an answer to the natural language query is known to a first LLM based on the determined intent of the natural language query. As mentioned earlier, the first LLM is a local LLM that is specifically trained on the particular domain. The term “Large Language Model (LLM)” refers to a type of artificial intelligence model designed to understand and generate human-like text based on vast amounts of natural language data. These models are built using deep learning architectures, particularly transformer-based architectures. LLMs are trained on massive datasets containing billions or even trillions of words from diverse sources such as books, articles, websites, and other textual content. The large-scale training enables the model to learn complex patterns, structures, and nuances of human language. LLMs are typically built using deep learning architectures, particularly transformer architectures.. LLMs undergo two main stages of training: pre-training and fine-tuning. During pre-training, the model is trained on a large corpus of text data using unsupervised learning techniques to learn general language patterns and semantics. In fine-tuning, the pre-trained model is further optimized on domain-specific or task-specific datasets to adapt its knowledge and capabilities to specific applications. LLMs have the ability to generate human-like text based on given prompts or inputs. They can produce coherent paragraphs, articles, stories, code, or responses to questions by predicting the next words or tokens in the sequence based on the context provided. LLMs exhibit a strong understanding of context and semantics in natural language. They can infer meaning, resolve ambiguity, and generate text that is contextually relevant and coherent with the given input.
[0127] The first LLM is specifically trained on the particular domain of the industrial environment and enriched with a knowledge base. The knowledge base comprises information from one or more data sources combined with domain knowledge from a second LLM. The generation of the knowledge base is further explained in FIG 7, FIG 8 and FIG 9. Also, the generation of the first LLM is further explained in FIG 10, FIG 11 , and FIG 12.
[0128] Based on the intent of the user query, it is decided whether the user query is passed through the first LLM, the second LLM or through the knowledge base. The intent of the query can be determined using the machine learning model as aforementioned. In a case when it is determined that the intent of the natural language query is generic domain related, then the second LLM is queried. In a case, when it is determined that the intent of natural language query is historic event related to assets, then the first LLM is queried. These cases are explained in detail later in the description.
[0129] At step 610, the first LLM is queried with the natural language query using the generated prompts, if it is determined that the answer is known to the first LLM. At step 612, a response to the natural language query is generated based on an output of the first LLM. In an embodiment, the response to the natural language query is output in at least one of the formats: textual, images, videos, reports, and a combination thereof. The response to the natural language query is structured based on the intent of the user query. In an example, if the intent of the user is to receive only a text-based response, then the answer is generated in a text format. In another example, if the intent of the natural language query is to receive a report or a video file to explain the requested data, then response is structured accordingly.
[0130] In an embodiment, the method further comprising identifying a confidentiality level of the natural language query of the user based on the determined entities of the natural language query identified based on the comparison with a database comprising a plurality of keywords that are annotated as confidential. The database is maintained with one or more keywords that are marked as confidential for a particular domain, industry, asset or by an operator. Such a database can vary for each industrial factory floor or use case based on the requirements of the operator / user / owner of the system. Advantageously, the level of confidentiality is determined such that the query is not triggered to a open / public LLM, thereby preventing a breach of confidential data.
[0131] In an embodiment, the method of answering the received natural language query comprises querying to the first LLM using the generated prompts, if it is identified that the natural language query comprises one or more entities marked as confidential in the database. Advantageously, in order to maintain confidentiality of customer data or plant data, the natural language query comprising confidential entities are only directed to the first LLM i.e. the local LLM enriched with knowledge base. It is to be understood that if the intent of the query is to retrieve answers related to generic domain knowledge and does not comprise any confidential entities, then the query is directed to the second LLM. In another case, when the intent of the query is to retrieve plant / asset related specific information, and / or comprises confidential entities, then the query is directed to the first LLM.
[0132] In an embodiment, the method of answering the received natural language query comprises executing one or more Al workflows in the industrial environment based on the intent of the natural language query, if it is determined that the answer to the natural language query is not known to the first LLM, and the natural language query comprises one or more entities marked as confidential in the database. It is to be understood that the Al workflows are executed to generate output from one or more Al models deployed in the industrial environment for various tasks such as predictive maintenance, condition monitoring, quality control, process optimization, energy management, supply chain optimization, resource allocation, safety monitoring, compliance monitoring and the like. The Al models and their task vary as per the requirements of the plant operators. The term “Al models” as used herein refer to various computational models and algorithms designed to mimic human intelligence or perform tasks that typically require human intelligence. These models are used to analyze data, make predictions, solve problems, and automate tasks across a wide range of domains. Various examples of Al models are machine learning models (supervised learning, unsupervised learning, reinforcement learning etc.), deep learning models (artificial neural networks AN Ns, convolutional neural networks CNNs, recurrent neural networks RNNs, transformer-based models etc.), probabilistic models (Bayesian networks, hidden Markov models HMMs, etc.), evolutionary algorithms like genetic algorithms, and so on. In an example, Al model prediction model given a set of training dataset. In general, each individual sample of the training data is a pair containing a dataset (e.g., operating conditions of the asset and corresponding parameter values of the asset) and a desired output value or dataset (e.g., parameter values of the asset). The machine learning model analyzes the training data and produces a predictor function. The predictor function, once derived through training, is capable of reasonably predicting or estimating the correct output value or dataset. The term ‘workflow’ as used in one or more “Al workflows” refers to a set of instructions or task set that is required to be executed in order to determine the outcome based on the intent of the natural language query. Notably, every action to be performed by the Al model is modeled as a workflow. The workflow is determined and initiated based on the intent of the user query. The workflow may comprise one or more tasks for the asset identified from the user query. In an embodiment, task is at least one of: anomaly detection in the asset, root cause analysis of the asset, remaining useful life estimation of the asset, performance optimization of the asset, forecasting of one or more parameters associated with the asset and energy optimization of the asset. The one or more tasks may vary for different assets and different technical installations based on the requirements. In one example, when the technical installation is a building, the task may be forecasting load, forecasting temperature, forecasting humidity, and so forth. Once the workflow is defined as per the requirement, the workflow is submitted for execution.
[0133] In an example, when the query is ‘Predict the efficiency of pumps in the system’, then the pumps in the technical installation, the corresponding sensors and the Al models for determining efficiency of the pumps are selected to determine the desired output i.e. the efficiency of the pumps. In another example, when the set of requirements is ‘Provide maintenance schedule of the chillers’, then the chillers installed in the outlet or factory and the corresponding Al models will be selected to initiate the workflow for predicting maintenance schedule of the chillers. In yet another example, when the requirement is ‘predict energy consumption in a building’, then the entire building comprising multiple assets and corresponding one or more Al models are selected for predicting energy consumption of the building. It will be appreciated that the disclosed system is capable of correctly identifying the Al workflows based on the intent of the user query.
[0134] The method comprises updating the knowledge base with output of the one or more Al workflows triggered. Once the Al workflows are executed, the output of the one or more Al models is then used to update the knowledge base for future queries with a similar intent. Further, the method comprises generating a second set of prompts based on the domain taxonomy. When the knowledge base is updated with the output of the Al models, the second set of prompts are generated to query the first LLM with updated knowledge base. The method comprises querying the natural language query to the first LLM enriched with the updated knowledge base using the second set of prompts. Further, the method comprises generating a response to the natural language query based on an output of the first LLM enriched with the updated knowledge base.
[0135] In an embodiment, the method of answering the received natural language query comprises querying to the second LLM using the generated prompts, if it is identified that the natural language query does not comprise one or more entities marked as confidential in the database the answer to the natural language query is not known to the second LLM. Advantageously, the system directly routes the queries to the public LLM is the intent of the query is to retrieve generic domain related information.
[0136] In an embodiment, the method of answering the received natural language query identified as comprising entities marked as both confidential and non-confidential comprises identifying one or more confidential entities in the natural language query. The method comprises generating a third set of prompts by replacing the confidential entities with general descriptions using the domain taxonomy, and a fourth set of prompts by retaining the confidential entities. Notably, when the query comprises both confidential and non-confidential entities and the intent of the query is not fulfilled by local LLM, then the query is broken into different prompts to be queried to local LLM and the public LLM. The third set of prompts are generated by replacing the keywords in the query marked as confidential with general world descriptions. The method comprises querying the second LLM with the third set of prompts. Further, the fourth set of prompts are generated by retaining the keywords in the query that are marked as confidential. The method comprises querying the first LLM with the fourth set of prompts. The method comprises structuring a response from the first LLM and the second LLM in order to generate a natural language response to the natural language query.
[0137] In an exemplary implementation, the intent of the query based on an initial classification model (DIET classifier) is used to classify the natural language query into two categories: 1. Intent to query historic events. 2. Intent to query generic domain related information
[0138] In a scenario, when the intent is querying historical events, then the below workflow is followed: a. User can specify the date and time specific to the query. b. Identify and display the events around that time. c. An event is an anomaly, maintenance activity, incident report, root cause report or any other operations related event that has occurred in the plant and was either detected by one of the underlying Al systems or entered manually by the operator. d. The event is verbalized and passed as context to the local LLM along with the domain information from knowledge base. The local LLM answers operator query about the event based on this context and the open world knowledge local LLM has. The prompts are engineered to get domain rich information about the event. e. User is then given a set of options in case he wants more information on the event. i. Full report - Give the full report in a structured format (template based) ii. Time Series Verbalization - Highlights of the time series data around the event in a verbalized format explaining the event to some extent. iii. Trigger Al models - User can trigger various Al workflows via text classification, forecasts, regression etc. to explore the event more.
[0139] In another scenario, if intent is querying from knowledge base, the user can query the knowledge base and ask domain related questions like information about the equipment, inspection assistance, explaining inspection results, maintenance steps etc. The knowledge base is a vector database. The vector database is queried using the operator query and the top n hits (configurable) are passed as context to the first LLM or the local LLM. The local LLM answers operator query based on this context and the open world knowledge local LLM was trained on. There is a feedback loop to check if the user is satisfied with the information available in the knowledge base. If not, the strong open world LLMs like ChatGPT are queried directly with the query. There is a confidentiality check enforced before sending the query to open world LLMs. If the user / operator is satisfied with the open world LLM answer, the knowledge from open world is then appended to the knowledge base.
[0140] Referring to FIG 7, illustrated is a block diagram 700 depicting different sources of a knowledge base 702 with one or more data sources, in accordance with an embodiment of the present invention. The knowledge base 702 is a heterogeneous database comprising information pertaining to the domain and the industrial environment. The knowledge base is a centralized repository or database that contains domain-specific information, asset specific information, process- specific information, sensor-specific information, etc. The knowledge base also contains information relevant to the specific industry, sector, or domain in which the organization operates. This may include technical specifications, manufacturing processes, equipment manuals, safety procedures, regulatory requirements, and industry standards. The knowledge base also comprises information retrieved from one or more Al models deployed in the industrial environment. The knowledge base also comprises information extracted from documentation of performance parameters, efficiency parameters, anomalies, root cause analysis, prediction and resolution anomalies, etc. The knowledge base also comprises information extracted from knowledge graphs comprising information of the plant and its assets in a hierarchical manner. The knowledge base also comprises images, videos, audios etc. having information of the industrial environment. Advantageously, the information in the knowledge base is stored in in a structured manner to facilitate easy navigation, retrieval, and access to information. Information is typically categorized into topics, sections, or modules based on the subject matter, making it easier for users to locate relevant content.
[0141] In an example, knowledge base is generated for system for maintenance of chillers. The knowledge base comprises all information pertaining to chillers, different types of chillers, condensers, evaporators, pumps, valves, control systems, and all sensor parameters, electrical parameters, sensing units related to the different assets in the chiller prediction system. The knowledge base will also contain all information pertaining to operation of chillers, operation parameters such as temperature of fluid entering evaporator coil, typically chilled water or a refrigerant, temperature of fluid leaving the evaporator coil after absorbing heat from the chilled space, pressure inside evaporator coil, pressure inside condenser coil, pressure of refrigerant, flow rate of chiller water, compressor power, pump power, refrigerant temperature, energy efficiency ratio and other relevant parameters. The knowledge base also comprises information related to predictive system of the chiller, sensor information, etc.
[0142] As shown, the knowledge base 702 extracts knowledge from the one or more data sources 704, 706, 708, 710, 712, and 714. In an example, the data source 704 is a sensor, the data source 706 is a RCA report, the data source 708 is an Al model, the data source 710 is a database including asset information and other asset related documentation, 712 is a knowledge graph and the data source 714 is a log file. The knowledge extracted form the one or more data sources is further processed to build the knowledge base. The generation of the knowledge is further explained in detail in FIG 8 and FIG 9.
[0143] Referring to FIG 8, depicted is a flowchart of a method 800 for generating the knowledge base pertaining to the particular domain, in accordance with an embodiment of the present invention.
[0144] At step 802, the information is acquired from one or more data sources in the industrial environment. The one or more data sources comprises at least one of: sensors, RCA reports, prediction models, one or more Al models, database including asset specification, predetermined knowledge graphs, images, videos, audios, or a combination thereof. The data is extracted from the one or more data sources based on the domain, and requirements of the user / operator of the system. At step 804, domain taxonomy is generated based on the extracted information. The domain taxonomy comprises a set of keywords and corresponding descriptions extracted from the acquired information, identifying a plurality of keywords from one or more sources pertaining to a particular domain. The sources are databases from where information pertaining to the particular domain is extracted. In an example, the one or more sources are product catalogs, product manuals, industry standards, regulatory documents, patents, technical specifications, and reference materials. The sources can also be knowledge graphs for the domain, ontologies, historical databases, etc. Further, the method comprises extracting a textual description of each of the keywords from the one or more sources. The method comprises defining a relationship between the keywords and the textual description by annotating the keywords. Furthermore, the method comprises generating the domain taxonomy with identified keywords and corresponding textual description based on a defined hierarchy. In an example, a hierarchy is selected or defined by the system along with human in the loop, and then information extracted from the one or more sources is arranged accordingly to generate the domain taxonomy.
[0145] At step 806, a set of questions are generated using the domain taxonomy based on a predefined template. The set of questions are prompts that are generated using the list of keywords and their corresponding descriptions in the domain taxonomy. The set of questions or prompts are textual input provided to the second LLM to generate possible responses from the second LLM. The set of question generated from the domain taxonomy are the possible questions that a user / operator can query the system. The prompts can include questions, statements, commands, or any other form of textual communication. The prompts convey the intent of the possible queries to the second LLM. In general, prompts may include contextual information relevant to the desired task or response. This could include background information, constraints, preferences, or other contextual cues that help guide the response of the LLM. The set of questions are prompts here are based on the predefined template. Advantageously, the set of questions generated vary significantly and aims to include all potential queries such that the knowledge base is accurately enriched with correct information.
[0146] At step 808, a confidentiality level of the generated set of questions is assessed. The confidentiality level of the generated set of questions is assessed based on the keywords in the each of the set of the questions marked as confidential for a particular domain. In case the generated set of questions do not contain confidential entities, the step 810 is executed. In another case, when the generated set of questions contains confidential entities, then the step 812 is executed.
[0147] At step 810, the second LLM is queried using the generated set of questions that do not comprise confidential entities. The second LLM is queried for providing responses to the generic domain related queries based on the set of questions. At step 812, Al workflows are triggered in the one or more Al models for the set of questions that comprise confidential entities. When the confidential entities are included in the questions, the responses are retrieved from the one or more Al models deployed in the industrial environment. The Al workflows are triggered based on the set of questions in order to determine output from the Al models that is then used to generate the knowledge base. At step 814, the knowledge base is generated by combining output from Al workflows and second LLM.
[0148] Referring to FIG 9, illustrated is a block diagram of an architecture 900 for generating the knowledge base pertaining to the particular domain, in accordance with an embodiment of the present invention. The architecture comprises knowledge base 902, an Al layer 904 comprising RCA report module 906, predictive maintenance module 908, analytics module 910, data sources 912, and documentation databases 914. The architecture 900 further comprises a local domain data corpus 916, a domain taxonomy 917, and domain a taxonomy generation module 918. The architecture further comprises a prompt database 919 and a prompt generation module 920. The architecture further comprises a confidentiality assessment module 922, a first report generation module 924, open world domain data corpus 926 and a public LLM 928. The architecture 900 further comprises a second report generation module 930.
[0149] The data from the Al layer 902 is extracted and a local domain data corpus 916 is generated. Based on this local domain data corpus 916 and the domain taxonomy generation module 918, the domain taxonomy 917 is generated. The domain taxonomy comprises data with domain tags and corresponding general world descriptions. The prompt generation module 920 is configured to generate the prompts or a set of open world questions using the domain taxonomy 917. Once the open-world questions or prompts are generated, the prompts are stored in the prompt database 919. The prompts from the prompt database 919 are then passed to the confidentiality assessment module 922 that filters out any questions that comprise confidential entities based on a classification model. All the prompts comprising only non-confidential entities are then queried to the public LLM 928 to retrieve domain knowledge available in the public domain. The responses from the public LLM 928 are used to generate the open world domain data corpus 926. Further, the data from the local domain data corpus 916 and the open world domain data corpus 926 is combined by the first report generation module 924 to generate descriptive textual reports by using natural language generation (NLG) techniques. The first report generation module 924 is capable of generating reports having responses to domain specific queries. Further, asset specific outputs from the Al layer 904 are combined with the output of the first report generation module 924 by the second report generation module 930 to generate asset specific reports comprising textual description for each asset. Further, the reports are stored in the knowledge base 902 with data index for optimal fetching and querying.
[0150] Referring to FIG 10, depicted a flowchart of a method 1000 for generating the first LLM, in accordance with an embodiment of the present invention. The first LLM is an artificial intelligence model designed to understand and generate human-like text based on industrial domain data and one or more Al model deployed in the industrial environment. These models are built using deep learning architectures, particularly transformer-based architectures. The first LLM is trained on massive datasets acquired from the one or more data sources such as sensors, RCA reports, prediction models, one or more Al models, database including asset specification, predetermined knowledge graphs, images, videos, audios, or a combination thereof. The large-scale training enables the model to learn complex patterns, structures, and nuances of human language. The first LLM is built using deep learning architectures, particularly transformer architectures. The first LLM is an in-house Al model that can exhibit the behavior and style of a large language model (LLM) while also performing a specific task with limited data. The first LLM is trained specifically on the domain data using knowledge distillation techniques. In an example, the first LLM is generated by tuning a locally hosted LLM like T5 leveraging domain-specific tuning in the industrial domain. The first LLM is generated using fine tuning of a locally hosted LLM (like T5) using a distillation approach to learn the behavior and style of a reference LLM (like ChatGPT). Advantageously, the technique for fine tuning ensures that the fine tuning of the first LLM is done on-premises and data need not be out to external sources (like ChatGPT API’s) mitigating privacy issues. Advantageously, the technique to generate the first LLM enables focused domain-specific training in an efficient manner and eliminating the need to use large corpus of data.
[0151] At step 1002, a set of questions are generated from the knowledge base using the domain taxonomy. The set of questions are prompts that are generated using the keywords and corresponding descriptions in the domain taxonomy. The set of questions are filtered for confidential and non-confidential entities. The questions that do not contain confidential entities are selected to be passed through the second LLM. The questions that contain confidential entities are further processed to replace the confidential entities with domain specific keywords / descriptions using the domain taxonomy. Further, these updated set of questions with non-confidential entities are ready to be passed through the second LLM or the public LLM. Notably, the set of questions with the confidential entities are only passed through the first LLM or the local LLM.
[0152] At step 1004, the second LLM is queried with the generated set of questions iteratively with different variations of the generated set of questions. A plurality of versions of the generated set of questions are created to query the second LLM. The second LLM is queried with each of the variation of questions and responses are recorded iteratively.
[0153] At step 1006, the first LLM is queried with the generated set of questions iteratively with different variations of the generated set of questions. A plurality of versions of the generated set of questions are created to query the first LLM. The first LLM is queried with each of the variation of questions and responses are recorded iteratively.
[0154] At step 1008, a distillation loss is iteratively calculated between the responses of the first LLM and the second LLM. The distillation loss is calculated based on a comparison of responses received from the first LLM and the second LLM for each iteration.
[0155] At step 1010, one or more parameters of the first LLM are tuned such that the distillation loss decreases after each iteration. The one or more parameters such as the weights of the nodes in architecture layers of the first LLM are tuned such that the distillation loss decreases after each iteration. In other words, the weights of the first LLM are tuned such that the response of the first LLM is close to the response of the second LLM. The distillation loss represents a deviation of the response from the first LLM when compared to the second LLM.
[0156] At step 1012, the first LLM is generated when the calculated distillation loss for each of the set of questions is below a predefined threshold. A threshold is defined for the distillation loss based on the deviation of the responses acceptable for a particular use case. In an example, for critical industrial scenarios like in a power generation plant, or a chemical plant, the threshold for distillation loss is low such that accuracy of the responses is increased. However, in non-critical scenarios like for condition monitoring of motors in a plant, the threshold is high. The first LLM is generated when the calculated distillation loss for each of the set of questions is below the predefined threshold.
[0157] Referring to FIG 11 , illustrated is a block diagram of an architecture 1100 for generating the first LLM, in accordance with an embodiment of the present invention. The architecture 1100 comprises the second LLM 1102, the first LLM 1104, the domain-specific data or knowledge base 1108, and a knowledge distillation model 1110. The second LLM 1102 is a public LLM such as ChatGPT, T5 (Text-to-Text-Transfer-Transformer), GPT-Neo, GPT-J, and GPT-NeoX, XLNet, Roberta - Robustly Optimized BERT Approach, DeBERT,, DistilBERT, etc. For the purpose of implementation of the invention, ChatGPT is selected as the second LLM 1102. The first LLM 1104 is a locally hosted LLM such as T5 (Text-to-Text-Transfer-Transformer), GPT-Neo, GPT-J, and GPT-NeoX, XLNet, Roberta: Robustly Optimized BERT Approach, DeBERT,, DistilBERT, etc. For the purpose of implementation of this invention, T5 model is selected as the first LLM 1104. Google’s T5 is one of the most advanced natural language models to date. It builds on top of previous work on T ransformer models in general. Unlike BERT, which had only encoder blocks, and GPT-2, which had only decoder blocks, T5 uses both. T5: Text-to-Text-Transfer-Transformer model proposes reframing all NLP tasks into a unified text- to- text-form at where the input and output are always text strings. This formatting makes one T5 model fit for multiple tasks. T5 was trained on a 700 GB dataset - on cleaned version of 04 (Colossal Clean Crawled Corpus) Dataset. It’s a high quality pre-processed English language corpus that they have made available for download. Also, the T5 model, pre-trained on C4, achieves state-of-the-art results on many NLP benchmarks while being flexible enough to be fine-tuned to a variety of important downstream tasks.
[0158] The knowledge base 1108 is generated using domain knowledge and Al models deployed in the industrial environment as explained above in FIG 7, FIG 8, and FIG 9.
[0159] The knowledge distillation model 1110 is a machine learning technique where the knowledge from a large, complex model (referred to as the teacher model, herein the second LLM 1102) is transferred to a smaller, simpler model (referred to as the student model, herein the first LLM 1104). This process involves training the first LLM to mimic the behavior of the second LLM by learning from its predictions or internal representations. In other words, knowledge distillation model 1110 is configured to captures and “distil” the knowledge in a complex machine learning model or an ensemble of models into a smaller single model that is much easier to deploy without significant loss in performance. “Knowledge distillation” refers to the process of transferring the knowledge from a large unwieldy model or set of models to a single smaller model that can be practically deployed under real-world constraints. The first LLM is trained to mimic the predictions or internal representations of the second LLM. Instead of directly learning from the original training data, the first LLM learns from the output of the second LLM. This is by minimizing the distillation loss function that measures the discrepancy between the second LLM predictions and the first LLM predictions on a set of training examples from the knowledge base 1108. Further, during training, the parameters of the first LLM are optimized to minimize the difference between its predictions and the second LLM predictions. After the distillation process, the first LLM is further fine-tuned on the knowledge base to improve its performance or adapt it to specific downstream tasks. The architecture for the knowledge distillation technique is explained further in FIG 12.
[0160] Referring to FIG 12, illustrated is a block diagram of another architecture 1200 for generating the first LLM using knowledge distillation, in accordance with another embodiment of the present invention. The architecture 1200 comprises a query database 1202 having domain-specific queries as generated from the domain taxonomy (not shown), the first LLM like T5 model 1204, the second LLM like ChatGPT 1206, a distillation loss calculator 1208. The first LLM 1204 is further fine-tuned on the domain-specific queries stored in a first database 1210 to generate the updated first LLM 1212. The responses from the updated first LLM 1212 are stored in a second database 1214. Advantageously, the first LLM 1204 trained on prompt-query pair on generic data learns to find / identify plausible response to a given query, which is generally through a semantic search in latent space. The generated response could either be discriminative (BertQA) or generative (T5), but the efficiency or coherency of generated responses is generally not on par with the model trained through RLHF(RL with Human feedback) like ChatGPT, thus, fine-tuning the first LLM with respect to the second LLM as teacher would lead to more natural and intuitive response, which when presented with a domain-specific data / prompt and query would lead to better responses. Advantageously, the fine-tune LLM 1212 is capable for answering closed domain question with respect to knowledge adaptation to domain prompt. Also, as the model is on-premise and finetuned internally, allows the users / operators to have more control over handling model biases as well as data privacy.
[0161] Referring to FIG 13, illustrated is a block diagram representing a system 1300 for answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention.
[0162] The system 1300 comprises a query database 1302 having domain-specific queries as generated from the domain taxonomy, the first LLM like T5 model 1304, the second LLM like ChatGPT 1306, a first query / response generator 1308 and a second query / response generator 1310, and a distillation loss calculator 1312, and a fine-tuned first LLM 1314. The first LLM 1304 is tuned to mimic the responses of the second LLM 1306 using response-based knowledge distillation. The distillation loss calculator 1312 calculates a deviation from the response between the second LLM 1306 and the first LLM 1304. The first LLM 1304 is iteratively fine-tuned on the responses until the output of the distillation loss calculator 1312 is below a predefined threshold loss value. The one or more parameters of the first LLM 1304 are optimized to generate the fine-tuned first LLM 1314.
[0163] The system 1300 an Al layer 1316 comprising RCA report module 1318, predictive maintenance module 1320, analytics module 1322, data sources 1324, and documentation databases 1326. The system 1300 further comprises a local domain data corpus 1328, a domain taxonomy 1330, and domain taxonomy generation module 1332. The architecture further comprises a prompt database 1334 and a prompt generation module 1336. The architecture further comprises a confidentiality assessment module 1338, a first report generation module 1340, open world domain data corpus 1342 and a public LLM 1344. The system 1300 further comprises a second report generation module 1346, an indexing module 1348 and an indexed report database 1350. The data from the Al layer 1316 is extracted and a local domain data corpus 1328 is generated. Based on this local domain data corpus 1328 and the domain taxonomy generation module 1332, the domain taxonomy 1330 is generated. The domain taxonomy 1330 comprises data with domain tags and corresponding general world descriptions. The prompt generation module 1336 is configured to generate the prompts or a set of open world questions using the domain taxonomy 1330. Once the open-world questions or prompts are generated, the prompts are stored in the prompt database 1334. The prompts from the prompt database 1334 are then passed to the confidentiality assessment module 1338 that filters out any questions that comprise confidential entities based on a classification model. All the prompts comprising only non-confidential entities are then queried to the public LLM 1344 (similar to the second LLM 1306) to retrieve domain knowledge available in the public domain. The responses from the public LLM 1344 are used to generate the open world domain data corpus 1342. Further, the data from the local domain data corpus 1328 and the open world domain data corpus 1342 is combined by the first report generation module 1340 to generate descriptive textual reports by using natural language generation (NLG) techniques. The first report generation module 1340 is capable of generating reports having responses to domain specific queries. Further, asset specific outputs from the Al layer 1316 are combined with the output of the first report generation module 1340 by the second report generation module 1346 to generate asset specific reports comprising textual description for each asset. Further, the reports are stored in the knowledge base 1348 by indexing the report using the indexing module 1348. The indexed reports are then stored in the knowledge base 1350.
[0164] The system 1300 further comprises a user interface 1358 for receiving queries from the user 1360.
[0165] The system also comprises an intent determination module 1356, a decision module 1354 and a query / prompt generation module 1352. The intent determination module 1356 is used to parse the query and determine the intent of the query input by the user 1360. The intent determination module is trained with labeled sample questions to distinguish between questions related to cement kiln and generic questions. It will also be trained to identify domain specific key words (“KN3TI3456A”) in the question. Based on the nature of the question, the intent identification model can trigger suitable actions such as querying the knowledge base, the first LLM , the second LLM, or trigger questions to the user if the original query is missing some information.
[0166] Once the intent of the query is identified, the decision module 1354 identifies whether the answer is available to the first LLM 1314 or not. The objective of the correct intent identification of the query is to route the user query to the appropriate data sources so as to provide an accurate answer. For example, imagine the knowledge base in the apparatus currently has information regarding a cement kiln and the events associated with it. If the user queries a question regarding the manufacturing process of paints, then the knowledge base won't have the answer for it. So, the intent identification model will identify that question as something that is out of the scope of the local knowledge base and will redirect the query to give an answer from open-source knowledge. If the answer is known to the first LLM, then the query is directed towards the prompt generation module 1352.
[0167] The prompt generation module 1352 is configured for generating prompts based on the query received. The prompt is then queried to the fine-tuned LLM 1314 to generate response to the query that is displayed on the user interface 1358. If the answer is not known to the fine-tuned first LLM 1314, and the intent is to query the local knowledge base, then one or more Al workflows are triggered in the Al layer 1316 in order to retrieve the relevant information. The knowledge base 1350 is updated with the output of the Al workflows, thereby enriching the fine-tuned LLM with updated knowledge for future queries.
[0168] In an exemplary implementation, the decision module 1354 may encounter below cases:
[0169] Case 1 : When the decision module 1354 decides from the intent of the query that the query is not known to the first LLM 1314, then the system triggers the Al workflow and then follow the process of using open world knowledge corpus and generating textual report - index it further and then generates prompt to respond to the query. In this case, the system also identifies and executed the correct Al workflows for retrieving correct response to the query. For instance - RCA for a particular event being an example query - the decision module 1354 decides if the knowledge corpus is already available for the event in the knowledge base 1348 - if yes it will go to case 2 - if no it will trigger the workflow for Al System which is the RCA module for the event. The RCA model output will then be transformed to detailed textual report using the open world knowledge corpus and then saved on DB (for future reference) and passed on to fine-tuned first LLM 1314 to answer the query.
[0170] Case 2: Taking a lead from Case 1 - if the data already exist in the knowledge base 1348 then, the decision module 1354 decides to fetch the data and pass the data as prompt to local hosted domain LLM (like T5) 1314 to answer the query.
[0171] Case 3: In case the query is generic enough and not has any confidential data and easel and case2 does not match with the intent of the query, then the decision module 1354 decides to query the second LLM 1306 for the answer. Case 4: In case the query comprises both confidential and non-confidential entities, then the decision module 1354 decides to trigger both case 1 and case 2 and then combining the other part of the answer from case 3. The completeness of the answer based on what is queried. For this case, it will be specific to use case and the decision module 1354 decides based on the use case and domain. The query is passed to the local corpus 1328 to see if we have the answer for the query in completeness. In case, the response is not complete - it will split the query and then pass the other half of the query to the second LLM 1306 as in case 3. The system 1300 will combine the response from both the LLMs to provide a response to the user query.
[0172] In an exemplary implementation, consider a cement kiln which is a critical equipment in the manufacturing of cement. To create the knowledge base 1348, the Al layer 1316 to extract information regarding the design, operation of the kiln through its documentation. From these documents, a domain taxonomy 1330 is generated. For example, every equipment sensor in the plant will have a unique identifier associated with it. The domain taxonomy 1330 will have all the list of sensor names and a small description of what it means. Consider a sensor named “KN3TI3456A”. The first two letters KN might refer to the equipment which is cement kiln, the number 3 will mean the kiln number 3. The letter Tl refers to a temperature indicator and the numbers 3456 will refer to the location of the sensor and the letter A means that there are more than one sensor associated with it. So after the creation of domain taxonomy, the ID “KN3TI3456A” will mean the outlet temperature sensor for kiln 3. The next step after the creation of domain taxonomy is augmentation of the data with publicly available knowledge. For example, publicly available information regarding the common issues, maintenance recommendations etc. for the temperature sensors that are located at the outlet of a kiln is gathered via public LLMs. The augmented knowledge base can then take in information regarding the specific events that has happened in the equipment. This information can come from analytics / AI tools like iPAE (reference to a previously filed application, mention application number) as well as past maintenance history of the equipment. For example, if the causality module in iPAE has identified the root causes behind an anomaly of the kiln, then that information along with supplementary information like the alarms at the time of event, time, date etc. is stored in the knowledge base.
[0173] Once the user asks a query and the intent identification model 1356 recognizes it comprises a relation to the cement kiln the information retrieval system will fetch the relevant section of the local knowledge base which will have answer to this question. This section along with the query can then be sent as a prompt for the local LLM 1314 to answer. This will reduce the need for sending huge prompts to the local LLMs and improves the efficiency of the process. For example, the user wants to know the root cause for an anomaly that occurred in kiln 3 on 30thDec 2022. The data retrieval system which currently uses vectorDB will launch a vector search to identify sections of the text that has similar content as that of the question. In the example mentioned, the data retrieval system will search for the RCA report on the kiln for the date of 30thDec 2022 and it will identify the section in the report that has information regarding the cause of the event. This knowledge is then sent to the LLM for updating the same. In another case, when the answer is not found in the local knowledge base, then an Al workflow is triggered in the system to generate RCA report for the anomaly that occurred in kiln 3 on 30thDec 2022. Once the RCA report is generated, the knowledge base 1348 is updated with the information corresponding to the anomaly, the first LLM 1304 is triggered is provide a response to the query.
[0174] Referring to FIG 14, illustrated is an exemplary system 1400 for answering a natural language user query in an industrial environment, in accordance with another embodiment of the present invention. The system 1400 illustrated here can be understood as an industrial intelligence layer combined with knowledge from a the trained first LLM (referred to as IIL_LLM_Agent in following examples) capable of providing accurate answers to the natural language query from users. In an exemplary implementation, the local data corpus 1408 is generated by extracting information from one or more sources like alarm reports 1402, RCA reports 1404, predictive model 1406. Further, a domain taxonomy 1410 is generated using the local data corpus 1408 and the domain knowledge 1412. A prompt generation module 1416 generates a set of question to be queried to the second LLM 1416 using open world data corpus 1414. The response from the second LLM 1416 is combined with output of the Al workflows 1418 to generate the knowledge base 1420 comprising full reports with textual descriptions to the set of questions. The system 1400 also comprises a user interface 1426, and the first LLM 1428 enriched with the knowledge base 1420. The system further comprises an intent identification module 1424 and a decision module 1422.
[0175] Some examples of how the system 1400 provides responses to the queries are provided below:
[0176] Example Scenario 1 :
[0177] 1. Plant Operator enters the query to the 11 L_LLM Agent “what is the root cause of the failure (with fault ID) on May 7th2022 for the Heat exchanger 11?”
[0178] 2. The system provides the causal analysis of the fault and provide the insights.
[0179] 3. IIL_LLM_Agent communicates with the second LLM (e.g. ChatGPT) by feeding the output of causal analysis to generate a root cause analysis report.
[0180] 4. The plant operator now gets both causal analysis as an output and also a root cause analysis report based on public knowledge.
[0181] Example Scenario2: 1. Plant Operator asks the question to the IIL_LLM_Agent “What is the possibility of an anomaly in my chiller?”
[0182] 2. This Chiller is being monitored by chiller monitoring and prediction Al model deployed in the environment, so it has all the available data for the chiller.
[0183] 3. The industrial predictive analytics engine (iPAE) will run anomaly detection workflow and optimization workflow.
[0184] 4. IIL_LLM_Agent will provide the insight - anomalies, optimization set points.
[0185] 5. The IIL_LLM_Agent will then call the second LLM with the anomalies, optimization set points to generate an analysis report and recommendations.
[0186] Example Scenarios:
[0187] 1. Plant Operator asks the question to the IIL_LLM_Agent “What is the possibility of operating my building?”
[0188] 2. This Building is being monitored by Desigo CC system so it has all the available data for the building.
[0189] 3. The industrial predictive analytics engine (iPAE) will run forecasting workflow, RL workflow.
[0190] 4. IIL_LLM_Agent will provide the insight what is the current status of rooms in the building / building as a whole in terms of temperature, humidity, occupancy, other metrics.
[0191] 5. The IIL_LLM_Agent will then call the second LLM with the anomalies, optimization set points to generate a analysis report and recommendations to run the building more efficiently
[0192] Example Scenario 4:
[0193] 1. The process operator asks a question: “Will there be an anomaly on the heat exchanger in my plant in the next week?”
[0194] 2. The IIL_LLM_Agentwill then break the question into different actions that can be executed in iPAE.
[0195] 3. It will first initiate a training workflow to train a model to detect anomaly in a HX.
[0196] 4. It will use neural architecture search to identify the best model architecture. Once different models are trained, it will identify the best model for anomaly detection.
[0197] 5. Then it will need to initiate the inferencing workflow for the past week and use it to check if an anomaly is going to happen.
[0198] 6. For this we need to train the model on a handbook of best data science practices along with the API document for iPAE to understand the question and process it. 7. Once the models are trained, iPAE provides the outputs and then the response is generated based on the output of iPAE by passing the output to the IIL_LLM_Agent
[0199] Example Scenarios:
[0200] 1. Siemens Sales / Marketing person / 3rdparty person looking to buy Siemens equipment; asks the question to the IIL_LLM_Agent “What are the different types of the SCADA system being offered from Siemens?”
[0201] 2. IIL_LLM_Agent will query its “dictionary of Siemens Catalogue / products / available public unrestricted material” and provide a clear answer which includes <product name>, <product type>, <pricing as available>, <availability>, <Key features / highlights>, <Social media material video, other info>, <contact information based on the region>, <License terms and conditions>
[0202] 3. It will further provide any necessary paperwork required such as <typical NDA>, <how to approach the suppliers <other>
[0203] 4. All the relevant information is structured in a natural language response and provided to the user. With this information any person keen on buying Siemens equipment will have all information at hand to make a quick decision.
[0204] Referring to FIG 15, illustrated is an exemplary method flow 1500 for answering a natural language user query in an industrial environment, in accordance with another embodiment of the present invention. The user 1502 asks a query “Tell me about maintenance steps for FCV” on a user interface 1504. The user query is provided to an intent identification module 1506 to identify the intent of the user query. The identified intent is then provided to a decision module 1508 to identify an accurate workflow to be executed. The decision module 1508 determines that the intent of the user query is to query the local knowledge base. The query is then passed through a parser 1510 to identify the entities from the query. In this case, the identified entity is “FCV”. The entity “FCV” and the intent “maintenance” is provided to a database retrieval system 1512. The database retrieval system 1512 selects data sources relevant for the query and provided to a data extraction module 1514. The data extraction module 1514 extracts relevant information from the selected data sources based on the intent of the query. In order to extract relevant information from the knowledge base based on optimal retrieval approach using a sophisticated pipeline for data retrieval. The data extraction module 1514 extracts all relevant documents associated with context of the query. The specific section of the context is extracted to make the response more precise and optimal. Further, data extraction module 1514 combines the context and rank based on the query context. The data extracted from the data extraction module 1514 is passed through a local LLM 1516 for contextualization and generating the response to the query, which is then provided to the user 1502.
[0205] Referring to FIG 16, illustrated is an exemplary method flow 1600 for answering a natural language user query in an industrial environment, in accordance with another embodiment of the present invention. At step 1604, the user 1602 asks a query “Tell me about the anomaly at 27 / 01 / 2024” on a user interface. At step 1604, the intent of the user query is determined. At step 1608, a decision is made on the identified intent based on the user query to determine an accurate workflow to be executed. Herein, it is that the user query is an event-based query. At step 1610, the query is then passed through a parser to identify the entities from the query. In this case, the identified entity is “anomaly” & “27 / 01 / 2024”. At step 1612, the entity “anomaly” and the intent “27 / 01 / 2024” are provided to a context retrieval system. The context retrieval system selects data sources relevant for the query. At step 1614, a decision is made by the decision module to check if the answer to the user query is available in the knowledge base. If the decision module decides that the answer is available to the first LLM, then at step 1616 context is extracted from the knowledge base and passed through the first LLM to generate the response. The structured response is then provided to the user 1602. If the decision module decides that the answer is not available to the first LLM, then at step 1618 the Al workflow to be executed to generate the answer is identified. At step 1620, the selected Al models in the Al workflow is triggered to detect the anomaly in the time ranges provided in the query. At step 1622, the Al model is executed to generate the output and more context for anomaly such as causality, outlier detection, etc. At step 1624, the output of the Al model is received. At step 1626, the anomaly report is generated. At step 1628, further inferencing from the knowledge base is added for the anomaly such as preventive and maintenance steps for causal variable in the anomaly. At step 1630, the report is stored in the knowledge base. Once the report is generated and stored, the step 1616 is executed and query along with the context from the anomaly report is passed through the first LLM to generate the response for the user 1602.
[0206] Referring to FIG 17, illustrated is a Graphical User Interface (GUI) 1700 of the user device for answering a natural language user query in an industrial environment, in accordance with an embodiment of the present invention. The GU1 1700 represent a chatbot like interface for querying the first LLM enriched with knowledge base for a particular industrial domain. As shown, the user is querying the first LLM about health of chillers, water temperature, outcomes of fault in a condenser, and so on. The first LLM provides descriptive human-like responses to the user.
[0207] The present invention provides a system for answering natural language query in an industrial environment in an accurate and efficient manner. This is achieved by locally training a LLM on domain data extracted from one or more sources. The invention provides a robust mechanism and method for deriving insights into any query related to the plant / process / equipment being monitored. Thereby, improving the performance and efficiency of decision making in the industrial environment. Furthermore, the invention also provides a mechanism to easily structure all publicly available data pertaining to an organization and provide the same to the user when queried in a structured manner.
[0208] The present invention serves as an industrial co-pilot for the plant operators, maintenance engineers, supervisors, or factory personnel who would like to interrogate the system from time to time to get insight on the state of the system, get specific answers regarding the problems being faced, specific KPI values at that time, causal analysis of the anomalies already happened, potential solutions to the problems already faced. Moreover, the present system is also capable of providing publicly available information and structure it with the context of the query along with additional information on the equipment deployed ranging from model make, type, age, ideal operating conditions to complete process in which the equipment are deployed, operated in, potential weaknesses and inefficiencies in the process. This invention enables the users / plant operators monitor the inventories, supply chain of raw materials, storage, logistics and others in s time-efficient and seamless manner.
[0209] The present invention provides an in-house large language model (LLM), herein, the first LLM capable of accurately answering queries in the industrial environment. The first LLM is trained using response-based knowledge distillation technique thereby a huge amount of data is not required for training the in-house first LLM. Advantageously, the approach of training is LLM agnostic and can be used to and applied on any existing LLM. As a result we can create a distillation chain to learn behavior and style from a reference LLM (e.g. ChatGPT) and have multiple intermediate LLMs(e.g. T5) that can learn specific aspects from the reference LLM. Furthermore, the approach works on the principle of on-premises distillation. Hence there is no need to use large corpus of data or even use that data to re-train the LLM. Hence, privacy remains intact as the data always remains on-premise. The approach can be applied to learn multiple aspects from multiple LLMs. For example, it is possible to perform a domain specific finetuning of an open source LLM like T5, with respect to answering style from ChatGPT, mathematical calculations from Minerva, and so on. As a result, the system is capable of combining multiple models together to create a diversified distillation, thereby enhancing the accuracy of the first LLM. Furthermore, the first LLM is a lightweight model and hence suitable for deployment on resource-constrained devices or environments. Furthermore, it will be appreciated that the fine-tuning of the system is done on-premise thereby eliminating any data privacy issues. Moreover, the training / fine tuning of the first LLM is done on the domain data eliminating the need to use large corpus of data.
[0210] The present invention focusses on intent identification of the query. Advantageously, the objective of the correct intent identification of the query is to route the user query to the appropriate data sources so as to provide an accurate answer. Advantageously, the system is capable of intelligently deciding the which workflow to be followed for accurately answering the natural language query.
[0211] Furthermore, once the query is received, the system is capable of automatically generate prompts for querying the LLMs. Advantageously, the prompts are generated for accurately querying the first LLM and / or second LLM. Furthermore, the correct prompt generation eliminated the need for sending huge prompts to the first LLM and improves the efficiency of the response. It will be appreciated that the prompts are generated by leveraging information in the domain taxonomy. Advantageously, the domain taxonomy is utilized to get more open world information which will enable the system to make more richer inferences. Furthermore, the descriptions in the domain taxonomy also ensure accurate generations of prompts.
[0212] The present invention also aides in automatically running multiple Al workflows based on potential future queries and store the same in the knowledge base in order to enrich the first LLM. Advantageously, the set of questions generated vary significantly and aims to include all potential queries such that the knowledge base is accurately enriched with correct information. Advantageously, the information in the knowledge base is stored in in a structured manner to facilitate easy navigation, retrieval, and access to information. Information is typically categorized into topics, sections, or modules based on the subject matter, making it easier for users to locate relevant content.
[0213] Furthermore, since the present system is capable of automatically triggering Al workflows based on the intent of the query, the system is capable of answering queries that are not present in the knowledge base or in the training database of the first LLM, thereby making sure that new queries are also answered accurately.
[0214] The present invention provides a mechanism to automatically assess the confidentiality level of the queries. Advantageously, the confidentiality check ensures that the queries having non- confidential entities are only forwarded to the second LLM or the public LLM. Advantageously, in order to maintain confidentiality of customer data or plant data, the natural language query comprising confidential entities are only directed to the first LLM i.e. the local LLM enriched with knowledge base. Advantageously, the system intelligently understands the intent of the user and structures a response most suitable for the user to interpret the response.
Claims
Patent claims l / We claim:
1. A method (600) for answering a natural language user query in an industrial environment, the method comprising: receiving, by the processing unit (135), the natural language query from the user (1360), the natural language query pertaining to a particular domain of the industrial environment; parsing, by the processing unit (135), the natural language query to determine one or more entities and an intent of the natural language query; generating, by the processing unit (135), a first set of prompts from the natural language query based on a defined taxonomy pertaining to the particular domain of the industrial environment and the determined intent of the user wherein the domain taxonomy (917, 1330, 1410) is a hierarchical structure comprising keywords in the domain mapped to corresponding descriptions pertaining to the domain; determining, by the processing unit (135), if an answer to the natural language query is known to a first LLM (102, 412, 1104, 1204, 1304, 1428) based on the determined intent of the natural language query, wherein the first LLM is specifically trained on the particular domain of the industrial environment and enriched with a knowledge base (112, 434, 532, 702, 902, 1108, 1420), wherein the knowledge base comprises information from one or more data sources (114- 1 to 114-N, 206-1 to 206-N, 912, 1324) combined with domain knowledge (1412) from a second LLM (104, 1102, 1206, 1306); querying, by the processing unit (135), the natural language query to the first LLM () using the generated prompts, if it is determined that the answer is known to the first LLM; generating, by the processing unit, a response to the natural language query based on an output of the first LLM (102, 412, 1104, 1204, 1304, 1428).
2. The method (600) as claimed in claim 1 , wherein the response to the natural language query is output in at least one of the formats: textual, images, videos, reports, and a combination thereof.
3. The method (600) as claimed in any of the claims 1 or 2, further comprising identifying a confidentiality level of the natural language query of the user based on the determined entities of the natural language query identified based on the comparison with a database (150, 710, 1210, 1214) comprising a plurality of keywords that are annotated as confidential.
4. The method (600) as claimed in any of the preceding claims, wherein answering the received natural language query comprises querying to the first LLM (102, 412, 1104, 1204, 1304, 1428)using the generated prompts, if it is identified that the natural language query comprises one or more entities marked as confidential in the database (150, 710, 1210, 1214).
5. A method (600) as claimed in any of the preceding claims, wherein answering the received natural language query comprises: executing, by the processing unit (135), one or more Al workflows in the industrial environment based on the intent of the natural language query, if it is determined that the answer to the natural language query is not known to the first LLM (102, 412, 1104, 1204, 1304, 1428), and the natural language query comprises one or more entities marked as confidential in the database (150, 710, 1210, 1214); updating, by the processing unit (135), the knowledge base (112, 434, 532, 702, 902, 1108, 1420) with output of the one or more Al workflows triggered; generating, by the processing unit (135), a second set of prompts based on the domain taxonomy (917, 1330, 1410); querying, by the processing unit (135), the natural language query to the first LLM enriched with the updated knowledge base using the second set of prompts; and generating, by the processing unit (135), a natural language response to the natural language query based on an output of the first LLM (102, 412, 1104, 1204, 1304, 1428) enriched with the updated knowledge base.
6. The method (600) as claimed in any of the preceding claims, wherein answering the received natural language query comprises querying to the second LLM (104, 1102, 1206, 1306); using the generated prompts, if it is identified that: the natural language query does not comprise one or more entities marked as confidential in the database (150, 710, 1210, 1214; and the answer to the natural language query is not known to the second LLM (104, 1102, 1206, 1306).
7. The method (600) as claimed in any of the preceding claims, wherein answering the received natural language query identified as comprising entities marked as both confidential and non- confidential comprises: identifying, by the processing unit (135), one or more confidential entities in the natural language query; generating, by the processing (135), a third set of prompts by replacing the confidential entities with general descriptions using the domain taxonomy (917, 1330, 1410), and a fourth set of prompts by retaining the confidential entities;querying, by the processing unit (135), the second LLM (104, 1102, 1206, 1306) with the third set of prompts and the first LLM (102, 412, 1104, 1204, 1304, 1428 with the fourth set of prompts; structuring, by the processing unit (135), a response from the first LLM (102, 412, 1104, 1204, 1304, 1428) and the second LLM (104, 1102, 1206, 1306) in order to generate a natural language response to the natural language query.
8. The method as claimed in any of the preceding claims, wherein generating the knowledge base (112, 434, 532, 702, 902, 1108, 1420) pertaining to the particular domain comprises: acquiring, by the processing unit (135), information from one or more data sources (114- 1 to 114-N, 206-1 to 206-N, 912, 1324) in the industrial environment, wherein the one or more data sources comprises at least one of: sensors, RCA reports, prediction models, one or more Al models, database including asset specification, predetermined knowledge graphs, images, videos, audios, or a combination thereof; generating, by the processing unit (135), domain taxonomy based on the extracted information, wherein the domain taxonomy (917, 1330, 1410) comprises a set of keywords and corresponding descriptions extracted from the acquired information; generating, by the processing unit (135), a set of questions using the domain taxonomy (917, 1330, 1410) based on a predefined template; identifying, by the processing unit (135), a confidentiality level of the generated set of questions; querying, by the processing unit (135), the second LLM (104, 1102, 1206, 1306) using the generated set of questions that do not contain confidential entities; triggering Al workflows, by the processing unit (135), in the one or more Al models for the set of questions that contain confidential entities; generating, by the processing unit (135), the knowledge base by combining output from Al workflows and the second LLM (104, 1102, 1206, 1306).
9. A method (600) as claimed in any of the preceding claims, wherein generating the first LLM (102, 412, 1104, 1204, 1304, 1428) comprises: generating, by the processing unit (135), a set of questions from the knowledge base using the domain taxonomy (917, 1330, 1410; querying, by the processing unit, the second LLM (104, 1102, 1206, 1306) with the generated set of questions iteratively with different variations of the generated set of questions; querying, by the processing unit, the first LLM (102, 412, 1104, 1204, 1304, 1428) with the generated set of questions iteratively with different variations of the generated set of questions;iteratively calculating, by the processing unit (135), a distillation loss between the responses of the first LLM (102, 412, 1104, 1204, 1304, 1428) and the second LLM (104, 1102, 1206, 1306); tuning, by the processing unit (135), one or more parameters of the first LLM (102, 412, 1104, 1204, 1304, 1428) such that the distillation loss decreases after each iteration; and generating, by the processing unit (135), the first LLM (102, 412, 1104, 1204, 1304, 1428) when the calculated distillation loss for each of the set of questions in below a predefined threshold.
10. A method (600) as claimed in any of the preceding claims, wherein generating prompts from the natural language query for querying the first LLM (102, 412, 1104, 1204, 1304, 1428) comprises: retrieving, by the processing unit (135), information that is similar to the keywords in the natural language query from domain taxonomy (917, 1330, 1410); and generating prompts for querying the first LLM (102, 412, 1104, 1204, 1304, 1428) based on the descriptions extracted from the domain taxonomy (917, 1330, 1410).
11. The method (600) as claimed in any of the preceding claims, wherein generating domain taxonomy (917, 1330, 1410) comprises: identifying, by the processing (135) unit, a plurality of keywords from one or more sources pertaining to a particular domain; extracting, by the processing unit (135), a textual description of each of the keywords from the one or more sources; defining, by the processing unit (135), a relationship between the keywords and the textual description by annotating the keywords; and generating, by the processing unit (135), the domain taxonomy (917, 1330, 1410) with identified keywords and corresponding textual description based on a defined hierarchy.
12. The method (600) as claimed in any of the preceding claims, wherein determining an intent of the natural language query comprises: parsing, by the processing unit (135), the natural language query to extract one or more entities and determine semantic and syntactic relationships therebetween; and determining, by the processing unit (135), an intent of the natural language query based on a machine learning model trained on a historical database of one or more entities annotated with their corresponding intent.
13. An apparatus (106) for answering a natural language user query in an industrial environment, the apparatus comprising: one or more processing units (135); and a memory unit (140) communicatively coupled to the one or more processing units, wherein the memory unit comprises a module stored in the form of machine-readable instructions executable by the one or more processing units (135), wherein the industrial intelligence module (165) is configured to perform method steps according to any of the claims 1 to 12.
14. A system (100) for answering a natural language user query in an industrial environment, the system (100) comprising: one or more client devices (110) for inputting the natural language query; a first LLM (102, 412, 1104, 1204, 1304, 1428) trained specifically on a particular domain and enriched with a knowledge base (112, 434, 532, 702, 902, 1108, 1420), wherein the first LLM is a local LLM, the first LLM being communicatively coupled to the one or more client devices; a second LLM (104, 1102, 1206, 1306), wherein the second LLM is a public LLM; one or more sources (114-1 to 114-N, 206-1 to 206-N, 912, 1324) for acquiring information pertaining to one or more asset installed in the industrial environment; and an apparatus (106) according to claim 12, communicatively coupled to the one or more client devices, (110) answering a natural language user query in an industrial environment as claimed in any of the claims 1 to 12.
15. A system (200) as claimed in claim 14, further comprising: a plurality of first LLMs (202-1 to 202-N) trained specifically on a corresponding plurality of industrial domains; a plurality of knowledge bases (204-1 to 204-N) generated based on specific domain knowledge, Al models, manuals, text documents for the plurality of industrial domains, the knowledge bases (204-1 to 204-N) of the plurality of knowledge bases communicatively coupled to corresponding LLMs of the plurality of first LLMs (202-1 to 202-N); and a domain selection module configured for: identifying a domain of the received natural language query; and selecting a suitable first LLM and corresponding knowledge base based on the identified domain of the received natural language query.
16. A computer-program product having machine-readable instructions stored therein, which when executed by one or more processing units (135), cause the processing units to perform a method according to any of the claims 1 to 12.
17. A computer-readable storage medium comprising instructions which, when executed by one or more processing units, cause the one or more processing units to perform a method according to any of the claims 1 to 12.
Citation Information
Patent Citations
System for handling workplace queries using online learning to rank
US11803556B1
Methods and systems for improved document processing and information retrieval
WO2024015323A1