Verification and validation of large language models with retrieval-augmented generation in industrial automation systems
A comprehensive verification and validation framework for LLMs with RAG systems in industrial automation addresses inconsistent outputs and real-time performance challenges, enhancing reliability and safety through integrated testing components and adaptive strategies.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SIEMENS AG
- Filing Date
- 2025-11-05
- Publication Date
- 2026-05-15
AI Technical Summary
The deployment of large language models (LLMs) with retrieval-augmented generation (RAG) systems in industrial automation environments faces unique challenges such as inconsistent outputs, unreliable responses, real-time performance requirements, environmental adaptability, and lack of comprehensive verification and validation frameworks, which can lead to safety hazards and operational inefficiencies.
A comprehensive verification and validation framework that integrates various testing components, including model consistency verification, data quality verification, performance verification, and robustness verification, with an API server coordinating these components to ensure reliable operation and adaptability across industrial environments.
Enhances defect detection capabilities and maintains compliance with industrial safety and reliability standards by systematically addressing complex failure modes and real-time performance requirements, ensuring transparent and auditable decision-making.
Smart Images

Figure US2025054175_15052026_PF_FP_ABST
Abstract
Description
202419876VERIFICATION AND VALIDATION OF LARGE LANGUAGE MODELS WITH RETRIEVAL-AUGMENTED GENERATION IN INDUSTRIAL AUTOMATION SYSTEMSBACKGROUND
[0001] Large language models (LLMs) are advanced artificial intelligence systems trained on vast amounts of text data to understand and generate human-like text. These models, such as GPT (Generative Pre-trained Transformer) and similar architectures, have demonstrated capabilities in natural language understanding, reasoning, and generation across diverse domains. LLMs can process complex queries, provide detailed explanations, and assist in decision-making processes by leveraging patterns learned from extensive training datasets. Retrieval-augmented generation (RAG) is a technique that enhances LLMs by combining their generative capabilities with external knowledge retrieval systems. In a RAG system, relevant information is first retrieved from external databases, knowledge graphs, or document repositories based on the input query, and then this retrieved information is provided as context to the LLM to generate more accurate, up-to-date, and factually grounded responses. This approach addresses limitations of standalone LLMs, such as knowledge cutoff dates, hallucination of facts, and inability to access real-time information.
[0002] Industrial automation systems traditionally rely on deterministic control logic and predefined decision trees. It is recognized herein that the incorporation of LLM and RAG technology can enable more sophisticated decision-making capabilities across various levels of manufacturing operations, from device-level control to enterprise-level management. For example, in industrial settings, LLM and RAG systems might assist operators with complex troubleshooting procedures, provide intelligent recommendations for process optimization, interpret sensor data in context, and facilitate natural language interfaces for system monitoring and control.
[0003] It is recognized herein, however, that the deployment of LLM and RAG systems in industrial automation environments presents unique technical challenges that are not adequately addressed by conventional software validation approaches. For example, various industrial automation systems demand exceptionally high levels of reliability, safety, and real-time performance, as failures can result in equipment damage, production losses, safety hazards, or regulatory compliance violations. Furthermore, technical challenges arise across multiple dimensions throughout the development and deployment lifecycle. Current verification and202419876 validation approaches for Al systems typically address various challenges in isolation, focusing on individual aspects such as model accuracy testing, performance benchmarking, or basic robustness testing. These isolated approaches, however, fail to capture the complex interdependencies between different aspects of system performance and reliability.BRIEF SUMMARY
[0004] Embodiments of the invention address and overcome one or more of the described- herein shortcomings by providing methods, systems, and apparatuses that can perform various verification and validation tests on a Large Language Model (LLM) with Retrieval- Augmented Generation (RAG) (LLM and RAG) system in various industrial automation environments.
[0005] In an example aspect, a computing system of an industrial automation system includes a large language model (LLM) and retrieval-augmented generation (RAG) system integrated with the industrial automation system that comprises sensors, controllers, and actuators. The system further includes a plurality of verification and validation components. Each component is configured to perform a different type of verification comprising a model consistency verification, data quality verification, performance verification, and robustness verification. The system can further include an API server configured to use results from a first verification and validation component of the plurality to inform testing parameters of a second verification and validation component of the plurality. The API server is further configured to generate targeted test cases based on the results from the first verification and validation component. The system can define a real-time monitoring system configured to continuously monitor performance of the LLM and RAG system and provide feedback for ongoing verification and validation.
[0006] In another example aspect, a Large Language Model (LLM) with Retrieval-Augmented Generation (RAG) (LLM and RAG system) in industrial automation can be cross-validated by obtaining a plurality of system inputs to the LLM and RAG system. A computing system can perform explainability analysis on the plurality of system inputs to identify a critical input feature, so as to define an identified critical input feature, wherein the identified critical input feature affects decisions of the LLM. The system can perturb the identified critical input feature while leaving other input features unchanged, so as to generate a focused test case. The focused test case can be executed to measure variations in response of the LLM and RAG202419876 system. The identified critical input feature can be modified, so as to detect instabilities in behavior of the LLM. The system can iteratively refine the focused test case based on the instabilities to create targeted test scenarios that enhance defect detection capabilities. Verification requests can be received from a plurality of verification components comprising a model consistency tester, a data consistency checker, a generation robustness tester, and a quality of service analyzer. In some cases, the plurality of verification components can be configured with respective outputs from a plurality of validation components comprising an explainability analyzer and a reliability lifecycle manager. The execution of the plurality of verification components can be coordinated through an API server, wherein results from a first verification component of the plurality of verification components inform testing parameters of a second verification component of the plurality of verification components. In various examples, real-time performance metrics of the LLM and RAG system is monitored during operation with industrial controllers, sensors, and actuators. Based on the monitoring, integrated assessment results can be generated by combining outputs from the plurality of verification components. The integrated assessment results can be compared against predefined performance baselines for industrial automation requirements. Feedback can be provided to the LLM and RAG system to maintain compliance with industrial safety and reliability standards.
[0007] In another example aspect, baseline performance metrics for the LLM and RAG system are established using a baseline checker component. The system can detect performance degradation by comparing current system outputs against the established baseline performance metrics. In particular, for example, the system can automatically trigger focused verification testing when performance degradation exceeds predetermined thresholds. The baseline performance metrics can be updated based on validated system improvements. In some examples, a first verification on the LLM and RAG system is executed to generate first verification results. The system can extract key insight parameters from the first verification results that indicate critical system dependencies, so as to define extracted key insight parameters. A second verification can be configured using the extracted key insight parameters as input constraints, wherein the second verification targets testing scenarios informed by the first verification results. The system can execute the second verification to generate enhanced test cases that are more targeted than either the first verification or the second verification achieve independently from each other. The system can also measure system performance changes resulting from the enhanced test cases. The system can identify failure modes revealed202419876 by the system performance changes, so as to define identified failure modes; and store the identified failure modes in a knowledge base for future test case generation.
[0008] In yet another example aspect, the system can receive operational data from industrial automation infrastructure including sensor readings, controller states, and actuator positions. The system can process the operational data through the LLM and RAG system to generate industrial control recommendations, so as to define generated recommendations. The system can simultaneously execute multiple verification analyses on generated recommendations. For example, the system can measure response latency of the LLM and RAG system against industrial real-time requirements using a quality of service analyzer. Additionally, or alternatively, the system can validate that the generated recommendations meet safety and reliability criteria for industrial automation applications. The system can detect anomalies in behavior of LLM and RAG system by comparing current performance against historical baseline patterns. In various examples, the system automatically adjusts system parameters when performance metrics fall below industrial automation thresholds.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS
[0009] The foregoing and other aspects of the present invention are best understood from the following detailed description when read in connection with the accompanying drawings. For the purpose of illustrating the invention, there is shown in the drawings embodiments that are presently preferred, it being understood, however, that the invention is not limited to the specific instrumentalities disclosed. Included in the drawings are the following Figures:
[0010] FIG. 1 is a block diagram of an example application domain, in particular an example industrial automation or production system, that can include or be integrated with a large language model (LLM) and retrieval-augmented generation (RAG) system in accordance with an example embodiment.
[0011] FIG. 2 is a block diagram of the LLM and RAG system that can include various verification and validation components, in accordance with an example embodiment.
[0012] FIG. 3 is a flow diagram that depicts example operations that can be performed by the LLM and RAG system shown in FIG. 2, in accordance with various embodiments.
[0013] FIG. 4 depicts a computing environment within which embodiments of the disclosure may be implemented.202419876DETAILED DESCRIPTION
[0014] As an initial matter, it is recognized herein that various technical challenges are involved in deploying large language models (LLM) with retrieval-augmented generation (RAG) systems in industrial automation environments. For example, model accuracy and consistency can present significant challenges during development and testing stages. LLMs can produce inconsistent outputs for similar inputs due to their probabilistic nature, and RAG systems may retrieve irrelevant or outdated information, leading to unreliable responses. In industrial settings, such inconsistencies can result in conflicting operational guidance or erroneous control decisions.
[0015] By way of further example, real-time performance requirements can create substantial constraints during integration and continuous testing phases. Industrial automation often requires millisecond-level response times for critical control functions. The computational overhead of LLM inference and the latency introduced by retrieval operations can compromise the real-time responsiveness essential for safety-critical applications. Environmental adaptability can also become crucial during staging and stress testing phases. Various industrial environments can define harsh conditions including temperature fluctuations, electromagnetic interference, network instability, and varying computational loads, such that LLM and RAG systems should maintain consistent performance across these variable conditions without degradation. In some cases, LLMs may experience performance drift over time as external data sources change, operational conditions evolve, or model parameters degrade. This drift can lead to gradually declining performance that may not be immediately apparent but can eventually result in system failures. Further still, user operability and interpretability requirements can become critical in post-deployment phases. Industrial operators often need to understand and trust the recommendations provided by LLM and RAG systems. The black box nature of many Al systems can create challenges for operators who make critical decisions based on system outputs, particularly in safety-critical situations where the reasoning behind recommendations should be transparent and auditable. Additionally, it is recognized herein that such reasoning should be comprehensively tested and explainable.
[0016] It is further recognized herein that existing solutions often lack comprehensive frameworks that can systematically address the critical aspects of LLM and RAG deployment in industrial settings, such as those described above. Moreover, conventional validation approaches often do not provide mechanisms for leveraging insights from one type of testing to202419876 enhance the effectiveness of other testing methodologies, resulting in suboptimal detection of potential failure modes. The absence of integrated verification and validation frameworks specifically designed for the unique requirements of industrial automation environments can create risks for organizations seeking to deploy LLM and RAG technologies in production settings. These risks can include, for example and without limitation, potential safety hazards, regulatory compliance failures, operational inefficiencies, and loss of operator confidence in Al-assisted systems.
[0017] In accordance with various embodiments, a system defines a comprehensive verification and validation framework specifically designed to address the complex challenges of integrating LLM and RAG systems into industrial automation environments, providing systematic risk mitigation across all stages of the development and deployment lifecycle while ensuring reliable, safe, and effective operation in dynamic industrial settings. For example, embodiments described herein can integrate diverse types of testing (e.g., robustness, consistency, QoS), while also enabling bidirectional cross-leveraging between procedural layers of verification and validation. Unlike conventional systems that treat verification and validation as isolated or sequential processes, the disclosed framework allows outputs from validation components (e.g., explainability analysis, lifecycle monitoring) to inform verification procedures (e.g., robustness testing, baseline checking), and visa-versa. This integration can enhance the detection of complex failure modes and supports adaptive, context-aware testing strategies tailored to industrial automation environments.
[0018] Referring initially to FIG. 1, an example industrial system 100 includes an office or corporate IT network 102 and an operational plant or production network 104 communicatively coupled to the IT network 102. It will be understood that the system 100 is illustrated and simplified as an example, and LLM and RAG systems can be deployed within a variety of industrial systems in accordance with various embodiments, and all such systems are contemplated as being within the scope of this disclosure. For example, and without limitation, embodiments can be implemented in an operational technology system, energy generation system (e.g., wind parks, solar parks, etc.), an energy distribution network, manufacturing systems, banking systems, insurance systems, medical applications, transportation applications, and infrastructure systems.
[0019] The production network 104 can include a computing system or LLM and RAG engine or system 106 that is connected to the IT network 102. The production network 104 can202419876 include various production machines configured to work together to perform one or more manufacturing operations. Example production machines of the production network 104 can include, without limitation, robots 108 and other field devices, such as sensors 110, actuators 112, or other machines, which can be controlled by a respective PLC 114. The PLC 114 can send instructions to respective field devices. In some cases, a given PLC 114 can be coupled to one or more human machine interfaces (HMIs) 116. The production network 104 can also define various application domain management systems or manufacturing domain management systems that can identify root causes (e.g., a first root cause or a root cause 1, further described herein). For example, an example manufacturing domain management system can manage a bill of material used in the production network 104.
[0020] The example system 100, in particular the production network 104, can define a fieldbus portion 118 and an Ethernet portion 120. For example, the fieldbus portion 118 can include the robots 108, PLC 114, sensors 110, actuators 112, and HMIs 116. The fieldbus portion 118 can define one or more production cells or control zones. The fieldbus portion 118 can further include a data extraction node 115 that can be configured to communicate with a given PLC 114 and sensors 110.
[0021] The PLC 114, data extraction node 115, sensors 110, actuators 112, and HMI 116 within a given production cell can communicate with each other via a respective field bus 122. Each control zone can be defined by a respective PLC 114, such that the PLC 114, and thus the corresponding control zone, can connect to the Ethernet portion 120 via an Ethernet connection 124. The robots 108 can be configured to communicate with other devices within the fieldbus portion 118 via a WiFi connection 126. Similarly, the robots 108 can communicate with the Ethernet portion 120, in particular a Supervisory Control and Data Acquisition (SCAD A) server 128, via the WiFi connection 126. The Ethernet portion 120 of the production network 104 can include various computing devices communicatively coupled together via the Ethernet connection 124. Example computing devices in the Ethernet portion 120 include, without limitation, a mobile data collector 130, HMIs 132, the SCADA server 128, the LLM and RAG system 106, a wireless router 134, a manufacturing execution system (MES) 136, an engineering system (ES) 138, and a log server 140. The ES 138 can include one or more engineering workstations. In an example, the MES 136, HMIs 132, ES 138, and log server 140 are connected to the production network 104 directly. The wireless router 134 can also connect to the production network 104 directly. Thus, in some cases, mobile users, for instance the202419876 mobile data collector 130 and robots 108, can connect to the production network 104 via the wireless router 134. In some cases, by way of example, the ES 138 and the mobile data collector 130 define guest devices that are allowed to connect to the computing system 106. The computing system 106 can define one or more LLMs or Al models configured to collect or obtain data related to the example industrial system 100.
[0022] Users of the system 100 can include, for example and without limitation, operators of an industrial plant or engineers that can update the control logic of a plant. By way of an example, an operator can interact with the HMIs 132, which may be located in a control room of a given plant. Alternatively, or additionally, an operator can interact with HMIs of the system 100 that are located remotely from the production network 104. Similarly, for example, engineers can use the HMIs 116 that can be located in an engineering room of the system 100. Alternatively, or additionally, an engineer can interact with HMIs of the system 100 that are located remotely from the production network 104.
[0023] Referring now to FIG. 2, the LLM and RAG system 106 can define a comprehensive verification and validation framework for integrating Large Language Models with Retrieval- Augmented Generation systems into industrial automation environments. The system 106 can define an architecture that is organized into service layers or components, for instance verification services or components 202, validation services or components 204, and fundamental services or components 206, which work together to ensure reliable, safe, and effective operation of the LLM and RAG system 106 in industrial settings. The system 106 can further include an API server 224 and an LLM agent 228, in a particular an LLM agent front end and back end, communicatively coupled to the API server 224. In various examples, the API server 224 and the LLM agent 228 can define the core computing system that integrates with industrial automation infrastructure such as various sensors, controllers, and actuators. The API server 224 and the LLM agent 228 can process inputs from the industrial environment and can generate outputs that can influence industrial control decisions and operator guidance. The layer architecture, for instance the validation services 202, verification services 204, and the fundamental services 206, can perform systematic verification and validation across all aspects of LLM and RAG system performance while maintaining clear separation of concerns.
[0024] With continuing reference to FIG. 2, the validation services layer 202 can define high- level validation components that assess the overall system performance and user acceptance. For example, the validation services 202 can include an unsupervised evaluator 206 configured202419876 to perform machine learning techniques to automatically assess system performance without requiring labeled ground truth data, thereby reducing the cognitive load on human operators. The unsupervised evaluator 206 can receive evaluation inputs from the API sever 224, and thus via the LLM agent 228. The unsupervised evaluator 206 can generate validation outputs that indicate system performance quality. The outputs can be provided to the API server 224, and thus to the LLM agent 228. The validation services 202 can also include an explainability analyzer 208 that ensures transparency and interpretability of LLM outputs by analyzing the reasoning processes behind system decisions. The explainability analyzer 208 can receive explainability inputs from the LLM agent 228 via the API server 204. Based on the inputs, the analyzer 208 can generate explainability outputs that provide human-readable explanations of system reasoning, which can be critical for building operator trust and enabling informed decision-making in industrial environments. The validation services 202 can also include a a reliability lifecycle manager 210 configured to continuously monitor the long-term performance of the LLM agent 228, and thus the LLM and RAG system, so as detect potential performance drift and degradation over time. The reliability lifecycle manager 210 can receive performance data inputs from the API server 224 and generate reliability assessment outputs, based on the performance data inputs, so as to inform or trigger maintenance and retraining decisions.
[0025] Still referring to FIG. 2, the verification services layer 204 can define various specialized verification components that test specific aspects of system reliability and performance. For example, and without limitation, the verification services 204 can include a data consistency checker 212, a generation robustness checker 214, a retrieval diversity analyzer 216, a model consistency tester 218, a quality of service analyzer 220, and a baseline checker 222. The data consistency checker 212 is configured to verify the integrity and consistency of input data feeding into the LLM agent 228 via the API server 224. Based on the data inputs to the data consistency checker 212, the checker 212 generates consistency verification outputs, so as to prevent errors caused by inconsistent or corrupted data sources. The verification services 204 further includes the generation robustness tester 214 configured to validate the LLM's ability to maintain consistent performance under varying environmental conditions and input perturbations. The robustness tester 214 can receive robustness test inputs from the LLM agent 228 via the API server 224, and generate robustness assessment outputs based on the test inputs, so as to ensure the system can operate reliably in the dynamic202419876 conditions typical of industrial environments. The retrieval diversity analyzer 216 is configured to enhance the adaptability of the RAG system 106 by ensuring diverse and contextually relevant information retrieval across varying operational scenarios. The diversity analyzer 216 can process retrieval inputs from the LLM agent 228 via the API server 224 and produces diversity analysis outputs based on the inputs. The model consistency tester 218 is configured to ensure that the LLM generates reliable and repeatable outputs that adhere to predefined performance standards. The tester 218 can receive model test inputs from the LLM agent 228 and can generate consistency verification outputs, which can be crucial for maintaining predictable system behavior in industrial applications. The Quality of Service (QoS) analyzer 220 is configured to monitors and validates that the LLM and RAG system 106 meets stringent real-time performance requirements essential for industrial automation. The QoS analyzer 220 can receive performance monitoring inputs from the LLM agent 228, via the API server 224, and generate QoS assessment outputs to the LLM agent 228, so as to ensure that the system can respond within required latency constraints. The baseline checker 222 is configured to maintain output consistency by comparing current system results against established reference baselines throughout the system lifecycle. The baseline checker 222 can receive baseline comparison inputs and can generate baseline verification outputs, so as to provide early warning of performance deviations.
[0026] The API server 224 can select the verification and validation services based on the unique demands of a given industrial automation systems, which can require high reliability, real-time responsiveness, and operator trust. Each service addresses a distinct dimension of system performance, for instance input integrity and model robustness to explainability and lifecycle stability. The combination of these services can ensure comprehensive coverage of potential failure modes, and their coordinated operation enables adaptive testing strategies that evolve with system behavior. This layered and synergistic approach, among other things, distinguishes the disclosed framework from conventional Al testing systems.
[0027] In various examples, the API server 224 defines the central orchestration component that manages communications between the verification and validation components or modules, and coordinates their interactions with the LLM agent 228. The API server 224 can receive various orchestration inputs from various components and generate coordination outputs that manage the overall verification and validation workflow, so as to enable seamless integration and coordination of all system components. The system 106 can further define a fundamental202419876 services layer 226 that includes the LLM agent 228 and a Neo4J database 230 communicatively coupled to the LLM agent 228, so as to define the core infrastructure for LLM and RAG operations. The LLM Agent backend and frontend 228 can manage the primary LLM operations and user interfaces, so as to receive agent inputs and producing agent outputs 249. The database 230 can define the knowledge graph repository for the RAG system. The database 230 can store and retrieve various contextual information to support RAG functionality.
[0028] Referring also to FIG. 3, example operations 300 can be performed by a computing system, for instance the LLM and RAG system 106. In various examples, insights from one verification and validation component directly inform and enhance the testing parameters of other components. It is recognized herein that this synergistic testing can significantly improve defect detection capabilities beyond what individual testing methods can achieve independently.
[0029] For example, at 302, the system 106, in particular the explainability analyzer 208, can receive or otherwise obtain system inputs. By way of example, in an industrial HVAC control scenario, the analyzer 208 might obtain various keywords from operators. At 304, the analyzer 208 can apply explainability frameworks (e.g., SHAP, LIME, or Integrated Gradients) to identify the most influential or critical features in the system inputs for given model decisions. Continuing with the example industrial HVAC control scenario, the explainability analyzer 208 might determine that the top five most influential keywords in an operator query about system performance are "HVAC," "chiller," "valves," "temperature," and "pressure." Based on that determination, at 306, the explainability analyzer 208 can generate explainability results that identify these critical input features. At 308, the system 106 can use the explainability results (through the coordinated operation of multiple components via the API server 224) as targeted inputs for the generation robustness tester 214. Thus, at 308, instead of applying random perturbations to system inputs, the robustness tester 214 can perturb the most influential features identified by the explainability analyzer 208. For instance, the system might systematically modify the keywords "HVAC" and "chiller" in operator queries while leaving other terms unchanged, creating focused robustness test cases.
[0030] With continuing reference to FIG. 3, at 310, the model consistency tester 218 and QoS analyzer 220 can measure the resulting system behavior changes by analyzing whether, for example, and without limitation: model outputs significantly change when perturbations are applied to influential regions, indicating potential system instability or brittleness; and key explainability scores change when perturbations are applied outside of influential regions,202419876 suggesting issues in model interpretation consistency. Thus, at 310, the model consistency tester 218 and the QoS analyzer 220 can generate integrated assessment results that combine multiple verification perspectives. By way of further example, if perturbing the heavily weighted terms "HVAC" and "chiller" significantly changes system predictions about equipment maintenance requirements, this might reveal that the LLM and RAG system is over- reliant on these specific terms, indicating potential overfitting. Conversely, if predictions remain stable but the explainability mapping radically shifts, this might indicate that the explainability results are too fragile or inconsistent for reliable operator use. At 312, the system 106 performs iterative refinement and continuous monitoring. In particular, for example, the reliability lifecycle manager 210 can use the discrepancies and failure points identified in the previous steps to inform continuous improvement of test case generation. The system can focus increasingly tighter testing on perturbations that yield the most significant failures or instabilities, creating an adaptive testing regime that evolves with system performance. Further, at 312, the baseline checker 222 can continuously compare the enhanced test results against established performance baselines, so as to provide ongoing validation that the cross-validation approach is maintaining or improving system reliability over time. In various examples, this defines a feedback loop where successful cross-validation insights are incorporated into the baseline expectations for future testing cycles. At 314, feedback is generated. In particular, for example, the real-time monitoring system that can be implemented through the coordinated operation of the reliability lifecycle manager 210, QoS analyzer 220, and API Server 224, can continuously monitor the performance of the LLM and RAG system 106, so as to provide ongoing feedback for verification and validation processes. Thus, the process can return to 302 where the feedback can define system inputs. This monitoring system can generate real-time status outputs that inform operators and system administrators about current system health and performance metrics.
[0031] Referring again to FIGs. 1 and 2, the LLM agent back end and front end 228 can integrate with industrial automation infrastructure, for instance the production network 104 that includes industrial controllers, sensors, and actuators, to provide real-world operational data to the verification and validation framework. This integration can ensure that the verification and validation processes are grounded in actual industrial operating conditions rather than theoretical test scenarios. Furthermore, without being bound by theory, the comprehensive architecture and cross-validation methodologies described herein can enables the systematic202419876 detection of complex failure modes that might be missed by conventional testing approaches, while providing the reliability, safety, and real-time performance characteristics essential for successful deployment of LLM and RAG systems in industrial automation environments.
[0032] Thus, as described herein, a computing system of an industrial automation system includes a large language model (LLM) and retrieval-augmented generation (RAG) system integrated with the industrial automation system that comprises sensors, controllers, and actuators. The system further includes a plurality of verification and validation components. Each component is configured to perform a different type of verification comprising a model consistency verification, data quality verification, performance verification, and robustness verification. The system can further include an API server configured to use results from a first verification and validation component of the plurality to inform testing parameters of a second verification and validation component of the plurality. The API server is further configured to generate targeted test cases based on the results from the first verification and validation component. The system can define a real-time monitoring system configured to continuously monitor performance of the LLM and RAG system and provide feedback for ongoing verification and validation.
[0033] In another example aspect of the disclosure, a Large Language Model (LLM) with Retrieval-Augmented Generation (RAG) (LLM and RAG system) in industrial automation can be cross-validated by obtaining a plurality of system inputs to the LLM and RAG system. A computing system can perform explainability analysis on the plurality of system inputs to identify a critical input feature, so as to define an identified critical input feature, wherein the identified critical input feature affects decisions of the LLM. The system can perturb the identified critical input feature while leaving other input features unchanged, so as to generate a focused test case. The focused test case can be executed to measure variations in response of the LLM and RAG system. The identified critical input feature can be modified, so as to detect instabilities in behavior of the LLM. The system can iteratively refine the focused test case based on the instabilities to create targeted test scenarios that enhance defect detection capabilities. Verification requests can be received from a plurality of verification components comprising a model consistency tester, a data consistency checker, a generation robustness tester, and a quality of service analyzer. In some cases, the plurality of verification components can be configured with respective outputs from a plurality of validation components comprising an explainability analyzer and a reliability lifecycle manager. The execution of the plurality of202419876 verification components can be coordinated through an API server, wherein results from a first verification component of the plurality of verification components inform testing parameters of a second verification component of the plurality of verification components. In various examples, real-time performance metrics of the LLM and RAG system is monitored during operation with industrial controllers, sensors, and actuators. Based on the monitoring, integrated assessment results can be generated by combining outputs from the plurality of verification components. The integrated assessment results can be compared against predefined performance baselines for industrial automation requirements. Feedback can be provided to the LLM and RAG system to maintain compliance with industrial safety and reliability standards.
[0034] In another example aspect, baseline performance metrics for the LLM and RAG system are established using a baseline checker component. The system can detect performance degradation by comparing current system outputs against the established baseline performance metrics. In particular, for example, the system can automatically trigger focused verification testing when performance degradation exceeds predetermined thresholds. The baseline performance metrics can be updated based on validated system improvements. In some examples, a first verification on the LLM and RAG system is executed to generate first verification results. The system can extract key insight parameters from the first verification results that indicate critical system dependencies, so as to define extracted key insight parameters. A second verification can be configured using the extracted key insight parameters as input constraints, wherein the second verification targets testing scenarios informed by the first verification results. The system can execute the second verification to generate enhanced test cases that are more targeted than either the first verification or the second verification achieve independently from each other. The system can also measure system performance changes resulting from the enhanced test cases. The system can identify failure modes revealed by the system performance changes, so as to define identified failure modes; and store the identified failure modes in a knowledge base for future test case generation.
[0035] In yet another example aspect, the system can receive operational data from industrial automation infrastructure including sensor readings, controller states, and actuator positions. The system can process the operational data through the LLM and RAG system to generate industrial control recommendations, so as to define generated recommendations. The system can simultaneously execute multiple verification analyses on the generated recommendations. For example, the system can measure response latency of the LLM and RAG system against202419876 industrial real-time requirements using a quality of service analyzer. Additionally, or alternatively, the system can validate that the generated recommendations meet safety and reliability criteria for industrial automation applications. The system can detect anomalies in behavior of LLM and RAG system by comparing current performance against historical baseline patterns. In various examples, the system automatically adjusts system parameters when performance metrics fall below industrial automation thresholds.
[0036] FIG. 4 illustrates an example of a computing environment within which embodiments of the present disclosure may be implemented. A computing environment 500 includes a computer system 510 that may include a communication mechanism such as a system bus 521 or other communication mechanism for communicating information within the computer system 510. The computer system 510 further includes one or more processors 520 coupled with the system bus 521 for processing the information. The computing system 106 may include, or be coupled to, the one or more processors 520.
[0037] The processors 520 may include one or more central processing units (CPUs), graphical processing units (GPUs), or any other processor known in the art. More generally, a processor as described herein is a device for executing machine-readable instructions stored on a computer readable medium, for performing tasks and may comprise any one or combination of, hardware and firmware. A processor may also comprise memory storing machine-readable instructions executable for performing tasks. A processor acts upon information by manipulating, analyzing, modifying, converting or transmitting information for use by an executable procedure or an information device, and / or by routing the information to an output device. A processor may use or comprise the capabilities of a computer, controller or microprocessor, for example, and be conditioned using executable instructions to perform special purpose functions not performed by a general purpose computer. A processor may include any type of suitable processing unit including, but not limited to, a central processing unit, a microprocessor, a Reduced Instruction Set Computer (RISC) microprocessor, a Complex Instruction Set Computer (CISC) microprocessor, a microcontroller, an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), a System-on-a-Chip (SoC), a digital signal processor (DSP), and so forth. Further, the processor(s) 520 may have any suitable microarchitecture design that includes any number of constituent components such as, for example, registers, multiplexers, arithmetic logic units, cache controllers for controlling read / write operations to cache memory, branch predictors, or the like. The microarchitecture202419876 design of the processor may be capable of supporting any of a variety of instruction sets. A processor may be coupled (electrically and / or as comprising executable components) with any other processor enabling interaction and / or communication there-between. A user interface processor or generator is a known element comprising electronic circuitry or software or a combination of both for generating display images or portions thereof. A user interface comprises one or more display images enabling user interaction with a processor or other device.
[0038] The system bus 521 may include at least one of a system bus, a memory bus, an address bus, or a message bus, and may permit exchange of information (e.g., data (including computer-executable code), signaling, etc.) between various components of the computer system 510. The system bus 521 may include, without limitation, a memory bus or a memory controller, a peripheral bus, an accelerated graphics port, and so forth. The system bus 521 may be associated with any suitable bus architecture including, without limitation, an Industry Standard Architecture (ISA), a Micro Channel Architecture (MCA), an Enhanced ISA (EISA), a Video Electronics Standards Association (VESA) architecture, an Accelerated Graphics Port (AGP) architecture, a Peripheral Component Interconnects (PCI) architecture, a PCI -Express architecture, a Personal Computer Memory Card International Association (PCMCIA) architecture, a Universal Serial Bus (USB) architecture, and so forth.
[0039] Continuing with reference to FIG. 4, the computer system 510 may also include a system memory 530 coupled to the system bus 521 for storing information and instructions to be executed by processors 520. The system memory 530 may include computer readable storage media in the form of volatile and / or nonvolatile memory, such as read only memory (ROM) 531 and / or random access memory (RAM) 532. The RAM 532 may include other dynamic storage device(s) (e.g., dynamic RAM, static RAM, and synchronous DRAM). The ROM 531 may include other static storage device(s) (e.g., programmable ROM, erasable PROM, and electrically erasable PROM). In addition, the system memory 530 may be used for storing temporary variables or other intermediate information during the execution of instructions by the processors 520. A basic input / output system 533 (BIOS) containing the basic routines that help to transfer information between elements within computer system 510, such as during start-up, may be stored in the ROM 531. RAM 532 may contain data and / or program modules that are immediately accessible to and / or presently being operated on by the processors 520. System memory 530 may additionally include, for example, operating system202419876534, application programs 535, and other program modules 536. Application programs 535 may also include a user portal for development of the application program, allowing input parameters to be entered and modified as necessary.
[0040] The operating system 534 may be loaded into the memory 530 and may provide an interface between other application software executing on the computer system 510 and hardware resources of the computer system 510. More specifically, the operating system 534 may include a set of computer-executable instructions for managing hardware resources of the computer system 510 and for providing common services to other application programs (e.g., managing memory allocation among various application programs). In certain example embodiments, the operating system 534 may control execution of one or more of the program modules depicted as being stored in the data storage 540. The operating system 534 may include any operating system now known or which may be developed in the future including, but not limited to, any server operating system, any mainframe operating system, or any other proprietary or non-proprietary operating system.
[0041] The computer system 510 may also include a disk / media controller 543 coupled to the system bus 521 to control one or more storage devices for storing information and instructions, such as a magnetic hard disk 541 and / or a removable media drive 542 (e.g., floppy disk drive, compact disc drive, tape drive, flash drive, and / or solid state drive). Storage devices 540 may be added to the computer system 510 using an appropriate device interface (e.g., a small computer system interface (SCSI), integrated device electronics (IDE), Universal Serial Bus (USB), or FireWire). Storage devices 541 , 542 may be external to the computer system 510.
[0042] The computer system 510 may also include a field device interface 565 coupled to the system bus 521 to control a field device 566, such as a device used in a production line. The computer system 510 may include a user input interface or GUI 561, which may comprise one or more input devices, such as a keyboard, touchscreen, tablet and / or a pointing device, for interacting with a computer user and providing information to the processors 520.
[0043] The computer system 510 may perform a portion or all of the processing steps of embodiments of the invention in response to the processors 520 executing one or more sequences of one or more instructions contained in a memory, such as the system memory 530. Such instructions may be read into the system memory 530 from another computer readable medium of storage 540, such as the magnetic hard disk 541 or the removable media drive 542. The magnetic hard disk 541 (or solid state drive) and / or removable media drive 542202419876 may contain one or more data stores and data files used by embodiments of the present disclosure. The data store 540 may include, but are not limited to, databases (e.g., relational, object-oriented, etc.), file systems, flat files, distributed data stores in which data is stored on more than one node of a computer network, peer-to-peer network data stores, or the like. The data stores may store various types of data such as, for example, skill data, sensor data, or any other data generated in accordance with the embodiments of the disclosure. Data store contents and data files may be encrypted to improve security. The processors 520 may also be employed in a multi-processing arrangement to execute the one or more sequences of instructions contained in system memory 530. In alternative embodiments, hard-wired circuitry may be used in place of or in combination with software instructions. Thus, embodiments are not limited to any specific combination of hardware circuitry and software.
[0044] As stated above, the computer system 510 may include at least one computer readable medium or memory for holding instructions programmed according to embodiments of the invention and for containing data structures, tables, records, or other data described herein. The term “computer readable medium” as used herein refers to any medium that participates in providing instructions to the processors 520 for execution. A computer readable medium may take many forms including, but not limited to, non-transitory, non-volatile media, volatile media, and transmission media. Non-limiting examples of non-volatile media include optical disks, solid state drives, magnetic disks, and magneto-optical disks, such as magnetic hard disk 541 or removable media drive 542. Non-limiting examples of volatile media include dynamic memory, such as system memory 530. Non-limiting examples of transmission media include coaxial cables, copper wire, and fiber optics, including the wires that make up the system bus 521. Transmission media may also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0045] Computer readable medium instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, statesetting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions may execute entirely on the user's computer, partly on the user's computer, as a202419876 stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0046] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, may be implemented by computer readable medium instructions.
[0047] The computing environment 500 may further include the computer system 510 operating in a networked environment using logical connections to one or more remote computers, such as remote computing device 580. The network interface 570 may enable communication, for example, with other remote devices 580 or systems and / or the storage devices 541, 542 via the network 571. Remote computing device 580 may be a personal computer (laptop or desktop), a mobile device, a server, a router, a network PC, a peer device or other common network node, and typically includes many or all of the elements described above relative to computer system 510. When used in a networking environment, computer system 510 may include modem 572 for establishing communications over a network 571, such as the Internet. Modem 572 may be connected to system bus 521 via user network interface 570, or via another appropriate mechanism.
[0048] Network 571 may be any network or system generally known in the art, including the Internet, an intranet, a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a direct connection or series of connections, a cellular telephone network, or any other network or medium capable of facilitating communication between computer system 510 and other computers (e.g., remote computing device 580). The network 571 may be202419876 wired, wireless or a combination thereof. Wired connections may be implemented using Ethernet, Universal Serial Bus (USB), RJ-6, or any other wired connection generally known in the art. Wireless connections may be implemented using Wi-Fi, WiMAX, and Bluetooth, infrared, cellular networks, satellite or any other wireless connection methodology generally known in the art. Additionally, several networks may work alone or in communication with each other to facilitate communication in the network 571.
[0049] It should be appreciated that the program modules, applications, computer-executable instructions, code, or the like depicted in FIG. 4 as being stored in the system memory 530 are merely illustrative and not exhaustive and that processing described as being supported by any particular module may alternatively be distributed across multiple modules or performed by a different module. In addition, various program module(s), script(s), plug-in(s), Application Programming Interface(s) (API(s)), or any other suitable computer-executable code hosted locally on the computer system 510, the remote device 580, and / or hosted on other computing device(s) accessible via one or more of the network(s) 571, may be provided to support functionality provided by the program modules, applications, or computer-executable code depicted in FIG. 4 and / or additional or alternate functionality. Further, functionality may be modularized differently such that processing described as being supported collectively by the collection of program modules depicted in FIGs. 2 and 4 may be performed by a fewer or greater number of modules, or functionality described as being supported by any particular module may be supported, at least in part, by another module. In addition, program modules that support the functionality described herein may form part of one or more applications executable across any number of systems or devices in accordance with any suitable computing model such as, for example, a client-server model, a peer-to-peer model, and so forth. In addition, any of the functionality described as being supported by any of the program modules depicted in FIGs. 2 and 4 may be implemented, at least partially, in hardware and / or firmware across any number of devices.
[0050] It should further be appreciated that the computer system 510 may include alternate and / or additional hardware, software, or firmware components beyond those described or depicted without departing from the scope of the disclosure. More particularly, it should be appreciated that software, firmware, or hardware components depicted as forming part of the computer system 510 are merely illustrative and that some components may not be present or additional components may be provided in various embodiments. While various illustrative202419876 program modules have been depicted and described as software modules stored in system memory 530, it should be appreciated that functionality described as being supported by the program modules may be enabled by any combination of hardware, software, and / or firmware. It should further be appreciated that each of the above-mentioned modules may, in various embodiments, represent a logical partitioning of supported functionality. This logical partitioning is depicted for ease of explanation of the functionality and may not be representative of the structure of software, hardware, and / or firmware for implementing the functionality. Accordingly, it should be appreciated that functionality described as being provided by a particular module may, in various embodiments, be provided at least in part by one or more other modules. Further, one or more depicted modules may not be present in certain embodiments, while in other embodiments, additional modules not depicted may be present and may support at least a portion of the described functionality and / or additional functionality. Moreover, while certain modules may be depicted and described as sub-modules of another module, in certain embodiments, such modules may be provided as independent modules or as sub-modules of other modules.
[0051] Although specific embodiments of the disclosure have been described, one of ordinary skill in the art will recognize that numerous other modifications and alternative embodiments are within the scope of the disclosure. For example, any of the functionality and / or processing capabilities described with respect to a particular device or component may be performed by any other device or component. Further, while various illustrative implementations and architectures have been described in accordance with embodiments of the disclosure, one of ordinary skill in the art will appreciate that numerous other modifications to the illustrative implementations and architectures described herein are also within the scope of this disclosure. In addition, it should be appreciated that any operation, element, component, data, or the like described herein as being based on another operation, element, component, data, or the like can be additionally based on one or more other operations, elements, components, data, or the like. Accordingly, the phrase “based on,” or variants thereof, should be interpreted as “based at least in part on.”
[0052] Although embodiments have been described in language specific to structural features and / or methodological acts, it is to be understood that the disclosure is not necessarily limited to the specific features or acts described. Rather, the specific features and acts are disclosed as illustrative forms of implementing the embodiments. Conditional language, such as, among202419876 others, “can,” “could,” “might,” or “may,” unless specifically stated otherwise, or otherwise understood within the context as used, is generally intended to convey that certain embodiments could include, while other embodiments do not include, certain features, elements, and / or steps. Thus, such conditional language is not generally intended to imply that features, elements, and / or steps are in any way required for one or more embodiments or that one or more embodiments necessarily include logic for deciding, with or without user input or prompting, whether these features, elements, and / or steps are included or are to be performed in any particular embodiment.
[0053] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the Figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions.
Claims
202419876CLAIMSWhat is claimed is:
1. A method for cross-validation testing of a Large Language Model (LLM) with Retrieval- Augmented Generation (RAG) (LLM and RAG) system in industrial automation, the method comprising: obtaining a plurality of system inputs to the LLM and RAG system; performing explainability analysis on the plurality of system inputs to identify a critical input feature, so as to define an identified critical input feature, the identified critical input feature affecting decisions of the LLM; perturbing the identified critical input feature while leaving other input features unchanged, so as to generate a focused test case; executing the focused test case to measure variations in response of the LLM and RAG system; modifying the identified critical input feature, so as to detect instabilities in behavior of the LLM; and iteratively refining the focused test case based on the instabilities to create targeted test scenarios that enhance defect detection capabilities.
2. The method as recited in claim 1, the method further comprising: receiving verification requests from a plurality of verification components comprising a model consistency tester, a data consistency checker, a generation robustness tester, and a quality of service analyzer.
3. The method as recited in claim 2, the method further comprising: configuring the plurality of verification components with respective outputs from a plurality of validation components comprising an explainability analyzer and a reliability lifecycle manager.
4. The method as recited in claim 2, the method further comprising: coordinating execution of the plurality of verification components through an API server, wherein results from a first verification component of the plurality of verification202419876 components inform testing parameters of a second verification component of the plurality of verification components.
5. The method as recited in claim 4, the method further comprising: monitoring real-time performance metrics of the LLM and RAG system during operation with industrial controllers, sensors, and actuators; generating integrated assessment results by combining outputs from the plurality of verification components; comparing the integrated assessment results against predefined performance baselines for industrial automation requirements; and providing feedback to the LLM and RAG system to maintain compliance with industrial safety and reliability standards.
6. The method as recited in claim 1, the method further comprising: establishing baseline performance metrics for the LLM and RAG system using a baseline checker component; and detecting performance degradation by comparing current system outputs against the established baseline performance metrics.
7. The method as recited in claim 6, the method further comprising: automatically triggering focused verification testing when performance degradation exceeds predetermined thresholds; and updating the baseline performance metrics based on validated system improvements.
8. The method as recited in claim 1, the method further comprising: executing a first verification on the LLM and RAG system to generate first verification results; extracting key insight parameters from the first verification results that indicate critical system dependencies, so as to define extracted key insight parameters; and configuring a second verification using the extracted key insight parameters as input constraints, wherein the second verification targets testing scenarios informed by the first verification results.2024198769. The method as recited in claim 8, the method further comprising: executing the second verification to generate enhanced test cases that are more targeted than the first verification or the second verification method achieve independently from each other; and measuring system performance changes resulting from the enhanced test cases.
10. The method ad recited in claim 9, the method further comprising: identifying failure modes revealed by the system performance changes, so as to define identified failure modes; and storing the identified failure modes in a knowledge base for future test case generation.
11. The method as recited in claim 1, the method further comprising: receiving operational data from industrial automation infrastructure including sensor readings, controller states, and actuator positions; processing the operational data through the LLM and RAG system to generate industrial control recommendations, so as to define generated recommendations; and simultaneously executing multiple verification analyses on the generated recommendations.
12. The method as recited in claim 11, the method further comprising: measuring response latency of the LLM and RAG system against industrial real-time requirements using a quality of service analyzer; and validating that the generated recommendations meet safety and reliability criteria for industrial automation applications.
13. The method as recited in claim 12, the method further comprising: detecting anomalies in behavior of LLM and RAG system by comparing current performance against historical baseline patterns; and automatically adjusting system parameters when performance metrics fall below industrial automation thresholds.20241987614. A computing system of an industrial automation system, the computing system comprising: a large language model (LLM) and retrieval-augmented generation (RAG) system integrated with the industrial automation system that comprises sensors, controllers, and actuators; a plurality of verification and validation components, each component configured to perform a different type of verification comprising a model consistency verification, data quality verification, performance verification, and robustness verification; and an API server configured to use results from a first verification and validation component of the plurality to inform testing parameters of a second verification and validation component of the plurality, the API server further configured to generate targeted test cases based on the results from the first verification and validation component.
15. The computing system as recited in claim 14, the system further comprising: a real-time monitoring system configured to continuously monitor performance of the LLM and RAG system and provide feedback for ongoing verification and validation.