System and method for incorporating user feedback into retrieval enhancement tool based on large language model

By implementing controllers and optimization logic on the host device, user feedback is received and the Large Language Model (LLM) output is optimized, solving the problem of inaccurate LLM output. This achieves improved output accuracy and consistency while reducing system complexity, ensuring the reliability and accuracy of LLM.

CN121765031APending Publication Date: 2026-03-31GM GLOBAL TECHNOLOGY OPERATIONS LLC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-15
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing Large Language Models (LLMs) lack systematic coverage metrics, have non-diversified test sets, and uneven test coverage when generating output, resulting in inaccurate and inappropriate output. A systematic approach is needed to ensure the accuracy, precision, consistency, and reliability of LLM output, while providing user feedback to further ensure the accuracy and consistency of the output.

Method used

By implementing a controller on the host device, receiving user feedback, and utilizing logic such as performance tuning optimizers, hybrid searchers, revised domain knowledge, and constraint regeneration, the reliance on human validators is gradually reduced, and the LLM output is optimized. This includes revising the domain knowledge base and system hints, and performing regression tests to ensure the accuracy and consistency of the output.

Benefits of technology

This approach improves the accuracy and consistency of LLM output while reducing system complexity, gradually reduces reliance on human validators, ensures the reliability and accuracy of LLM output, and provides user feedback for further optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765031A_ABST
    Figure CN121765031A_ABST
Patent Text Reader

Abstract

A system for incorporating user feedback into a large language model (LLM)-based retrieval enhancement (RAG) tool includes an application for incorporating user feedback (UFA). The UFA receives an LLM input from a host device user, generates an LLM output as a response to the input such that the system user verifies the LLM output and provides user feedback via a human machine interface (HMI) of the host device, and in response to the user feedback, enables a performance adjustment optimizer to adjust performance of the host device. The performance adjustment optimizer modifies the LLM output by revising domain knowledge, revising system cues, and performing constrained regeneration of the LLM output. Over time, the performance adjustment optimizer gradually reduces computing resource utilization, gradually increases computing efficiency, and gradually reduces dependency on human verifier and user feedback, and wherein the LLM output is a command to one or more systems of the host device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to systems and methods for generating validation test suites for artificial intelligence-driven tools, and more specifically, to automatically generating validation and evaluation test suites for large language model-driven tools. Background Technology

[0002] Artificial intelligence (AI) models, including large language models (LLMs), are increasingly being used to perform tasks for end users across a variety of technical and non-technical pursuits. The output of AI models (including LLMs) can be hampered by a lack of systematic coverage metrics, non-diversified test sets, and uneven test coverage and coverage measurements. Therefore, LLMs are trained on large amounts of data (typically from diverse sources) and then optimized using a retrieval-enhanced generation (RAG) process to ensure accuracy. However, even RAG-assisted LLMs can generate inaccurate, inappropriate, or compromised outputs for various reasons.

[0003] Therefore, while current systems and methods for validating generative AI-driven tools have achieved their intended purpose, there is still a need for new and improved systems and methods that can provide a systematic approach to automatically determine the accuracy, relevance, and consistency of RAG tools-assisted LLM, ensure the accuracy, precision, consistency, and reliability of LLM outputs, provide redundancy and consistency checks to ensure that the accuracy, precision, consistency, and reliability of LLM outputs are maintained, while maintaining or reducing system complexity, and provide user feedback to further ensure the accuracy, relevance, and consistency of LLM outputs. Summary of the Invention

[0004] According to several aspects, a system for incorporating user feedback into a Large Language Model (LLM)-based Retrieval Enhancement (RAG) tool includes a host device with a controller. The controller has a processor, memory, and input / output (I / O) ports. The I / O ports communicate with a Human-Machine Interface (HMI) and one or more databases. The processor executes program control logic stored in memory. The program control logic includes an application for incorporating user feedback (UFA) into the LLM-based RAG tool. The UFA includes at least a first control logic, a second control logic, a third control logic, and a fourth control logic. The first control logic receives input to the LLM from a user on the host device. The second control logic generates LLM output as a response to the input. The third control logic enables a system user to verify the LLM output and provide user feedback via the host device's HMI. In response to the user feedback, the fourth control logic enables a performance tuning optimizer, which modifies the LLM output by revising domain knowledge, revising system hints, and performing constraint regeneration of the LLM output. The performance tuning optimizer gradually reduces computational resource utilization, gradually increases computational efficiency, and gradually reduces reliance on human verifiers and user feedback over time. Furthermore, the LLM output consists of commands to one or more systems on the host device.

[0005] In another aspect of this disclosure, the first control logic also includes control logic for receiving input via a human-machine interface (HMI) of the host device.

[0006] In another aspect of this disclosure, the second control logic further includes control logic for enabling a hybrid retrieval unit. The hybrid retrieval unit determines the similarity between the input and predetermined data in one or more databases stored in memory. The second control logic also includes control logic for causing the hybrid retrieval unit to generate LLM outputs and output confidence scores, control logic for causing a human verifier to prioritize and review the LLM outputs based on the output confidence scores, and control logic for assigning the output confidence scores to the LLM outputs. The LLM output commands one or more actuators of the host device to adjust the performance of the relevant host device system. The second control logic also includes control logic that causes the human verifier to prioritize and review the LLM outputs based on the output confidence scores and ranking context, and control logic that causes the human verifier to selectively update one or more of the text vector database and the original text database using data obtained from the input, the LLM output, and the output confidence scores. The host device is a vehicle, and the LLM output commands one or more actuators of the vehicle to change the vehicle's functionality based on input received from the user.

[0007] In another aspect of this disclosure, the third control logic also includes control logic that prompts the system user to provide feedback by requesting confirmation that the LLM output is a sufficiently accurate response to the input or user command. The confirmation request also includes: visual, auditory, tactile, verbal, numeric, or alphanumeric confirmation requests, and measures the accuracy of the response based on the user's personal preferences.

[0008] In another aspect of this disclosure, the fourth control logic also includes control logic for enabling subroutines for revising the domain knowledge base (SRDK). The SRDK receives LLM output and user feedback and determines whether the LLM output has been modified by the user, and when it is determined that the LLM output has been modified by the user, executes control logic for optimizing the domain knowledge base (DKB).

[0009] In another aspect of this disclosure, the control logic for optimizing the DKB further includes control logic that retrieves the confidence score of the user-modified LLM output and determines whether the user-modified LLM output already exists in the DKB. When it is determined that the user-modified LLM output does not yet exist in the DKB, the control logic for optimizing the DKB obtains input from a human expert to verify that the user-modified LLM output has been correctly added to the DKB. When it is determined that the user-modified LLM output already exists in the DKB, the control logic for optimizing the DKB revises the domain knowledge embedded in the DKB to generate an updated DKB containing the user-modified LLM output.

[0010] In another aspect of this disclosure, the fourth control logic also includes control logic for enabling a subroutine for revising the system prompt (SRSP). The SRSP receives a prompt revision verification request, LLM output, and user feedback, and determines whether there is sufficient evidence to implement a revision to the system prompt. To determine if sufficient evidence exists, the SRSP compares the host device prompt with the user feedback and determines whether there is a threshold level of similarity between the host device prompt and the user feedback. When it is determined that sufficient evidence does exist, the fourth control logic executes the control logic to rewrite the system prompt according to the performance test.

[0011] In another aspect of this disclosure, the control logic for rewriting system prompts also includes control logic that performs regression testing on the rewritten system prompts. The regression testing utilizes test inputs stored in the memory of the DKB. The regression testing verifies that the new information in the rewritten system prompts allows the system to continue operating without negatively impacting system responsiveness. When it is determined that the rewritten system prompts are functioning correctly, the system executes control logic to update the system prompts in the DKB. When it is determined that the rewritten system prompts are not functioning correctly, the system continues to recursively and continuously rewrite system prompts using user feedback, performing regression testing on the rewritten system prompts and testing system functionality until the regression tests indicate that the new information allows the system to continue operating without negatively impacting system responsiveness.

[0012] In another aspect of this disclosure, the fourth control logic also includes control logic for executing a subroutine for constraint regeneration (SCR). The SCR responds to the LLM output to utilize user feedback to revise the LLM output based on user constraints.

[0013] In another aspect of this disclosure, the SCR also includes control logic for reviewing the context used by the LLM to generate the LLM output, and control logic for determining the presence of low-quality contexts. A low-quality context defines a low-quality match between the context used by the LLM to generate the LLM output and the context of user input. Low-quality contexts are defined based on user preferences. The SCR also includes control logic that receives user feedback via the HMI indicating the presence of a low-quality context, deleting the low-quality context, and re-prioritizing the contexts before executing the control logic to regenerate the LLM output based on the context constrained by the user feedback.

[0014] In another aspect of this disclosure, a method for incorporating user feedback into a Large Language Model (LLM)-based Retrieval Enhancement (RAG) tool includes program control logic stored in the controller's memory, executed by a processor of a vehicle controller. The controller also has input / output (I / O) ports that communicate with a Human-Machine Interface (HMI) and one or more databases. The program control logic includes an application for incorporating user feedback (UFA) into the LLM-based RAG tool. The UFA includes control logic for receiving LLM input from a vehicle user via the vehicle's HMI, generating LLM output as a response to the input, enabling the user to verify the LLM output and provide user feedback via the vehicle's HMI; and, in response to the user feedback, activating a performance tuning optimizer that modifies the LLM output by revising domain knowledge, revising system hints, and performing constraint regeneration of the LLM output. The performance tuning optimizer gradually reduces computational resource utilization, gradually increases computational efficiency, and gradually reduces reliance on human verifiers and user feedback over time, wherein the LLM output is a command to one or more systems of the vehicle.

[0015] In another aspect of this disclosure, the method further includes enabling a hybrid retrieval system. The hybrid retrieval system determines the similarity between the input and predetermined data in one or more databases stored in memory, causing the hybrid retrieval system to generate an LLM output and an output confidence score. The method also includes causing a human verifier to prioritize and review the LLM output based on the output confidence score, and assigning the output confidence score to the LLM output. The LLM output commands one or more actuators of the vehicle to adjust the performance of the relevant vehicle system. The method further includes causing a human verifier to prioritize and review the LLM output based on the output confidence score and ranking context, and causing the human verifier to selectively update one or more of a text vector database and a raw text database using data obtained from the input, the LLM output, and the output confidence score. The LLM output commands one or more actuators of the vehicle to change the vehicle's functionality based on input received from a user.

[0016] In another aspect of this disclosure, the method also includes prompting the system user for feedback by requesting confirmation that the LLM output is a sufficiently accurate response to the input or user command. The request for confirmation also includes: visual, auditory, tactile, verbal, numeric, or alphanumeric confirmation requests, and measures the accuracy of the response based on the user's personal preferences.

[0017] In another aspect of this disclosure, the method also includes enabling a subroutine for revising the domain knowledge base (SRDK). The SRDK receives LLM output and user feedback and determines whether the LLM output has been modified by the user, and when it is determined that the LLM output has been modified by the user, executes control logic for optimizing the domain knowledge base (DKB).

[0018] In another aspect of this disclosure, optimizing the DKB also includes retrieving the confidence score of the user-modified LLM output and determining whether the user-modified LLM output already exists in the DKB. When it is determined that the user-modified LLM output does not yet exist in the DKB, the method obtains input from a human expert to verify that the user-modified LLM output has been correctly added to the DKB. When it is determined that the user-modified LLM output already exists in the DKB, the method revises the domain knowledge embedded in the DKB to generate an updated DKB containing the user-modified LLM output.

[0019] In another aspect of this disclosure, the method also includes enabling a subroutine for revising the system prompt (SRSP). The SRSP receives a prompt revision verification request, LLM output, and user feedback, and determines whether there is sufficient evidence to implement a revision to the system prompt. The SRSP also includes using the SRSP to compare the vehicle prompt with the user feedback to determine whether there is sufficient evidence; and determining whether there is a threshold level of similarity between the vehicle prompt and the user feedback. When it is determined that sufficient evidence does exist, the method rewrites the system prompt based on performance testing.

[0020] In another aspect of this disclosure, rewriting system prompts also includes performing regression testing on the rewritten system prompts. The regression testing utilizes test inputs stored in the memory of the DKB. The regression testing verifies that the new information in the rewritten system prompts allows the vehicle to continue operating without negatively impacting vehicle responsiveness. When it is determined that the rewritten system prompts are functioning correctly, the regression testing executes the control logic for updating the system prompts in the DKB. When it is determined that the rewritten system prompts are not functioning correctly, the regression testing continues to recursively and continuously rewrite the system prompts using user feedback, performing regression testing on the rewritten system prompts and testing vehicle functionality until the regression tests indicate that the new information allows the vehicle to continue operating without negatively impacting system responsiveness.

[0021] In another aspect of this disclosure, the method also includes a subroutine for executing constraint regeneration (SCR). The SCR responds to the LLM output to utilize user feedback to revise the LLM output based on user constraints.

[0022] In another aspect of this disclosure, the SCR also includes control logic for reviewing the context used by the LLM to generate the LLM output, and control logic for determining the presence of a low-quality context. A low-quality context defines a low-quality match between the context used by the LLM to generate the LLM output and the context of user input. Low-quality contexts are defined based on user preferences. The SCR also includes control logic for receiving user feedback via the HMI, indicating the presence of a low-quality context, deleting the low-quality context, and re-prioritizing the contexts before executing the control logic to regenerate the LLM output based on the context constrained by the user feedback.

[0023] In another aspect of this disclosure, a method for incorporating user feedback into a Large Language Model (LLM)-based Retrieval Enhancement (RAG) tool includes: executing program control logic stored in controller memory by a processor of a vehicle controller. The controller also includes input / output (I / O) ports communicating with a Human-Machine Interface (HMI) and one or more databases. The program control logic includes an application for incorporating user feedback (UFA) into the LLM-based RAG tool, the UFA including control logic for receiving input to the LLM from a vehicle user via the vehicle's HMI and control logic for generating LLM outputs as a response to the inputs, including: enabling a hybrid retrieval unit. The hybrid retrieval unit determines the similarity between the input and predetermined data stored in one or more databases in memory. The method also includes control logic for causing the hybrid retrieval unit to generate LLM outputs and output confidence scores, such that a human verifier prioritizes and reviews the LLM outputs based on the output confidence scores; and assigns the output confidence scores to the LLM outputs. The LLM outputs command one or more actuators of the vehicle to adjust the performance of the relevant vehicle systems. The method also includes control logic for enabling a human validator to prioritize and review LLM outputs based on output confidence scores and ranking context, and for enabling the human validator to selectively update one or more of the text vector database and the original text database using data obtained from the inputs, LLM outputs, and output confidence scores. The LLM outputs command one or more actuators of the vehicle to change the vehicle's functionality based on input received from the user. The method also includes control logic for enabling the user to validate the LLM outputs and prompt the user for feedback via the vehicle's HMI, including providing a request for confirmation that the LLM outputs are a sufficiently accurate response to the inputs or user commands. The request for confirmation also includes: visual, auditory, tactile, verbal, numeric, or alphanumeric request for confirmation, and wherein a sufficiently accurate response is measured based on the user's personal preferences. In response to user feedback, a performance tuning optimizer is activated, which modifies the LLM outputs by revising domain knowledge, including enabling a subroutine for revising the domain knowledge base (SRDK). The SRDK receives the LLM outputs and user feedback and determines whether the LLM outputs have been modified by the user, and when it is determined that the LLM outputs have been modified by the user, executes control logic for optimizing the domain knowledge base (DKB). The method also includes control logic for retrieving the confidence score of the user-modified LLM output and determining whether the user-modified LLM output already exists in the DKB. When it is determined that the user-modified LLM output does not yet exist in the DKB, the method executes control logic for obtaining input from a human expert to verify that the user-modified LLM output has been correctly added to the DKB.When it is determined that the user-modified LLM output already exists in the DKB, the method revises the domain knowledge embedded in the DKB to generate an updated DKB containing the user-modified LLM output. The method also includes control logic for revising system prompts, including enabling a subroutine for Revising System Prompts (SRSP). The SRSP receives a prompt revision verification request, LLM output, and user feedback, and determines whether there is sufficient evidence to implement a revision to the system prompt. The method also includes control logic for using the SRSP to compare the vehicle prompt with the user feedback to determine whether there is sufficient evidence; and for determining whether there is a threshold level of similarity between the vehicle prompt and the user feedback. When it is determined that sufficient evidence does exist, the method executes the control logic for rewriting the system prompt based on performance testing and performs regression testing on the rewritten system prompt. The regression testing utilizes test inputs stored in the DKB's memory. The regression testing verifies that the new information in the rewritten system prompt allows the vehicle to continue operating without negatively impacting vehicle response. When it is determined that the rewritten system prompt is functioning correctly, the method executes control logic to update the system prompt in the DKB. When it is determined that the rewritten system prompt is not functioning correctly, the method continues to rewrite the system prompt recursively and continuously using user feedback, performs regression testing on the rewritten system prompt, and tests vehicle functionality until the regression tests show that the new information allows the vehicle to continue operating without negatively impacting system response. The method also includes control logic for performing constraint regeneration (SCR) on LLM output, including: executing a subroutine for constraint regeneration (SCR). The SCR responds to LLM output by using user feedback to revise the LLM output based on user constraints, examines the context used by the LLM to generate the LLM output, and determines whether a low-quality context exists. A low-quality context defines a low-quality match between the context used by the LLM to generate the LLM output and the context of user input. Low-quality contexts are defined based on user preferences. The method also includes control logic for receiving user feedback via the HMI, which indicates the existence of a low-quality context, deletes the low-quality context, and re-prioritizes the contexts before executing the control logic to regenerate the LLM output based on the context of user feedback constraints. The performance tuning optimizer gradually reduces computational resource utilization, gradually increases computational efficiency, and gradually reduces reliance on human verifiers and user feedback over time. The LLM output is a command to one or more systems of the vehicle.

[0024] Further areas of application will become apparent from the description provided herein. It should be understood that these descriptions and specific examples are for illustrative purposes only and are not intended to limit the scope of this disclosure. Attached Figure Description

[0025] The accompanying drawings described herein are for illustrative purposes only and are not intended to limit the scope of this disclosure in any way.

[0026] Figure 1 This is a schematic diagram depicting a system for calculating confidence scores of a retrieval enhancement (RAG) tool based on a large language model (LLM) according to an exemplary embodiment;

[0027] Figure 2 It describes a method for computing according to an exemplary embodiment. Figure 1 A flowchart illustrating the logical flow of an application system based on the confidence score of the LLM-based RAG tool;

[0028] Figure 3 It is a description of a method for merging according to an exemplary embodiment. Figure 1 A flowchart of the logical flow of the User Feedback (UFA) application for the system;

[0029] Figure 4 It is a description of the revision according to exemplary embodiments. Figure 3 A flowchart of the logical flow of subroutines of the domain knowledge (SRDK) within the UFA;

[0030] Figure 5 It is a description of the revision according to exemplary embodiments. Figure 3 A flowchart of the logical flow of the System Prompt (SRSP) subroutine within the UFA; and

[0031] Figure 6 It is a description of a method for using according to an exemplary embodiment. Figure 3 The flowchart shows the logical flow of the constraint regeneration (SCR) subroutine within the UFA. Detailed Implementation

[0032] The following description is merely exemplary in nature and is not intended to limit this disclosure, its application, or its uses.

[0033] Reference Figure 1A system 10 for calculating confidence scores for a Retrieval Augmentation (RAG) tool based on a Large Language Model (LLM) is illustrated schematically. System 10 generally runs in or on a host device 12. Host device 12 can take any of a variety of forms, including a vehicle 14. However, it should be understood that system 10 of this disclosure is not required to be tied to such a vehicle 14. Rather, vehicle 14 is merely an exemplary, non-limiting embodiment to which system 10 of this disclosure is described herein. System 10 can operate in any hardware and software configuration in which a generative AI-driven tool is used to receive input from user 16, such as user command 18, and generate output 20 that alters the functionality of the hardware and / or software configuration or system in which the generative AI-driven tool is used. Furthermore, while the vehicle 14 shown is a car, it should be understood that vehicle 14 can be any type of vehicle without departing from the scope or intent of this disclosure. In several non-limiting examples, vehicle 14 can be: automobiles, trucks, sport utility vehicles (SUVs), semi-trailer trucks, tractor trailers, tractors, combine harvesters or other such agricultural equipment; powered and unpowered aircraft, such as airplanes, helicopters, gliders or rotorcraft; powered and unpowered vessels, such as boats, sailboats, motorboats, yachts, jet skis, sailboats, etc. In other non-limiting embodiments, it should be understood that system 10 described herein can be adapted to operate with host device 12, such as manned or unmanned spacecraft, for example: satellites, rockets, space stations, and other orbital and extraorbital satellite communication enabled devices, without departing from the scope or intent of this disclosure. In further non-limiting examples, host device 12 can include a mobile computing platform, such as a laptop, mobile phone, tablet, or any other such host device 12 through which a user can interact with generative AI-driven tools.

[0034] System 10 also includes a controller 22, which is a non-general-purpose electronic control device having a pre-programmed digital computer or processor 24, a non-transitory computer-readable medium or memory 26 for storing data such as control logic, software applications, instructions, computer code, data, lookup tables, etc., and a transceiver or input / output (I / O) port 28. The computer-readable medium or memory 26 includes any type of media accessible by a computer, such as read-only memory (ROM), random access memory (RAM), hard disk drive, optical disc (CD), digital video disc (DVD), or any other type of memory 26. "Non-transitory" computer-readable memory 26 excludes wired, wireless, optical, or other communication links that transmit transient electrical or other signals. Non-transitory computer-readable memory 26 includes media that can permanently store data and media that can store data and subsequently overwrite it, such as rewritable optical discs or erasable storage devices. Computer code includes any type of program code, including source code, object code, and executable code. The processor 24 is configured to execute code or instructions.

[0035] When system 10 operates on vehicle 14, controller 22 may include a dedicated Wi-Fi controller or engine control module, transmission control module, body control module, infotainment control module, etc. Transceiver or I / O port 28 is configured to wirelessly communicate with backend 30 using cellular protocols including Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Enhanced Data Rate Evolution of GSM (EDGE), Universal Mobile Telecommunications Service (UMTS), High-Speed ​​Packet Access (HSPA), Code Division Multiple Access (CDMA), Evolved Data Optimized (EV-DO / EVDO / 1xEV-DO), Short Message Service (SMS), Wi-MAX, Manufacturing Message Specification (MMS), 2G, 3G, 4G, 5G, and wireless and cellular standards defined under IEEE 802.1X, IEEE 802LAN / MAN, and the IEEE Mobile Communications Networks Standards Committee (MobiNet-SC) standards. Backend 30 may include one or more controllers 22 and / or one or more human experts or verifiers 32, which are shown and described in more detail in the following figures.

[0036] Controller 22 also includes one or more applications 34. Application 34 is a software program configured to perform a specific function or set of functions. Application 34 may include one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or portions thereof suitable for implementation in appropriate computer-readable program code. Application 34 may be stored within memory 26 or in additional or separate memory. Examples of applications 34 include audio or video streaming services, games, browsers, social media, algorithms for calculating confidence scores of RAG tools 38 (hereinafter referred to as CLR application 40) based on Large Language Model (LLM) 36, computational generative AI-driven tools or LLM 36 tools, and applications 34 for incorporating user feedback (hereinafter referred to as UFA 41). Generative AI-driven tools or LLM 36 define applications 34 stored locally on vehicle 14 controller 22 and / or in a remote backend 30 or cloud-based computing device. CLR application 40 includes multiple subroutines or control logic sections.

[0037] In examples where System 10 operates in or on Vehicle 14, a vehicle operator or vehicle development engineering user 16 can use System 10 and CLR application 40 to validate tools for developing and testing systems that dynamically adjust the operating modes of Vehicle 14 and its characteristics. In some examples, Vehicle 14 may be equipped with a navigation system 42, one or more drive motors 44 that provide and change the amount of torque transmitted to the wheels 46 of Vehicle 14 to cause Vehicle 14 to move, stop, etc., and a steering system 48 that can adjust the direction and heading of Vehicle 14. In other examples, Vehicle 14 is equipped with a braking system 50 that, when activated by controller 22, decelerates the movement of Vehicle 14. Vehicle 14 may be equipped with various other vehicle motion control systems that can be used to alter or otherwise control the dynamic performance of Vehicle 14, including but not limited to aerodynamic control surfaces and actuators, active and / or semi-active suspension systems and actuators, without departing from the scope or intent of this disclosure.

[0038] The LLM 36 is trained on large amounts of data in one or more online and offline modes. The training data provides an efficient way for the LLM 36 to generate outputs for tasks such as answering questions, performing mathematical calculations, performing language translation, and / or summarizing. Retrieval-enhanced generation (RAG) is an architectural approach used to improve the quality of responses generated by the LLM 36 by supplementing its internal information representation by building the LLM 36 model on external knowledge sources. Implementing RAG in an LLM 36-based question-answering system offers at least two main benefits: ensuring that the LLM 36 has access to up-to-date, reliable data, and that user 16 has access to the sources of the LLM 36, ensuring that the LLM 36's responses to user 16's input can be trusted. Furthermore, by building the LLM 36 on a set of externally verifiable facts, RAG minimizes the opportunity for the LLM 36 to extract information into its parameters, thus reducing the chance of the LLM 36 leaking sensitive data or providing incorrect or misleading responses to user 16's input. More specifically, the LLM 36-based RAG tool 38 of this disclosure extends the capabilities and improves the efficiency of LLM 36 applications by utilizing custom data. Custom data can relate to any of a wide range of topics, but should generally be understood to specifically relate to a particular hardware or software application of the host device 12. In a non-limiting example, custom or domain-specific knowledge data stored in the Domain Knowledge Base (DKB) 51 can relate to the vehicle 14's in-vehicle systems, including but not limited to navigation systems, powertrain control systems, suspension control systems, heating, ventilation, and air conditioning (HVAC) systems, etc. LLM 36 utilizes the custom data from the RAG tool 38 related to commands, questions, or tasks generated by user 16, and provides the custom data as context for LLM 36 to generate responses. In many respects, RAG is an effective method for improving LLM 36 performance and has successfully supported chatbots and question-and-answer (Q&A) systems that require access to domain-specific information. Domain-specific information may include any data in a broad range of data relating to the functions of a specific host device 12, vehicle 14, and backend 30, and additionally to the onboard hardware, software, and applications of each of the host device 12, vehicle 14, and backend 30. Domain-specific information may also relate to specific proprietary data held by the original equipment manufacturer (OEM) that relates to specific hardware and software products manufactured by the OEM. Domain-specific information may define an embedded model stored in the memory 26 of the onboard controller 22 of the host device 12.

[0039] Now refer to Figure 2 And continue to refer to Figure 1A more detailed exemplary schematic logic flowchart of the CLR application 40 is shown. The CLR application 40 receives input 100 in offline mode. Input 100 may be received via a human-machine interface (HMI) 53 of the host device 12, including but not limited to audio or video or audiovisual receivers, such as microphones or other audio sensors 52, cameras or other visual sensors 54, and haptic interfaces including but not limited to buttons and touchscreens 56. In several examples, input 100 is a user command 18 of the user 16, which is subsequently processed in a series of subroutines by an LLM-based RAG tool 38 before generating output 20. Input 100 may take any of a variety of forms without departing from the scope or intent of this disclosure. In some non-limiting examples, input 100 may include a spoken or written user command 18 in any language, such as: a command to enable the navigation system of the vehicle 14, a command to search for points of interest, a command to change the cabin temperature of the vehicle 14, a command to pause or wait or delay, a request to obtain a mathematically stated solution, or any other such command 18. Input 100 is processed by multiple subroutines within the LLM 36-based RAG tool 38.

[0040] Output 20 is the response of host device 12 to input 100 from user 16. In several aspects, output 20 directly or indirectly alters the function of one or more systems of host device 12, including but not limited to: altering one or more functions of the navigation system, powertrain control system, suspension control system, HVAC system, etc. Therefore, output 20 may include LLM output commands that modify the destination or route planning function of the navigation system, or change the powertrain, suspension control, or HVAC system mode or operation by directly or indirectly altering the position of the host device actuator within the relevant host device 12 system, in response to user command 18 input 100 to adjust the performance of the relevant systems.

[0041] RAG tool 38 includes a hybrid retriever 102. Hybrid retriever 102 is a complex retrieval algorithm that improves the relevance of the retrieved contextual information by aggregating results from multiple different and parallel retrievers. By using multiple parallel retrievers, the advantages of each can be leveraged to obtain results more accurately related to user command 18 or input 100 than individual data retrieval algorithms. Input 100 is received within hybrid retriever 102 by at least two parallel subroutines (i.e., semantic tool 200 and syntactic tool 300) of the LLM36-based RAG tool 38. Semantic tool 200 receives input 100 in semantic retriever 202. The semantic retriever 202 subroutine first extracts semantic information from input 100. Semantic retriever 202 then calculates a semantic similarity score 204 of the semantic information extracted from input 100. To calculate the semantic similarity score 204, the semantic retrieval unit 202 converts the semantic information extracted from input 100 into a semantic text vector in a vector space, such that the vector is a mathematical graphical representation of the semantic information extracted from input 100. As used herein, the meaning of any particular semantic vector is implicitly defined by its position and orientation in the vector space. In a non-limiting example, input 100 may include a verbal command from user 16. It should be understood that the verbal command can be any of a variety of different commands, including but not limited to: “increase cabin temperature,” which user 16 intends for input 100 to command vehicle 14 to use the heating, ventilation, and air conditioning (HVAC) system to change the temperature of the passenger compartment of vehicle 14. Each word in the command input 100 by user 16, namely “increase,” “cabin,” and “temperature,” is parsed and plotted in the vector space and subsequently compared with a predetermined text vector in the text vector database 206 stored in memory 26.

[0042] Because the language input commands 100 from different users 16 may differ in the actual wording or terminology chosen, the vector representing input 100 may not be precisely represented by text vectors stored in text vector database 206. Text vector database 206 contains multiple semantic text vectors that define the mathematical, graphical, vectorized representation of predefined semantic input 100 programmed by host device 12 to accept, respond to, and understand. These multiple text vectors in text vector database 206 can be manually and / or automatically selected during pre-programming of system 10, and can be updated with new information when specific events occur, or can be manually or automatically, continuously, periodically, etc., updated or modified. In a non-limiting example, text vector database 206 is updated by a human expert or verifier 32.

[0043] Therefore, the semantic retrieval unit 202 calculates a semantic similarity score 204 based on the text vector representing input 100 and predefined text vector information accessed within the text vector database 206. In several aspects, the semantic similarity score 204 is a numerical, graphical, vectorized representation of the similarity level between the semantic structure of input 100 and the semantic structure of the text vector that most closely corresponds to the semantic text vector of input 100. After calculating the semantic similarity score 204, the semantic tool 200 performs a ranking fusion calculation 208, which filters the context of input 100 based on a threshold T to avoid "low-quality" contexts or low-quality matches between data in the text vector database 206 and the text vector representing input data 100. More specifically, the CLR application 40 utilizes an inverse ranking fusion algorithm to reorder and merge the results from each of the semantic tool 200 and the syntactic tool 300.

[0044] In other words, hybrid retrieval system 102 (Ret i Obtain the retrieved context C i1 And vector similarity score S i1 And if and only if the similarity score C i Make the vector similarity score S i Context C is eliminated only when the value is less than the threshold T. i In contrast, the "high-quality" context is defined based on an importance score, where the importance weights are based on different search engine CCs. i Context C i Counting, different retrieval RC i Context C i Ranking and Context C i It is calculated using the importance score, and the formula is IC. i =f(CC) i ,RC i Then, the contexts are ranked based on importance scores. Therefore, in some unrestricted examples, a series of ranked semantic similarity scores are generated according to the following formula:

[0045] Ret a =[(C a1 ,S a1 ),(C a2 ,S a2 ),…(C an ,S an )]

[0046] Ret b =[(C b1 ,S b1 ),(C b2 ,S b2 ),…(C bn,S bn )]

[0047] Ret c =[(C c1 ,S c1 ),(C c2 ,S c2 ),…(C cn ,S cn )]

[0048] Wherein, context {C ij}, j≥1 are ranked according to similarity score, i.e.: S i1 >S i2 …>S in Importance score (IC) i Context C is defined i The relative importance of the type of the received input 100. Therefore, in a non-limiting example, for context C that does not involve critical host device 12 functions (e.g., the powertrain, suspension, or safety-critical systems of vehicle 14). i (e.g., HVAC function), importance score IC i It has a low value. Conversely, when context C i When the indication input 100 is more closely related to or directly involves the function of the critical host device 12, the importance score IC i Higher than the low value.

[0049] Conversely, the syntactic tool 300, operating in parallel with the semantic tool 200, receives input 100 within the syntactic retrieval unit 302. Like the semantic retrieval unit 202, the syntactic retrieval unit 302 subroutine first extracts syntactic information from input 100. Then, the syntactic retrieval unit 302 subroutine calculates a syntactic similarity score 304 based on the raw text of input 100 and a raw text database 306 stored in memory 26. The raw text database 306 contains multiple raw text vectors that define the mathematical, graphical, vectorized representation of the predefined syntactic input 100, as programmed to accept and understand by the host device 12. These raw text vectors in the raw text database 306 can be manually and / or automatically selected during pre-programming of the system 10, and can be updated with new information when specific events occur, or can be updated or modified manually or automatically, continuously, periodically, etc. As used herein, the meaning of any particular raw text vector is implicitly defined by its position and orientation in the vector space. In the unrestricted example, the original text database 306 is updated by a human expert or validator 32.

[0050] To compute the syntactic similarity score 304, the syntactic retrieval subroutine 302 converts the syntactic information extracted from input 100 into a raw text vector in a vector space, making the vector a mathematical graphical representation of the syntactic information extracted from input 100. The syntactic similarity score 304 is a numerical representation of the level of similarity between the syntactic structure of input 100 and the data in the raw text database 36. More specifically, the syntactic similarity score 304 is computed based on the Jaccard index and the Levenshtein distance between them. The Jaccard index... This is used to measure the similarity and diversity between input 100 information and predefined original text information stored in the original text database 306. Similarly, the Levenshtein distance is a string metric used to measure the difference between two sequences, or in this example, a string metric used to measure the difference between the original text of input 100 and the original text stored in the original text database 306. In the example, the Levenshtein distance between two words is the minimum number of single-character edits (i.e., insertions, deletions, or replacements) required to change one word to another. The Levenshtein distance can be mathematically represented as follows:

[0051]

[0052] It should be understood that the use of the Jaccard index and Levenshtein distance is intended only as an exemplary, non-limiting example of an algorithm or function type that can be used to compute the syntactic similarity score 304 according to the purposes of this disclosure. The semantic similarity score 204 and the syntactic similarity score 304 may cover a range of values ​​that vary depending on the application, but in one non-limiting example, the semantic similarity score 204 and the syntactic similarity score 304 vary between zero (0) and one (0) such that the semantic similarity score 204 and / or the syntactic similarity score 304 equals 0 when there is no similarity whatsoever, and the semantic and / or syntactic similarity scores equal 1 when the semantic similarity score 204 and / or the syntactic similarity score 304 are completely consistent with the information in the text vector database 206 or the original text database 306.

[0053] Subsequently, the combined sorting and fusion calculations of the outputs of semantic similarity score 208 and syntactic similarity score 304 are performed, and a confidence score of 400 is calculated. The confidence score 400 is a normalized weighted sum of the semantic similarity score 204 and the syntactic similarity score 304. The confidence score can be expressed as:

[0054] C out =Z(w i *Score sem +w j *Score syn)

[0055] Where Z(...) is the normalization function, W i W j The weight is used for the semantic similarity score of 204. sem = f(rankfusion score), the syntactic similarity score of 304 is: Score syn =f(Jaccard Index,LevenshteinDistance,…).

[0056] In several aspects, weight W i W j It can vary significantly depending on the application. Weight W i W j This can be selected by application developers, original equipment manufacturers, suppliers, etc. It should also be understood that weight W... i W j It can be dynamic, variable, or constant, depending on the type and structure of the query that system 10 receives as input 100.

[0057] As described herein, the LLM 36-based RAG tool 38 generates output 20, which requires verification and validation by a human expert or validator 32, at least during the pre-production process. However, the system 10 of this disclosure offers several advantages, including, but not limited to, automatically generating output confidence scores for the LLM 36 output 20 via the CLR application 40 of this disclosure, allowing the human expert or validator 32 to effectively prioritize validation, thereby significantly reducing the human effort and time required to validate the LLM 36 output 20, while improving the accuracy of the LLM 36, reducing computational work and resource consumption, and reducing the likelihood of human-introduced typographical, syntactic, or other such errors, ranging from a first number to a second number significantly smaller than the first number. In the example, the expert or validator 32 may choose to first validate the LLM 36 output 20 with low confidence values ​​and then take action to validate the LLM 36 output 20 with a confidence level higher than the low confidence value. Therefore, confidence scores can be automatically calculated by utilizing the retrieval's vector-based similarity score and the input 100 similarity score (e.g., Jaccard distance) of a given input 100. It should also be understood that, in pre-production or production settings, as the CLR 40 is continuously used, evaluated, and updated over time, the amount of interaction and input from human experts or verifiers 32 decreases. That is, even in production applications where non-engineering end-users or customers interact with the LLM 36, the CLR application 40 can operate to accurately, consistently, reliably, and robustly interpret end-user or customer input to the system 10 and generate responses accordingly, with computational resource utilization gradually decreasing, computational efficiency gradually increasing, and reliance on verification by human verifiers 32 gradually decreasing.

[0058] Turn now Figure 3 And continue to refer to Figure 1 and Figure 2Once host device 12 or vehicle 14 is put into production, system 10, and more specifically, UFA 41, is shown in more detail in the form of a flowchart. Starting at box 500, the output 20 of LLM 36 is received by one or more human experts or verifiers 32. In some examples, the human expert or verifier 32 may be an engineer or expert remotely located in the background 30 of host device 12 or vehicle 14, or the human expert or verifier 32 may be a user of host device 12, such as a customer. Thus, in a non-limiting example, the human expert or verifier 32 may be a user 16 of vehicle 14, such as a driver, passenger, etc. The human expert or verifier 32 uses performance tuning optimizer 502 to verify that the output 20 of LLM 36 is a correct and accurate response to input 100 from user 16 and to user feedback 504, such as providing a correct and accurate response from host device 12 to user command 18. In other words, after validating output 20 using performance tuning optimizer 502, UFA 41 generates a validated LLM output 20' after receiving user feedback 504 from human expert or validator 32. User feedback 504 can take any of a variety of forms without departing from the scope or intent of this disclosure. In a non-limiting example, user feedback 504 may include one or more inputs to HMI 53 of host device 12, including but not limited to audio, visual and / or haptic user inputs to a microphone or other audio sensor 52, a camera or other visual sensor 54, or a haptic interface, including but not limited to buttons and touchscreen 56, etc. User feedback 504 can be positive and / or corrective in nature, depending on the level of similarity between input 100 or user command 18 and the LLM output 20 based thereon.

[0059] The performance tuning optimizer 502 includes at least three distinct subroutines or control logic: a subroutine for revising Domain Knowledge (SRDK) 600, a subroutine for revising System Tips (SRSP) 700, and a subroutine for Constraint Regeneration (SCR) 800. Now refer to... Figure 4 And continue to refer to Figure 1-3 The SRDK 600 is shown in more detail in the form of a flowchart.

[0060] SRDK 600 begins by receiving a verification prompt 602 from the performance tuning optimizer 502. The verification prompt 602 includes LLM output 20 and user feedback 504. In some non-limiting examples, the verification prompt 602 is provided to user 16 via HMI 53. The verification prompt 602 may include an audiovisual, tactile, verbal, numeric, or alphanumeric, or other such request to confirm that the LLM output 20 accurately and adequately represents the type of information or system 10 response desired by user 16 via input 100 or user command 18. Confirmation requests may include, but are not limited to, audiovisual, tactile, verbal, numeric, or alphanumeric confirmation requests via HMI 53. The adequacy of the response is measured based on user 16's personal preferences. SRDK 600 then determines at box 604 whether the output 20 from LLM 36 has been modified by user 16. When it is determined that output 20 has not been modified, SRDK 600 proceeds to box 606, where SRDK 600 exits and its results are combined with those from SRSP 700 and SCR 800, then returns to box 500 of UFA41. However, when it is determined that output 20 has been modified by user 16, SRDK 600 initiates the DKB 51 optimization process 608, which begins at box 610. At box 610, SRDK 600 retrieves the confidence score of output 20 modified by user 16. Subsequently, at box 612, SRDK 600 determines whether output 20 modified by user 16 already exists in DKB 51. When it is determined that the output 20 modified by user 16 is not yet present in DKB 51, SRDK 600 proceeds to box 614, where SRDK 600 prompts one or more human experts or validators 32 to review and provide input, response, or other such guidance to refine the output of system 10 to respond more accurately, precisely, and consistently to user 16's input 100 or command 18. Once the human expert or validator 32 has reviewed and provided a response to refine system 10's output 20, SRDK 600 returns to box 612 to reassess whether the output 20 modified by user 16 exists in DKB 51. When it is determined at box 612 that the output 20 modified by user 16 exists in DKB 51, SRDK 600 proceeds to box 616, where SRDK 600 revises the knowledge embedded in DKB 51, and updates DKB 51' to include the new output 20 modified by user 16 and validated by expert 32.

[0061] Turn now Figure 5 And continue to refer to Figure 1-4The flowchart illustrates SRSP 700 in more detail. SRSP 700 begins by receiving a prompt revision verification request 702 from the performance tuning optimizer 502. The prompt revision verification request 702 includes LLM output 20 and user feedback 504. SRSP 700 then determines at box 704 whether sufficient evidence has been provided to indicate that a prompt change is appropriate. Specifically, at box 704, system 10 and SRSP 700 compare the host device 12 prompt with user feedback 504 to determine if a threshold level of similarity exists between the host device 12 prompt and user feedback 504. When it is determined that a threshold level of similarity has been reached, indicating that the host device 12 prompt and user feedback 504 indicate that no change is needed, SRSP 700 proceeds to box 706, where SRSP 700 exits and the results of SRSP 700 are combined with the results of SRDK 600 and SCR 800, before returning to box 500 of UFA 41. However, when it is determined that the similarity has not yet reached a threshold level, SRSP 700 proceeds to box 708. At box 708, SRSP 700 rewrites the system 10 prompt. While rewriting the system 10 prompt, SRSP 700 attempts to increase the similarity level between user feedback 504 and the host device 12 system 10 prompt from a first level to a second level greater than the first level. In several examples, the rewriting of the system 10 prompt can be performed electronically and automatically, or, as shown in the figure, manually by one or more human experts or validators 32. Subsequently, at box 710, SRSP 700 uses test input 712 to perform a regression test on the rewritten system 10 prompt from box 708. The regression test performed at box 710 ensures that the new information in the rewritten system 10 prompt functions correctly and that the response of system 10 is not negatively affected by the rewritten system 10 prompt. At box 714, SRSP 700 determines whether the newly modified system 10 prompt performs satisfactorily based on predetermined data such as predetermined and / or variable metrics and / or thresholds. If performance is still determined to be unsatisfactory, SRSP 700 returns to box 708, where the System 10 prompt is rewritten to more closely align user feedback 504 with the System 10 prompt. It should be understood that the threshold used to determine whether performance is satisfactory or unsatisfactory can be predetermined, variable, or dynamic, depending on the System 10 prompt and the host device 12 functionality involved in user command 18. At box 714, once performance is determined to be satisfactory, SRSP 700 proceeds to box 716, where the System 10 prompt is recorded in DKB 51 or another such updated DKB 51', including the System 10 prompt modified by new user 16 and verified by expert 32. It should be understood that SRSP 700 is only required to change the System 10 prompt in very rare cases, as the System 10 prompt is an integral part of the LLM 36 functionality.Therefore, the review of the SRSP700 process by human experts 32 was intended to reduce the likelihood of erroneous data being added to the updated DKB 51' system 10 prompt.

[0062] Turn now Figure 6 And continue to refer to Figure 1-5 The SCR 800 is illustrated in more detail as a flowchart. The SCR 800 regenerates the context and / or redefines the priority of existing contexts to regenerate the LLM 36 output based on user 16's feedback 504. That is, the SCR 800 examines the context from the original LLM output 20 and uses progressively smaller subsets of context to filter and more accurately regenerate the LLM output 20 that is more relevant to the input 100 or command 18 from user 16. Figure 6 As shown, SCR 800 begins at box 802, where user 16 and / or human expert or validator 32 review the context used by LLM 36 to generate LLM output 20. Subsequently, at box 804, SCR 800 determines whether a low-quality context exists. As previously described with respect to CLR application 40, a "low-quality" context defines a low-quality match between one set of data and another, specifically the context used by LLM 20 and user 16 input 100 in a non-restrictive example. When a low-quality context is determined to exist at box 804, SCR 800 proceeds to box 806, where the low-quality context is removed from use within DKB 51. Subsequently, at box 808, SCR 800 re-prioritizes the contexts within DKB 51 to take into account the removed low-quality context. Referring again to box 804, when it is determined that no low-quality context exists, SCR 800 proceeds directly from box 804 to box 808, where, if necessary, the context repriority ordering in DKB 51 is performed. Subsequently, SCR 800 proceeds to box 810, where SCR 800 creates a regenerated LLM output 20”, which explains the context repriority ordering in DKB 51.

[0063] In a non-restrictive example of the SCR 800 in use, user 16 input command 18 could be a command to host device 12 or vehicle 14 to navigate to fast food restaurants within ten miles of user 16's current location. LLM 36 processes user 16's input command 18, referencing various databases and / or DKB 51 to determine which fast food restaurants are popular within a ten-mile radius of user 16's current location. LLM 36 then generates LLM output 20, which includes a predetermined number of top-ranking contexts, such as a ranking list of the most popular fast food restaurants locally, nationally, or internationally. Based on LLM output 20 and user 16's preferences, as shown in box 802, user 16 can determine that LLM output 20 is not accurate or specific enough for user 16's own expectations and can request LLM 36 to regenerate LLM output 20 based on additional constraints, such as user 16's preferences for ethnic foods, health foods, etc. Subsequently, SCR 800 causes LLM 36 to create a new, regenerated LLM output 20 constrained by user 16's preferences. In other words, SCR 800 causes LLM 36 to regenerate LLM output 20 based on the context of user feedback 504's constraints. Therefore, user 16 can continuously rearrange the priorities of LLM output 20 and the regenerated LLM output 20 based on user 16's preferences to further refine and specify LLM output 20. That is, SCR 800 responds to LLM output 20 by utilizing user feedback 504 to revise LLM output 20 based on user constraints.

[0064] System 10 and UFA 41, including the performance tuning optimizer 502 of this disclosure, offer several advantages, including providing a systematic approach to automatically determining the accuracy, relevance, and consistency of the RAG tool-assisted LLM 36, ensuring the accuracy, precision, consistency, and reliability of the LLM output 20, and providing redundant and consistent checks to ensure that the accuracy, precision, consistency, and reliability of the LLM output 20 are maintained, while maintaining or reducing the complexity of System 10, and providing user feedback 504 to be received and implemented to further ensure the accuracy, relevance, and consistency of the LLM output 20, while effectively prioritizing verification and significantly reducing the manpower and time required to verify the output 20 of the LLM 36. Simultaneously, System 10 and UFA 41 of this disclosure improve the accuracy of the LLM 36 while reducing computational workload and computational resource consumption, and simultaneously improve computational efficiency and reduce the likelihood of human-introduced typographical, syntactic, semantic, or other such errors, ranging from a first number to a second number significantly less than the first number.

[0065] The descriptions in this disclosure are merely exemplary in nature, and variations thereof that do not depart from the spirit and scope of this disclosure are intended to fall within its scope. Such variations should not be considered as departing from the spirit and scope of this disclosure.

Claims

1. A system for incorporating user feedback into a large language model (LLM)-based retrieval-augmented generation (RAG) tool, the system comprising: a host device having a controller with a processor, a memory, and input / output (I / O) ports in communication with a human-machine interface (HMI) and one or more databases, the processor executing program control logic stored in the memory, the program control logic including an application for incorporating user feedback (UFA) into an LLM-based RAG tool, the UFA comprising: first control logic that receives input to the LLM from a host device user; second control logic that generates an LLM output as a response to the input; third control logic that causes a system user to validate the LLM output and provide user feedback via a human-machine interface (HMI) of the host device; and fourth control logic that, in response to the user feedback, enables a performance adjustment optimizer that modifies the LLM output by revising domain knowledge, revising system prompts, and performing constraint regeneration of the LLM output, wherein the performance adjustment optimizer gradually reduces computational resource utilization, gradually increases computational efficiency, and gradually reduces reliance on human validators and user feedback over time, and wherein the LLM output is a command to one or more systems of the host device.

2. The system of claim 1, wherein, the first control logic further comprising: control logic for receiving the input via a human-machine interface (HMI) of the host device.

3. The system of claim 2, wherein, the second control logic further comprising: control logic for enabling a hybrid retriever that determines a similarity between the input and predetermined data stored in one or more databases in the memory; control logic for causing the hybrid retriever to generate the LLM output and an output confidence score; control logic for causing a human validator to prioritize and review the LLM output based on the output confidence score; control logic for assigning the output confidence score to the LLM output, wherein the LLM output commands one or more actuators of the host device to adjust performance of a related host device system; control logic that causes the human validator to prioritize and review the LLM output based on the output confidence score and a ranking context; and causing the human validator to selectively update one or more of a text vector database and a raw text database with data obtained from the input, the LLM output, and the output confidence score, and wherein the host device comprises a vehicle, and the LLM output commands one or more actuators of the vehicle to change a function of the vehicle based on input received from the user.

4. The system of claim 3, wherein, the third control logic further comprising: control logic that prompts the system user for feedback by requesting confirmation that the LLM output is a sufficiently accurate response to the input or user command, wherein the requesting confirmation further comprises an audio-visual, haptic, verbal, digital, or alphanumeric requesting confirmation, and wherein the sufficiently accurate response is measured based on the user's personal preferences.

5. The system of claim 4, wherein, The fourth control logic further comprises: control logic for enabling a subroutine for a revised domain knowledge, SRDK, wherein the SRDK receives the LLM output and the user feedback and determines whether the LLM output has been modified by the user, and when it is determined that the LLM output has been modified by the user, executes control logic for optimizing a domain knowledge base, DKB.

6. The system of claim 4, wherein, The control logic for optimizing the DKB further comprises: control logic that retrieves a confidence score of the LLM output that has been modified by the user and determines whether the LLM output that has been modified by the user already exists in the DKB, wherein, when it is determined that the LLM output that has been modified by the user does not already exist in the DKB, obtains input from a human expert to verify that the LLM output that has been modified by the user is correctly added to the DKB; and wherein, when it is determined that the LLM output that has been modified by the user already exists in the DKB, revises domain knowledge embedded in the DKB, thereby generating an updated DKB containing the LLM output that has been modified by the user.

7. The system of claim 4, wherein, The fourth control logic further comprises: control logic for enabling a subroutine for a revised system prompt, SRSP, wherein the SRSP receives a revised verification request, the LLM output, and the user feedback, and determines whether there is sufficient evidence to implement a revision to a system prompt, wherein to determine whether there is sufficient evidence, the SRSP compares a host device prompt to the user feedback and determines whether there is a threshold level of similarity between the host device prompt and the user feedback, wherein when it is determined that there is sufficient evidence, control logic is executed to rewrite the system prompt according to performance testing.

8. The system of claim 7, wherein, The control logic that rewrites the system prompt further comprises: control logic that performs regression testing on the rewritten system prompt, wherein the regression testing utilizes test inputs stored in memory of the DKB, the regression testing verifies that the new information in the rewritten system prompt allows the system to continue to function without negatively impacting system responses, and wherein, when it is determined that the rewritten system prompt functions properly, control logic is executed to update the system prompt in the DKB, and wherein, when it is determined that the rewritten system prompt does not function properly, the system prompt is recursively and continuously rewritten with user feedback, the rewritten system prompt is regression tested and tested for system functionality until the regression testing indicates that the new information allows the system to continue to function without negatively impacting system responses.

9. The system of claim 4, wherein, The fourth control logic further comprises: control logic for executing a subroutine for constrained regenerative SCR, wherein the SCR utilizes user feedback to revise LLM output based on user constraints in response to LLM output.

10. The system of claim 9, wherein, The SCR further comprises: control logic for reviewing a context in which the LLM generated the LLM output; control logic for determining whether a low quality context exists, wherein the low quality context defines a low quality match between a context in which the LLM generated the LLM output and a context of the user input, wherein a low quality context is defined according to user preferences; and control logic that receives user feedback via the HMI, the control logic indicating that a low quality context exists, deleting the low quality context, and re-determining a prioritization of contexts to regenerate LLM output according to a context constrained by user feedback prior to executing control logic.