GENERATION OF DIVERSIFIED VALIDATION TEST SUITES FOR GENERATIVE ARTIFICIAL INTELLIGENCE-BASED TOOLS
The system automates the generation of diversified validation test suites for generative AI-based tools, addressing inefficiencies in current validation methods by reducing human effort and increasing computing efficiency, ensuring robust and reliable LLM tool performance.
Patent Information
- Application Number
- DE102024129555
- Authority / Receiving Office
- DE · DE
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-27
- Filing Date
- 2024-10-12
- Publication Date
- 2026-03-05
AI Technical Summary
Current systems for validating generative AI-based tools, particularly large language models (LLMs), face inefficiencies due to lack of systematic coverage metrics, non-diversified test sets, and inconsistent test coverage, requiring time-consuming and labor-intensive human development that may result in incomplete validation.
A system and method for generating a diversified validation test suite using a controller with programmatic control logic that includes a diversified validation test suite application (DVTSA) to automate the generation of validation tests, reducing human effort and improving coverage metrics, and progressively increasing computing efficiency while reducing resource utilization and human reliance.
The solution automates the generation of validation test suites, enhancing coverage metrics, reducing human involvement, and increasing computing efficiency, thereby improving the robustness and reliability of LLM tools before deployment.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
INTRODUCTION
[0001] This disclosure relates to systems and methods for generating validation test suites for tools that utilize artificial intelligence, and in particular to the automatic generation of validation and evaluation test suites for tools that utilize large language models. Artificial intelligence (AI) models, including large language models (LLMs), are increasingly used to perform tasks for end users in a wide variety of technical and non-technical fields. The results of AI models, including LLMs, can be hampered by the lack of systematic coverage metrics, non-diversified test sets, and inconsistent test coverage and coverage measurements.Furthermore, validating the results of AI models, including LLMs, requires the development of validation tests by humans, which are both time-consuming and labor-intensive, and have a high probability of being incomplete.
[0002] While current systems and procedures for validating generative AI-based tools fulfill their intended purpose, there is a need for a new and improved system and procedure for generating a diversified validation test suite for generative AI-based tools. This would reduce human time and effort by automating the generation of the validation test suite, improving or optimizing systematic coverage metrics, and diversifying a set of tests with coverage metrics. Resource utilization would be reduced from a first level to a second level that is lower than the first, and the validation process speed would be streamlined and increased without increasing the complexity of the system design or hardware. Human involvement would also be reduced from a first level to a second level that is lower than the first.and is streamlined with improved coverage in order to gain increased confidence in the generative AI-based tool before its deployment. SUMMARY
[0003] According to several aspects of the present disclosure, a system for generating a diversified validation test suite for generative artificial intelligence (AI)-based tools comprises a controller with a processor, memory, and input / output (I / O) ports. The processor executes the programmatic control logic stored in memory. The programmatic control logic comprises a diversified validation test suite application (DVTSA) for a generative AI-based tool or a large language model (LLM). The DVTSA includes at least a first, a second, and a third control logic. The first control logic receives seed test input or user input via the I / O ports, analyzes the seed test input or user input, and extracts key elements for variations. The second control logic performs a coverage measurement of the outputs of the first control logic.The third control logic prompts a human reviewer to evaluate the outputs of the second control logic against predefined coverage metrics and to selectively and continuously iterate to ensure that the outputs of the second control logic increase the LLM tool input and output robustness from a first level to a second level that is higher than the first, in both production and pre-production LLM tool processes. The system progressively reduces the utilization of computing resources, progressively increases computing efficiency, and progressively reduces reliance on the human reviewer.
[0004] In another aspect of the present disclosure, the first control logic also includes a control logic for extracting syntactic information from the seed test input or user input, and a control logic that identifies and fills information gaps in the extracted syntactic information. The first control logic also synthesizes variations that actively adapt to close identified information gaps.
[0005] In another aspect of the present disclosure, the control logic for extracting syntactic information further comprises: control logic that collects all seed test input or user input coverage terms, including the extraction of individual keywords; and control logic that quantifies the output of coverage points based on corresponding test outputs and language information related to the extracted individual keywords. The control logic for extracting the syntactic information also automatically generates a script that uses each extracted keyword and employs supplementary documents as well as the quantified output of the coverage points to collect all coverage points.
[0006] In another aspect of the present disclosure, the supplementary documents further include: predefined language information, including linguistic databases, mathematical databases, syntactic and semantic databases, such as GridXML, ROBOT Script, natural language databases, mathematical terminology, and databases of variables, instructions, and terms defined in relation to the system hardware.
[0007] In another aspect of this disclosure, the input coverage terms include predefined terms and instructions for known inputs and outputs of the system. The coverage metrics are variable and are actively and automatically updated. The coverage metrics include an input robustness score and an output robustness score. Each input and output robustness score defines a percentage of the coverage of a list of, or all, possible output instructions of a specific type and a specific mapping of names to variables.
[0008] In another aspect of this disclosure, the input robustness score further considers variations in input terminology, variations in the nature of input instructions, and variations in substring interpretation coverage. The output robustness score comprises a percentage of the coverage of a list of, or all, possible output instructions and the mapping of names to variables.
[0009] In another aspect of this disclosure, the control logic that identifies and closes information gaps further comprises: control logic that reads all terms and phrases relevant to the coverage metrics from the outputs of the control logic for extracting syntactic information and divides the terms and phrases into key elements. The control logic that identifies and closes information gaps also comprises a first and a second loop control logic. The first loop control logic identifies variations for each of the key terms in a given context and assigns a confidence score to each of the variations for each of the key terms. The second loop control logic identifies input variations for each seed test input instruction and each user input instruction and assigns a confidence score to each of the seed test input and user input instructions.
[0010] In another aspect of the present disclosure, the first control logic further includes control logic that causes human reviewers to manually evaluate the outputs of each of the first and second loop control logics in order to identify variations for each of the key terms and for each of the seed test input and user input instructions.
[0011] In another aspect of the present disclosure, the second control logic also includes: control logic that applies the coverage metrics to outputs of the first control logic, comprising: measuring a seed extension for variation of input terms; measuring a seed extension for variation of input instruction types; measuring a seed extension for substring interpretation coverage; measuring a seed extension for output robustness; and aggregating and normalizing measured seed extensions for variation of input terms, variation of input instruction types, substring interpretation, and output robustness, and generating a coverage measurement score.
[0012] In another aspect of the present disclosure, the third control logic further comprises: control logic that causes the human reviewer to evaluate outputs of the second control logic by selectively and continuously adding or removing examples from the supplemental documents; and control logic that causes the human reviewer to convert the coverage measurement into an improved seed file and selectively and continuously iterating inputs into the control logic, which synthesizes variations that actively adapt to fill identified information gaps in order to increase the LLM tool input and output robustness from the first level to the second level, which is greater than the first level, in both production and pre-production LLM tool processes.
[0013] Further aspects of this disclosure include a method for generating a diversified validation test suite for generative artificial intelligence (AI)-based tools: executing programmatic control logic stored in the memory of a controller that has a processor, memory, and input / output (I / O) ports. The programmatic control logic comprises a diversified validation test suite application (DVTSA) for a generative AI-based tool or a large language model (LLM).The procedure further includes: receiving a seed test input or user input via the I / O ports; analyzing the seed test input or user input; extracting key elements from the seed test input or user input for variations; performing a coverage measurement of the outputs of the analysis and extraction steps; and evaluating outputs with respect to predefined coverage metrics and selectively and continuously iterating to increase the robustness of the LLM tool input and output from a first level to a second level that is higher than the first, in both production and pre-production LLM tool processes. The procedure progressively reduces the use of computing resources, progressively increases computing efficiency, and progressively reduces the reliance on a human reviewer.
[0014] In another aspect of the present disclosure, the method further includes extracting syntactic information from the seed test input or user input; identifying and filling information gaps in the extracted syntactic information; and synthesizing variations that actively adapt to close identified information gaps.
[0015] In another aspect of the present disclosure, the method further comprises: collecting all seed test input or user input coverage terms, including the extraction of individual keywords; quantifying the output of coverage points based on corresponding test outputs and language information for the extracted individual keywords; and automatically generating a script using each extracted keyword. The method further includes using supplementary documents and the quantified outputs of the coverage points to collect all coverage points.
[0016] In another aspect of the present disclosure, the use of supplementary documents further includes: accessing and referencing predefined language information, including linguistic databases, mathematical databases, syntactic and semantic databases, such as GridXML, ROBOT Script, natural language databases, mathematical terminology, and databases of variables, instructions, and terms defined in relation to the system hardware.
[0017] In another aspect of this disclosure, collecting all seed test input or user input coverage terms also includes collecting predefined terms and instructions for known inputs and outputs of the procedure. The coverage metrics are variable and are actively and automatically updated, and they include an input robustness score and an output robustness score. Each input and output robustness score defines a percentage of the coverage of a list of, or all, possible output instructions of a given type and a given mapping of names to variables.
[0018] In another aspect of this disclosure, the evaluation of outputs with respect to predefined coverage metrics also includes the evaluation of outputs with respect to the input robustness score. The input robustness score takes into account: variations in input terminology; variations in the nature of the input instructions; and variations in the coverage of substring interpretation; and the evaluation of outputs with respect to the output robustness score. The output robustness score is a percentage of the coverage of a list of, or all, possible output instructions and name-variable mappings.
[0019] In another aspect of the present disclosure, the method further comprises reading all terms and phrases relevant to the coverage metrics from outputs of the extracted syntactic information and splitting the terms and phrases into key elements, identifying variations for each of the key terms in a given context with a first loop control logic and assigning a confidence score to each of the variations for each of the key terms; and identifying input variations for each seed test input instruction and user input instruction with a second loop control logic and assigning a confidence score to each of the seed test and user input instructions.The procedure further involves having human reviewers manually evaluate the outputs of each of the loop control logics to identify variations for each of the key terms and for each of the seed test inputs and user input instructions.
[0020] In another aspect of the present disclosure, the method further comprises applying the coverage metrics to outputs, including: measuring a seed extension for variation of input terms; measuring a seed extension for variation of input instruction types; measuring a seed extension for substring interpretation coverage; measuring a seed extension for output robustness; and aggregating and normalizing measured seed extensions for variation of input terms, variation of input instruction types, substring interpretation, and output robustness, and generating a coverage measurement score.
[0021] In another aspect of the present disclosure, the method further comprises causing the human reviewer to evaluate the results by selectively and continuously adding or removing examples from the supplementary documents; and causing the human reviewer to convert the coverage measurement into an improved seed file and selectively and continuously iterate inputs into the control logic, which synthesizes variations that actively adapt to fill identified information gaps in order to increase the LLM tool input and output robustness from the first level to the second level, which is greater than the first level, in both production and pre-production LLM tool processes.
[0022] In further aspects of the present disclosure, a method for generating a diversified validation test suite for generative artificial intelligence (AI)-based tools comprises the execution of programmatic control logic stored in the memory of a controller comprising a processor, memory, and input / output (I / O) ports, wherein the processor executes the programmatic control logic. The programmatic control logic comprises a diversified validation test suite application (DVTSA) for a generative AI-based tool or a large language model (LLM).The DVTSA contains programmatic control logic for receiving seed test input or user input via the I / O ports, analyzing the seed test input or user input, and extracting key elements for variations. This includes: extracting syntactic information from the seed test input or user input, including: collecting all seed test input or user input coverage terms, including extracting individual key terms; quantifying the output of coverage points based on corresponding test outputs and language information for the extracted individual key terms; and automatically generating a script using the extracted key terms. The DVTSA also includes control logic for using supplemental documents and the quantified output of the coverage points to collect all coverage points.The supplementary documentation also includes: predefined language information, including linguistic databases, mathematical databases, syntactic and semantic databases, such as GridXML, ROBOT Script, natural language databases, mathematical terminology, and databases of variables, instructions, and terms defined in relation to hardware. Input coverage terms also include: predefined terms and instructions for known inputs and outputs of the procedure. The coverage metrics are variable and are actively and automatically updated, and they include an input robustness score and an output robustness score. Each input and output robustness score defines a percentage of the coverage of a list of, or all, possible output instructions of a specific type and a specific mapping of names to variables.The input robustness score further considers: variations in input terminology; variations in the nature of input instructions; and variations in substring interpretation coverage. The output robustness score defines a percentage of coverage for a list of, or all, possible output instructions and name-variable mappings. The DVTSA also includes control logic for identifying and filling information gaps in extracted syntactic information, including: reading all terms and phrases relevant to the coverage metrics from the outputs of the syntactic information extraction control logic and splitting the terms and phrases into key elements.The DVTSA also includes control logic to: identify variations for each of the keywords in a given context using a first loop control logic and assign a confidence score to each variation for each of the keywords; and identify input variations for each seed test input instruction and user input instruction using a second loop control logic, and assign a confidence score to each of the seed test and user input instructions. The DVTSA also includes control logic to instruct human reviewers to manually evaluate the outputs of the first and second loop control logics to identify variations for each of the keywords and for each of the seed test inputs and user input instructions; and to synthesize variations that actively adapt to fill any identified information gaps.The DVTSA also includes control logic for performing a coverage measurement of the extracted syntactic information, including: applying the coverage metrics to the extracted syntactic information, including: measuring a seed extension for variation of input terms; measuring a seed extension for variation of input instruction types; measuring a seed extension for substring interpretation coverage; measuring a seed extension for output robustness; and aggregating and normalizing measured seed extensions for variation of input terms, variation of input instruction types, substring interpretation, and output robustness, and generating a coverage measurement score.The DVTSA also includes control logic to instruct a human reviewer to evaluate the coverage measurement relative to predefined coverage metrics and to iterate selectively and continuously to cause the coverage measurement to increase the LLM tool input and output robustness from a first level to a second level that is greater than the first level, in both production and pre-production LLM tool processes, by instructing the human reviewer to evaluate outputs of the second control logic by selectively and continuously adding or removing examples from the supplemental documents.The DVTSA also includes control logic to instruct the human verifier to convert the coverage measurement into an improved seed file and selectively and continuously iterate inputs into the control logic, which synthesizes variations that actively adapt to fill identified information gaps. This raises the LLM tool input and output robustness from the first level to the second level, which is greater than the first, in both production and pre-production LLM tool processes. The method progressively reduces the use of computing resources, progressively increases computing efficiency, and progressively reduces reliance on the human verifier.
[0023] Further areas of application will become apparent from the description given here. It is understood that the description and the specific examples serve only for illustration and are not intended to limit the scope of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The drawings described here are for illustrative purposes only and are not intended to limit the scope of the present disclosure in any way. Fig. Figure 1 is a schematic diagram of a system for generating a diversified validation test suite for generative artificial intelligence (AI)-based tools according to an exemplary embodiment; Fig. Figure 2 is a flowchart that provides an overview of the control logic of a diversified validation test suite application (DVTSA) for generative AI-powered tools. Fig. 1 according to an exemplary embodiment; Fig. 3 is a flowchart showing the first part of the control logic of the DVTSA. Fig. 2 according to an exemplary embodiment; Fig. 4 is a flowchart that shows a second control logic part of the DVTSA from Fig. 2 according to an exemplary embodiment; and Fig. 5 is a flowchart that represents a third control logic part of the DVTSA from Fig. 2 represents an exemplary embodiment. DETAILED DESCRIPTION
[0025] The following description is for illustrative purposes only and is not intended to limit the present disclosure, its application or uses.
[0026] Fig. Figure 1 schematically shows a system 10 for generating a diversified validation test suite for generative artificial intelligence (AI)-based tools. The system 10 generally operates in or on a host device 11. The host device 11 can have a variety of forms, for example, a vehicle 12. However, it is clear that the system 10 of the present disclosure need not be bound to such a vehicle 12. Rather, the vehicle 12 is merely an exemplary, non-limiting embodiment with respect to which the system 10 of the present disclosure is described here. The system 10 can be operated in any hardware and software configuration in which a generative AI-based tool is used to process input from a user 14, such as...User commands 15, to receive and generate an output 16 that modifies the function of the hardware and / or software configuration or system in which the generative AI-based tool is used. Although the vehicle 12 shown is a passenger car, the vehicle 12 can be any type of vehicle 12 without this deviating from the scope or purpose of this disclosure. In some non-limiting examples, the vehicle 12 can be: a passenger car, truck, sport utility vehicle (SUV), semi-trailer truck, tractor trailer, tractor, combine harvester or other agricultural equipment, motorized and unmotorized aircraft such as airplanes, helicopters, gliders or autogyros, motorized and unmotorized watercraft such as: a ship, sailboat, motorboat, sports boat, jet ski, sailboat or the like.In further, non-limiting embodiments, the system 10 described herein can be adapted to function with host devices 11, such as manned and unmanned spacecraft, e.g., satellites, rockets, space stations, and other orbital and extraorbital satellite-based communication devices, without deviating from the scope or purpose of this disclosure. In further, non-limiting examples, the host devices 11 may include mobile computing platforms such as laptops, mobile phones, tablets, or other host devices 11 through which a user can operate a generative AI-based tool.
[0027] System 10 also includes a controller 18, which is a non-generalized electronic control device with a pre-programmed digital computer or processor 20, with non-transient computer-readable medium or memory 22 used for storing data such as control logic, software applications, instructions, computer code, data, lookup tables, etc., and a transceiver or input / output (I / O) ports 24. Computer-readable media or memory 22 includes all types of media that a computer can access, such as read-only memory (ROM), random access memory (RAM), hard disk drive, compact disc (CD), digital video disc (DVD), or any other type of memory 22. "Non-transient" computer-readable memory 22 excludes wired, wireless, optical, or other communication links that carry transient electrical or other signals.Non-transient computer-readable memory 22 includes media on which data can be stored permanently and media on which data can be stored and later overwritten, such as a rewritable optical disk or an erasable storage device. Computer code includes any type of program code, including source code, object code, and executable code. The processor 20 is configured to execute the code or instructions.
[0028] When the system 10 is operated in a vehicle 12, the controller 18 may include a dedicated Wi-Fi controller or an engine control module, a transmission control module, a body control module, an infotainment control module, etc. The transceiver or I / O ports 24 are configured to communicate wirelessly with a back office 26 using cellular protocols, including Global System for Mobile Communication (GSM), General Packet Radio Service (GPRS), Enhanced Data Rates for GSM Evolution (EDGE), Universal Mobile Telecommunications Services (UMTS), High Speed Packet Access (HSPA), Code-Division Multiple Access (CDMA), Evolution-Data Optimized (EV-DO / EVDO / 1xEV-DO), Short Message Services (SMS), Wi-MAX, Manufacturing Messages Specification (MMS), 2G, 3G, 4G, 5G, and wireless and cellular standards as defined in IEEE 802.1X, IEEE 802 LAN / MAN and IEEE Mobile Communication Networks Standards Committee (MobiNet-SC) standards, and the like. The back office 26 can include one or more controllers 18 and / or one or more human testers 28, which are shown and described in more detail in the following figures.
[0029] The controller 18 also contains one or more applications 30. An application 30 is a software program configured to perform a specific function or group of functions. The application 30 may contain one or more computer programs, software components, instruction sets, procedures, functions, objects, classes, instances, associated data, or a portion thereof, suitable for implementation in appropriate machine-readable program code. The applications 30 may be stored in memory 22 or in additional or separate memory. Examples of applications 30 include audio or video streaming services, games, browsers, social media, etc., as well as a diversified validation test suite application (DVTSA) 32 for a generative AI-based tool or a large language model (LLM) 34.The generative AI-based tool, or LLM tool 34, defines an application 30 that is stored either locally on the controller 18 of the vehicle 12 and / or in a remote back office 26 or a cloud-based computing device. The DVTSA 32 comprises a variety of subroutines or control logic parts.
[0030] In examples where the system 10 is operated in or on a vehicle 12, the system 10 and the DVTSA 32 can be used by the user 14, who is responsible for vehicle development, to validate the tools used to develop and test the system, which dynamically adjusts how the vehicle 12 and its functions are operated. In some examples, the vehicle 12 may be equipped with a navigation system 36, one or more drive motors 38 that provide and modify the torque supplied to the wheels 40 of the vehicle 12 to move, stop, or otherwise operate the vehicle 12, and a steering system 42 that can adjust the direction of travel of the vehicle 12. In other examples, the vehicle 12 is equipped with a braking system 44 that, when activated by the controller 18, decelerates the movement of the vehicle 12.The vehicle 12 may be equipped with a variety of other body motion control systems that may be used to modify or otherwise control the dynamic performance of the vehicle 12, including, but not limited to, aerodynamic control surfaces and actuators, active and / or semi-active suspension systems and actuators, and the like, without this deviating from the scope or purpose of the present disclosure.
[0031] In Fig. 2 and again in Fig. Figure 1 shows the DVTSA 32 in more detail in the form of a flowchart. The DVTSA 32 receives a set of seed test inputs 46 or user inputs 15, uses supplementary documentation 48 and coverage metrics 50 to analyze the seed test inputs 46 or user inputs 15 within an initial control logic 100 that synthesizes a validation test suite for the LLM-enabled tool 34. The seed inputs are typically incomplete with respect to the coverage metrics, and therefore the system 10 synthesizes the extended set of test suites with the coverage metrics in scope. The system 10 then employs a checker 54 to analyze the synthesized validation test suite 52 according to the predefined coverage metrics and to confirm the completion of the synthesized validation test suite. Then, System 10 can use the final validation test suite 56 to validate the outputs of tool 34.The generated validation test suite can also be used as a regression test suite for the tool should future updates be released. Coverage metrics can include information on edge cases, variations such as typos, repetitions, acronyms, and paraphrases in the inputs, as well as the structural scope of the outputs. The seed test inputs 46 and / or user inputs 15 can be received by a variety of hardware and software systems, such as text-based inputs (including keyboards, touchscreen interfaces, and the like) and / or audiovisual inputs to a human-machine interface (HMI), such as a microphone, touchscreen, or similar device. The seed test inputs 46 and / or user inputs 15 are received by the hardware and software systems and / or HMIs via the I / O ports 24.A second control logic 200 then performs a coverage measurement of the outputs of the first control logic 100, using the coverage metrics 50, and a third control logic 300 then prompts one or more human reviewers or validators 28 to evaluate the outputs of the second control logic 200. In some examples, the human reviewers 28 may determine that a change is required and issue a change request 58 before the first control logic 100 is reactivated. In other examples, the human validators 28 may determine that no changes are required and approve the outputs of the second control logic 200 at 60, forwarding the approved outputs to the reviewer 54 for completeness testing, and the validation test suite 56 is released for further use.
[0032] Based on Fig. 3 will be further referred to Fig. 1 and Fig. 2 The first control logic 100 is described in detail. The first control logic 100 comprises several subroutines or partial control logics. The first control logic 100 initially executes an extraction control logic 102, which extracts syntactic information from seed test inputs 46 and / or user inputs 15. More precisely, the extraction control logic 102 obtains data from several different sources, including the set of seed test inputs 46 and / or user inputs 15. The seed test inputs 46 and / or user inputs 15 can contain a variety of information, including, but not limited to, user commands to the generative AI-powered tool or the LLM tool 34.Accordingly, the seed test inputs 46 can contain language commands with which a user intends to activate various functions of the vehicle 12, such as: commands to turn on a vehicle navigation system, commands to search for a point of interest, commands to change the temperature in the vehicle interior, commands to stop, wait, or delay, requests to receive a response to a mathematical instruction, or other commands of this type. The extraction control logic 102 collects all input terms with coverage range in block 104, including the extraction of syntactic information from the seed test inputs 46. In a non-restrictive example, a seed test input 46 or a user input 15 such as "Model Delay 1200 milliseconds" or "Model Delay 1200 milliseconds" can be extracted separately as input coverage terms 105 "Model", "Delay", "1200", and "milliseconds".That is, the first control logic 100, more precisely the extraction control logic 102, decomposes the data set of seed test inputs 46 into different terms.
[0033] From block 104, the extraction control logic 102 passes to block 106, where it uses the corresponding test outputs for all seed tests and the language information from block 108 to quantify the output coverage points 109. The output coverage points 109 from block 106 extract all language constructs available as part of the expected output for the given seed inputs, e.g., all types of "action" statements, in particular: 'Action': 'Model.Delay', 'MILLI_SECOND' : '1200'. 'ACTION': 'Execute. Test", "TEST_FILE': '\\Functions\\PreTest.gridXmI', 'INPUT VALUES': 'return_vars'
[0034] From block 106, the extraction control logic 102 transitions to block 110, where it uses the supplemental documentation 48 from block 112 to gather and define required coverage points 113. The supplemental documentation 48 in block 112 can contain a variety of information, such as the names of software variables and their descriptions. In several non-restrictive examples, the supplemental documentation 48 can be described as a variably modifiable or dynamically updated and modified dictionary or glossary of terms that define specific variable names and meanings within the set of seed test inputs 46. In block 110, names, descriptions, and types of software variables, calibrations, constraints, and the like are extracted and / or applied to the automatically generated script.Once the extraction control logic 102 is complete, the automatically collected information is sent from the extraction control logic 102 to an identification control logic 114.
[0035] Based on Fig. 4 will be further referred to Fig. Figures 1-3 describe the identification control logic 114 in detail. In general terms, the identification control logic 114 identifies and fills information gaps, or, if necessary, generates additional test inputs based on the seed inputs to close the gap in the coverage metrics of the original seed test inputs 46. The identification control logic 114 begins in block 116 by retrieving all extracted information from 102. In block 118, the identification control logic 114 reads all terms and phrases relevant to the coverage metrics 50 and divides the terms and phrases into key elements.To determine which terms and phrases are relevant for the coverage metrics 50, and to divide such terms and phrases into key elements, the identification control logic 114 refers to input coverage terms 105, output coverage terms 109, and variables stored in a database 120 containing predefined but updatable terms. In some non-restrictive examples, the database 120 contains predefined linguistic, mathematical, syntactic, and semantic databases with language information, such as GridXML, ROBOT Script, or similar. The database 120 may also contain databases with natural language, mathematical terminology, and / or variables, or similar.
[0036] The identification control logic 114 then proceeds to a first and a second processing loop 122 and 124, which can be executed in parallel, concurrently, sequentially, periodically, or upon the occurrence of a triggering condition or event, without this deviating from the scope or purpose of the present disclosure. Starting from block 126, the first processing loop 122 calls the LLM engine 34 or any natural language processing tool to identify the keywords extracted in block 102 and read by block 118. In a non-restrictive example, the extracted terms 129 can include a subject, an object, an action, unit terms, and the like, such as: {Model, Delay, 1200, Milliseconds}.Subsequently, in block 128, the first processing loop 122 calls the LLM tool 34 to identify variations 131 for each of the extracted terms in the relevant contextual situation and assigns a confidence score to the variations. Continuing the example above {Model, Delay, 1200, Milliseconds}, plausible descriptive variations could be identified equivalent terms such as: • Model: Simulation, system, prototype • Delay: Backlog, waiting, pause • 1200: Twelve hundred, one thousand two hundred, 1.2 k • Milliseconds: 0.001 s, ms, 1 / 1000 s or similar.
[0037] From block 128, the first processing loop 122 moves to block 130, where human reviewers 28 perform a manual evaluation and then allow the first processing loop 122 to proceed to block 132. The confidence score of the generated terms helps the human reviewers 28 to effectively filter the terms. The human reviewers 28 have the ability to add to, edit, or remove terms from the generated list. The human reviewers 28 then instruct the first processing loop 122 to continue execution and iterate until an acceptable number of valid variations of the terms have been generated, at which point the first processing loop 122 exits block 132. It becomes clear that the confidence level required for a given set of seed test inputs 46 varies depending on the criticality of the systems on which the seed test inputs 46 are intended to act.In a non-restrictive example, for inputs relating to autonomous driving behavior or advanced driver assistance systems (ADAS or advanced driver assistance systems), the threshold for the confidence score may be significantly higher than for non-critical or non-safety-related system functions, such as the settings of the vehicle's entertainment system.
[0038] The second processing loop 124 begins in block 134, where the identification control logic 114 lists all missing syntactic instructions and variable descriptions. In a non-restrictive example, an input 135 can be represented as follows: {"Action": "if", "IF_CONDITION": <varname>== Val'}. The missing syntactic elements in block 134 are identified in the outputs of extraction control logic 102 with respect to seed test input 46. Subsequently, in block 136, identification control logic 104 identifies key input elements for variations, including algebraic expressions, variable names, and the like. With identified key input elements, system 10 uses, for example, LLM tool 34 to generate more plausible variations such as "delay" instead of "wait," and can also simplify complex mathematical expressions. Continuing the example {Action...} above, in block 136, LLM tool 34 is used to check at 137 according to "Check <vars>for Values" or similar. In combination with the missing syntactic elements identified in Block 134 and the variations identified in Block 136, the identification control logic 114 obtains the coverage metrics 50 from a database or other such storage. The coverage metrics 50 may be fixed, or in some exemplary embodiments, the coverage metrics 50 are variable and are actively and automatically updated. The coverage metrics 50 may also vary between the applications or systems involved in responding to the seed test input data 46. In additional aspects, the coverage metrics 50 include an input robustness score and an output robustness score. The input robustness score accounts for variations in terms, variations in the types of input instructions, and the coverage of the interpretation substring or substring.In some examples, variations in terms may include typos, abbreviations, synonyms, and the like. In some specific, non-restrictive examples, synonyms may be wait / delay or read / receive, or similar. Variations in input instructions include variations in algebraic expressions, which may range from simple expressions to complex expressions, multi-line output, and the like. A specific, non-restrictive example of variation in input instructions is "state A to B with input X," which can also be expressed as "given input X, transition to state B from state A," or similar. Substring interpretation coverage may include instructions for performing specific tasks, such as "replacing 'set' with '=".However, substring interpretation coverage is also specifically designed so that terms containing the substring within a larger string are not replaced. In the example where "set" is replaced by "=", the terms "reset", "dataset", "preset", and "subset" are neither replaced nor changed. In several respects, the output robustness score is a percentage of the coverage of the list, or of all possible output statements of a given type, and the mapping of names to variables.
[0039] For safety-critical systems, such as vehicle powertrain systems, braking systems, steering systems, advanced driver assistance systems (ADAS), and the like, the Coverage Metric 50 may require that the output Robustness Rating of System 10 be at least 95 percent (95%), whereas for non-safety-critical systems, an accuracy of at least 90 percent (90%) may only be required. It is understood that the Robustness Ratings of 95% and 90% listed above are merely non-restrictive examples of possible Coverage Metric 50.
[0040] From block 136, the second processing loop 124 passes to block 130, where the human reviewers 28 perform a manual evaluation of the outputs of the second processing loop 124 and either allow the second processing loop 124 to pass to block 132 if the confidence score for the identified input variations in block 136 is sufficiently high, or, if a threshold for the confidence score has not been reached, the human reviewers 28 cause the second processing loop 124 to be executed and repeated until the identified input variations in block 136 have reached the threshold for the confidence score, at which point the second processing loop 124 passes to block 132.As mentioned previously, the confidence level required for a given set of seed test inputs 46 varies depending on the criticality of the systems on which the seed test inputs 46 are intended to act. In a non-restrictive example, for inputs relating to autonomous driving behavior or advanced driver assistance systems (ADAS or ADAS), the threshold for the confidence score may be significantly higher than for non-critical or non-safety-related system functions, such as the settings of the vehicle's entertainment system 12.
[0041] As in Fig. 3 and Fig. 5 and with further reference to the Fig. 1, Fig. 2 and Fig. As shown in Figure 4, after completion of the identification control logic 114, the first control logic transitions to a variation-synthesis control logic 138, which in Fig. 5 is shown in detail.
[0042] The variation-synthesis control logic 138 receives term variations and context information from block 128 of the first control loop 122 and input variations or variables from block 136 of the second control loop 124. In block 140, variations or rewrites are synthesized using the LLM tool 34, along with permutations and combinations generated as shown in output block 142, where: {Model, Delay, 1200, Milliseconds} is defined as equivalent to: • Simulation Waiting 1200 ms • System pause 1200 ms • <vars>check values •<Var1 * Var2 + Konstante> check values • Check if <var1>and <var2>are the same.
[0043] Subsequently, the DVTSA 32 executes the second control logic 200 to perform a coverage measurement 202 of the outputs of the first control logic 100 using the coverage metrics 50. The second control logic 200 begins in block 204. In block 206, the coverage measurement takes into account variations in input terms by measuring the extent or extension of the seed term 46. In one example, the extension of the seed term 46 can range from a first seed term 46 to a hundredth seed term 46; however, other embodiments and sets of seed terms 46 are said to fall within the scope of this disclosure. From block 206, the second control logic 200 proceeds to block 208, where variations in the input instruction types are taken into account with a multitude of seed instructions.The seed instructions can range from a first to a fiftieth seed instruction, although additional or fewer seed instructions are deemed to fall within the scope of this disclosure. Subsequently, in block 210, the second control logic 200 performs substring interpretation coverage to verify that substrings are accurately represented and not erroneously replaced by a substring seed extension, which may contain up to five substrings. However, additional or fewer substring seeds are also conceivable. In block 212, the second control logic 200 computes output robustness to verify the quality and reliability of the outputs of the first control logic 100 and the second control logic 200 with respect to the seed input terms 46. In several examples, output robustness is measured via an output robustness seed extension, which includes up to five output robustness seeds.However, additional or fewer output robustness seeds are also conceivable. In block 214, the second control logic 200 aggregates and normalizes the values of the seed extensions performed in blocks 206, 208, 210, and 212, and receives the result of the coverage measurement score as output in block 216, where the second control logic 200 terminates. The coverage of each category can be restricted to generate a limited number of variations, and it can be weighted for coverage measurement and can also be a qualitative measure based on the user or applications. The coverage measurement score 216 and the coverage metrics 50 are then compared both within the second control logic 200 and within the third control logic 300, with the coverage measurement score 216 and the coverage metrics 50 being manually evaluated by one or more human reviewers 28.The human checkers 28 can perform a variety of functions, but in some specific, non-restrictive examples, the human checkers 28 eliminate examples or duplicates, add information or examples, and convert or otherwise implement improved seed files, which are then used within the first and second control logics 100, 200, and especially within the extraction control logic 102 and the identification control logic 114.
[0044] By using System 10 and the methodology, including the use of DVTSA 32 of this disclosure, a variety of advantages are realized in systems that use LLM tools 34 to perform or assume a variety of activities and tasks. DVTSA 32 provides a continuously updated, robust method for validating the responses of the LLM-based tool 34 to user inputs 15 in both the production and pre-production processes. That is, in a pre-production or prototype application, DVTSA 32 can be used by technical users to ensure that the analysis of the seed test inputs 46 is thorough and that key elements are identified and extracted despite variations in syntax, language, expression, grammar, or the like.The DVTSA 32 enables a set of robust processes that quantify output elements and ensure a correct and accurate match with plausible inputs, as well as synthesize coherent inputs based on individual variations. Furthermore, the DVTSA 32 performs the aforementioned processes automatically but includes a human evaluation loop in which the inputs and outputs of the DVTSA 32 are refined to eliminate potential sources of error or duplication of effort. Accordingly, in a pre-production environment or application, System 10 and the DVTSA 32 of this disclosure refine the LLM tool 34 to ensure that engineers communicating with System 10 or Vehicle 12 via the LLM tool 34 are correctly understood by the LLM tool 34, regardless of differing communication styles and content.Engineers can interact with the systems 10 of a host device 11 or vehicle 12 via the DVTSA 32 and the LLM tool 34 to train or program potential production databases for responses such as supplemental documentation 48, coverage metrics 50 and the like, which will be used in future production-based use by customers.
[0045] Furthermore, the DVTSA 32 offers automation benefits that reduce the amount of human effort and man-hours required to create validation and evaluation test suites for LLM-based tools 34, in particular by: automatically generating scripts based on seed test inputs 46, using the identification control logic 104 to identify and fill information gaps in the input instructions of system users 10, using the variation synthesis control logic 138 to generate expected inputs to fill the information gaps identified in the identification control logic 104, and applying coverage metrics 50 and the in-loop checker 28 to ensure that the LLM-based tool 34 accurately understands the instructions of the user input 15 and produces a suitable, accurate, reliable, and robust output in response to the instructions of the user input 15.Furthermore, the automatic scripting of DVTSA 32 reduces the potential for human-caused typographical, syntactic, or other errors from a first level to a second level that is significantly lower than the first. The automatically generated test case or script also enables more thorough validation, higher accuracy, faster processing, and less time spent validating responses from the LLM-supported tool 34. Because DVTSA 32 is continuously evaluated and updated in both pre-production and production, the number of interactions and inputs from the human reviewer 28 is reduced.This means that even in a production application where an end user or customer who is not an engineer interacts with the LLM-based tool 34, the DVTSA 32 operates in such a way that it accurately, consistently, reliably and robustly interprets the inputs of the end user or customer into the system 10 and accordingly generates an appropriate response, with progressively reduced use of computing resources, progressively increased computing efficiency and progressively reduced dependence on validation processes of human reviewers 28.
[0046] The description of the present revelation is merely exemplary, and variations that do not deviate from the core of the present revelation shall fall within its scope of protection. Such variations are not to be considered a deviation from the spirit and scope of the present revelation. < / vars> < / vars> < / varname>
Claims
[1] System for generating a diversified validation test suite for generative artificial intelligence (AI)-based tools, wherein the system comprises: a controller comprising a processor, memory and input / output (I / O) ports, wherein the processor executes programmatic control logic stored in the memory, wherein the programmatic control logic comprises a diversified validation test suite application (DVTSA) for a generative AI-based tool or a large language model (LLM), wherein the DVTSA comprises: a first control logic that receives a seed test input or user input via the I / O ports, analyzes the seed test input or user input, and extracts key elements for variations; a second tax logic that performs a coverage measurement of expenditures from the first tax logic; and A third control logic that causes a human reviewer to evaluate outputs of the second control logic relative to predefined coverage metrics and to iterate selectively and continuously to cause outputs of the second control logic to increase the LLM tool input and output robustness from a first level to a second level that is greater than the first level, in both production and pre-production LLM tool processes, with the system progressively reducing the use of computing resources, progressively increasing computing efficiency, and progressively reducing the dependence on the human reviewer. [2] System according to claim 1, wherein the first control logic further comprises: a control logic for extracting syntactic information from the seed test input or the user input; a control logic that identifies and fills information gaps in extracted syntactic information; and a control logic that synthesizes variations that actively adapt to close identified information gaps. [3] System according to claim 2, wherein the control logic for extracting syntactic information further comprises: Control logic that collects all seed test input or user input coverage terms, including the extraction of individual keywords; Control logic that quantifies the output of coverage points based on corresponding test outputs and language information for the extracted individual keywords; Control logic that automatically generates a script using each extracted keyword; and Control logic that uses supplementary documents and the quantified output of coverage points to collect all coverage points. [4] System according to claim 3, wherein the supplementary documents further comprise: Predefined language information, including linguistic databases, mathematical databases, syntactic and semantic databases, such as GridXML, ROBOT Script, natural language databases, mathematical terminology, and databases of variables, instructions, and terms defined in relation to the system hardware. [5] System according to claim 3, wherein the input coverage terms comprise: predefined terms and instructions for known inputs and outputs of the system; and wherein the coverage metrics are variable and are actively and automatically updated, and wherein the coverage metrics include an input robustness score and an output robustness score, each of the input and output robustness scores defining a percentage of the coverage of a list of or all possible output instructions of a given type and a given name-to-variable mapping. [6] System according to claim 5, wherein the input robustness rating further takes into account: Variations in input terminology; Variations in the type of input instructions; and Variations in the coverage of the substring interpretation; and where The output robustness rating includes: a percentage of the coverage of a list of or all possible output instructions and the mapping of names to variables. [7] System according to claim 3, wherein the control logic that identifies and fills information gaps further comprises: Control logic that reads all terms and phrases relevant to the coverage metrics from the outputs of the control logic for extracting syntactic information and divides the terms and phrases into key elements; a first loop control logic that identifies variations for each of the keywords in a given context and assigns a confidence score to each of the variations for each of the keywords; and a second loop control logic that identifies input variations for each seed test input instruction and user input instruction and assigns a confidence score to each of the seed test input and user input instructions. [8] System according to claim 5, further comprising: Control logic that causes human reviewers to manually evaluate the outputs of each of the first and second loop control logics to identify variations for each of the key terms and for each of the seed test input and user input instructions. [9] System according to claim 5, wherein the second control logic further comprises: Tax logic that applies the coverage metrics to outputs of the first tax logic, including: Measuring a seed extension for varying input terms; Measuring a seed expansion for a variation of input statement types; Measuring a seed extension to cover the substring interpretation; Measuring a seed extension for output robustness; and Aggregating and normalizing measured seed extensions for variation of input terms, variation of input instruction types, substring interpretation, and output robustness, and generating a coverage measurement score. [10] System according to claim 4, wherein the third control logic further comprises: Tax logic that causes the human auditor to evaluate expenditures of the second tax logic by selectively and continuously adding or removing examples from the supplemental documents; and Control logic that causes the human tester to convert the coverage measurement into an improved seed file and selectively and continuously iterating inputs into the control logic that synthesizes variations that actively adapt to fill identified information gaps in order to increase the LLM tool input and output robustness from the first level to the second level, which is greater than the first level, in both production and pre-production LLM tool processes.