LLM-Based Software QA for Automated Security and Reliability Testing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Maintaining quality and security of feature-rich and user-interactive software has become cumbersome and complex, requiring more time and effort in quality-assurance testing and hardening against attackers.
Innovation Solution
An AI/ML system using large language models (LLMs) performs automated software quality, safety, and security assurance by simulating interactions with test software to identify vulnerabilities and defects, evaluating software behavior without human intervention.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If traditional quality-assurance testing methods are used on feature-rich and user-interactive software, then testing coverage can be achieved, but the process becomes cumbersome and time-consuming
Solution Approach 1:
The patent replaces traditional mechanical/manual quality assurance processes with an AI-based automated system. The AI actor and AI evaluator automatically perform testing tasks that previously required human testers, thereby reducing testing time while maintaining comprehensive coverage of feature-rich software systems.
Solution Approach 2:
The AI-based quality assurance system performs self-service by automatically generating test inputs, executing tests, evaluating results, and identifying vulnerabilities without requiring continuous human intervention. This automation enables the system to handle complex software testing independently, reducing both time and manual effort.
2Adaptability or versatility
If software becomes more feature-rich and user-interactive, then functionality and user experience improve, but maintaining quality and security becomes increasingly difficult and complicated
Solution Approach 1:
The AI-based quality assurance system provides universal testing capabilities that can handle diverse software types and functionalities through a single platform. The AI actor can adapt to different software interfaces and behaviors, while the AI evaluator can assess multiple quality dimensions (security, functionality, user experience) simultaneously, simplifying the complex task of testing feature-rich software.
Solution Approach 2:
The system dynamically adjusts testing parameters and strategies based on the specific software being tested. The AI can modify test input characteristics, evaluation criteria, and exploration depth according to the software's features and complexity, enabling effective quality assurance across varied software types without requiring separate specialized processes for each.
3Object-affected harmful factors
If comprehensive security hardening is performed against attackers, then software security improves, but the process increases in complexity
Solution Approach 1:
The AI actor performs preliminary security testing by simulating attacker behaviors and identifying vulnerabilities before actual attacks can occur. The system proactively explores security weaknesses, tests authentication mechanisms, and evaluates vulnerability responses in advance, enabling security hardening to be performed systematically rather than reactively.
Solution Approach 2:
The AI-based system acts as an intermediary between potential attackers and the software being tested. It safely simulates attack scenarios, evaluates security responses, and provides feedback on vulnerabilities without requiring actual malicious attacks. This intermediary approach simplifies security testing by providing a controlled, automated mechanism for assessing and hardening software security.
Data Source
AI summary
Systems and methods are provided for implementing quality assurance for digital technologies using language model (“LM”)-based artificial intelligence (“AI”) and/or machine learning (“ML”) systems. In various embodiments, a first prompt is provided to an LM actor or attacker to cause the LM actor or attacker to generate interaction content for interacting with test software. Responses from the test software are then evaluated by an LM evaluator to produce evaluation results. In some examples, a second prompt is generated that includes the responses from the test software along with the evaluation criteria for the test software. When the second prompt is provided to the LM evaluator, the LM evaluator generates the evaluation results.


