Large Language Model Response Sequence Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Large language models (LLMs) require significant computational resources and face challenges in reliably performing natural language processing tasks without necessitating additional follow-up inputs.
Innovation Solution
Implementing techniques that allow an LLM to generate a sequence of responses in a single inference call, using an attention mechanism or other memory to improve each subsequent response, and optionally generating critique responses to further refine outputs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If an LLM generates responses iteratively through multiple inference calls, then the quality of responses can be improved, but the computational resources and time consumed increase significantly
Solution Approach 1:
The patent segments the response generation process into multiple candidate responses generated in parallel within a single inference call, rather than generating responses sequentially through multiple inference calls. This allows the system to explore multiple response paths simultaneously while consuming fewer computational resources overall.
Solution Approach 2:
The patent performs preliminary actions by generating multiple candidate responses and selecting the best one before final output, all within a single inference call. This preliminary exploration of multiple response options enables quality improvement without the need for iterative refinement through additional inference calls.
2Reliability
If an LLM generates comprehensive responses, then the likelihood of requiring follow-up inputs decreases, but the time and computational resources required to generate the response increase
Solution Approach 1:
The patent segments response generation into multiple parallel candidate responses, allowing the model to explore different levels of detail and approaches simultaneously. This enables comprehensive coverage of possible answer spaces without sequentially iterating through each possibility, thus reducing overall generation time.
Solution Approach 2:
The patent maintains continuous useful action by generating multiple candidate responses in parallel within a single inference call, rather than stopping after one response and requiring additional calls for refinement. This continuous parallel processing ensures comprehensive responses are generated efficiently without time loss from iterative calls.
3Manufacturing precision
If an LLM generates multiple candidate responses in parallel, then the quality and completeness of the final response improve, but the device complexity increases
Solution Approach 1:
The patent applies universality by using the same LLM infrastructure to perform multiple functions: generating candidate responses, evaluating them, and selecting the best one, all within a single inference call. This multi-functionality approach improves response quality without requiring separate specialized systems for each task, thus limiting the increase in device complexity.
4Reliability
If the LLM processes more data through additional inference calls, then the reliability of task completion improves, but the productivity decreases
Solution Approach 1:
The patent segments the task completion process into multiple candidate responses generated in parallel within a single inference call, ensuring thorough exploration of solution spaces while maintaining high productivity. This segmentation allows reliable task completion without the productivity loss associated with sequential iterative calls.
Solution Approach 2:
The patent ensures continuous useful action by processing multiple candidate responses in parallel within a single inference call, maximizing the utilization of computational resources. This continuous processing maintains high productivity while achieving reliable task completion through comprehensive evaluation of multiple response options.
Data Source
AI summary
Implementations relate to a method implemented by one or more processors, the method including: receiving natural language (NL) based input associated with a client device; generating, using a large language model (LLM) and based on processing the NL based input, LLM output; determining, based on the LLM output, a sequence of LLM responses, the sequence of LLM responses including at least one intermediate LLM response and a final LLM response. In some implementations, the method may further include causing the final LLM response to be rendered at the client device. In additional or alternative implementations, the method may further include storing, as an instance of training data for fine-tuning the LLM or an additional LLM, the NL based input along with the final LLM response.


