Large Language Model Response Sequence Generation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Large language models (LLMs) require significant computational resources and face challenges in reliably performing natural language processing tasks without necessitating additional follow-up inputs.

Innovation Solution

Implementing techniques that allow an LLM to generate a sequence of responses in a single inference call, using an attention mechanism or other memory to improve each subsequent response, and optionally generating critique responses to further refine outputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If an LLM generates responses iteratively through multiple inference calls, then the quality of responses can be improved, but the computational resources and time consumed increase significantly

Engineering Contradiction:
Improveresponse qualityVSAvoidcomputational resources
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The patent segments the response generation process into multiple candidate responses generated in parallel within a single inference call, rather than generating responses sequentially through multiple inference calls. This allows the system to explore multiple response paths simultaneously while consuming fewer computational resources overall.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary actions by generating multiple candidate responses and selecting the best one before final output, all within a single inference call. This preliminary exploration of multiple response options enables quality improvement without the need for iterative refinement through additional inference calls.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If an LLM generates comprehensive responses, then the likelihood of requiring follow-up inputs decreases, but the time and computational resources required to generate the response increase

Engineering Contradiction:
Improvecompleteness of responseVSAvoidresponse generation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent segments response generation into multiple parallel candidate responses, allowing the model to explore different levels of detail and approaches simultaneously. This enables comprehensive coverage of possible answer spaces without sequentially iterating through each possibility, thus reducing overall generation time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent maintains continuous useful action by generating multiple candidate responses in parallel within a single inference call, rather than stopping after one response and requiring additional calls for refinement. This continuous parallel processing ensures comprehensive responses are generated efficiently without time loss from iterative calls.

Inventive Principle:
Principle #20Continuity of useful action

3Manufacturing precision

If an LLM generates multiple candidate responses in parallel, then the quality and completeness of the final response improve, but the device complexity increases

Engineering Contradiction:
Improveresponse qualityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent applies universality by using the same LLM infrastructure to perform multiple functions: generating candidate responses, evaluating them, and selecting the best one, all within a single inference call. This multi-functionality approach improves response quality without requiring separate specialized systems for each task, thus limiting the increase in device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Reliability

If the LLM processes more data through additional inference calls, then the reliability of task completion improves, but the productivity decreases

Engineering Contradiction:
Improvetask completion reliabilityVSAvoidresponse generation efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments the task completion process into multiple candidate responses generated in parallel within a single inference call, ensuring thorough exploration of solution spaces while maintaining high productivity. This segmentation allows reliable task completion without the productivity loss associated with sequential iterative calls.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent ensures continuous useful action by processing multiple candidate responses in parallel within a single inference call, maximizing the utilization of computational resources. This continuous processing maintains high productivity while achieving reliable task completion through comprehensive evaluation of multiple response options.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS20250045534A1Efficient training and utilization of large language models
Publication Date: 2025.02.06 GOOGLE LLC
  • US20250045534A1 patent drawing
  • US20250045534A1 patent drawing
  • US20250045534A1 patent drawing

AI summary

Implementations relate to a method implemented by one or more processors, the method including: receiving natural language (NL) based input associated with a client device; generating, using a large language model (LLM) and based on processing the NL based input, LLM output; determining, based on the LLM output, a sequence of LLM responses, the sequence of LLM responses including at least one intermediate LLM response and a final LLM response. In some implementations, the method may further include causing the final LLM response to be rendered at the client device. In additional or alternative implementations, the method may further include storing, as an instance of training data for fine-tuning the LLM or an additional LLM, the NL based input along with the final LLM response.