Soft-prefix-based large language model question answering method and device

By constructing a vector retrieval library and utilizing attention weight information, soft prefixes are retrieved and injected to correct the focus of large language models, thus solving the logical error problem, improving output accuracy, and reducing computational costs.

CN121766461BActive Publication Date: 2026-06-19BEIJING UNIV OF POSTS & TELECOMM
View PDF -1 Cites -1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING UNIV OF POSTS & TELECOMM
Filing Date
2026-03-05
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing large language models are prone to logical errors when processing complex multi-step logic, resulting in inaccurate output results. Furthermore, existing correction methods are inefficient or rely on high-cost computing resources, making them difficult to apply effectively in high-precision fields.

Method used

By constructing a vector retrieval library and utilizing the attention weight information in the reasoning process of large language models, lightweight soft prefixes are retrieved and injected to directly correct the model's focus and improve output accuracy.

Benefits of technology

It improves the output accuracy of large language models, reduces computational resource consumption, adapts to different downstream tasks, and lowers correction costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121766461B_ABST
    Figure CN121766461B_ABST
Patent Text Reader

Abstract

This invention provides a large language model question answering method and apparatus based on soft prefixes. A specific implementation of the method includes: inputting a user question into a large language model; obtaining an initial answer from the large language model containing multiple intermediate inference steps; determining whether errors exist in the multiple intermediate inference steps; if so, identifying a target erroneous step based on the erroneous steps, and obtaining target attention weight information corresponding to the target erroneous step; retrieving matching retrieval attention weight information from a vector retrieval library based on the target attention weight information corresponding to the target erroneous step, and obtaining a soft prefix corresponding to the retrieval attention weight information matching the target attention weight information as the target soft prefix; fusing the target soft prefix with the vector representation corresponding to the user question to obtain a mixed sequence; inputting the mixed sequence into the large language model, and having the large language model output a corrected answer.
Need to check novelty before this filing date? Find Prior Art