System and method for accelerating large model question and answer output

By introducing a hot question intelligent search module and a Q&A cache server in the Q&A system, the problem of users consuming a lot of computing power each time they ask questions in the existing technology is solved, and the fast matching and cache of hot questions is achieved, which improves the Q&A output efficiency.

CN120337942APending Publication Date: 2025-07-18ZHONGKE URBAN BRAIN DIGITAL TECH (WUXI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510507366.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

The existing question-and-answer system needs to conduct multi-level and multi-level reasoning and analysis every time a user asks questions, which consumes a lot of computing power and cannot effectively cache hot issues, resulting in inefficient Q&A output.

Method used

A combined system of intelligent terminal devices and large-scale Q&A servers is adopted, which includes a hot question intelligent search module and a Q&A cache server to realize the cache and fast search of hot questions, and extract answers from the cache through question matching.

Benefits of technology

It improves the user's Q&A response speed, reduces the consumption of large-scale inference computing power, realizes fast matching and caching of hot issues, and improves the Q&A output efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337942A_ABST
    Figure CN120337942A_ABST
Patent Text Reader

Abstract

The invention discloses a system and method for accelerating large-model question and answer output, the system comprises an intelligent terminal device and a large-model question and answer server, the intelligent terminal device comprises an intelligent question and answer client module, the intelligent question and answer client module is in interactive connection with a question receiving and classification preprocessing module through a wireless network, and the question receiving and classification preprocessing module is connected with the large-model question and answer server. The question receiving and classification preprocessing module is arranged in the large model question and answer server, the question receiving and classification preprocessing module is interactively connected with the intelligent hotspot question retrieval module, and the intelligent hotspot question retrieval module is interactively connected with the question and answer cache server; according to the system, hot questions and answers can be cached, meanwhile, the hot questions can be quickly retrieved, the corresponding answers can be accurately extracted from the cache through question matching, the question answering response speed of a user is greatly increased, and the user experience is improved. And the calculation power consumption of large model reasoning is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information retrieval and query, and particularly relates to a system and method for accelerating the question-and-answer output of a large model. Background Art

[0002] In the current question-and-answer system during use, a user's question is sent through an intelligent terminal device, and then the large model question-and-answer server conducts knowledge and network searches on the question as needed. Finally, the question format of the inference output is transmitted to the intelligent terminal device to complete the question-and-answer operation.

[0003] However, in the current question-and-answer system, every time a user asks a question, it is necessary to perform multi-level and multi-layer reasoning and analysis on the question. During the re-analysis and reasoning process, a large amount of computing power is consumed, and hot issues cannot be buffered and stored, nor can hot issues be quickly matched, which affects the question-and-answer output efficiency. For this reason, we propose a system and method for accelerating the question-and-answer output of a large model. Summary of the Invention

[0004] The purpose of the present invention is to provide a system and method for accelerating the question-and-answer output of a large model, so as to solve the problem proposed in the above background art that every time a user asks a question in the current question-and-answer system, it is necessary to perform multi-level and multi-layer reasoning and analysis on the question. During the re-analysis and reasoning process, a large amount of computing power is consumed, and hot issues cannot be buffered and stored, nor can hot issues be quickly matched, which affects the question-and-answer output efficiency.

[0005] To achieve the above purpose, the present invention provides the following technical solution: A system for accelerating the question-and-answer output of a large model includes an intelligent terminal device and a large model question-and-answer server. An intelligent question-and-answer client module is included in the intelligent terminal device. The intelligent question-and-answer client module is interactively connected to a question receiving and classification preprocessing module through a wireless network. The question receiving and classification preprocessing module is arranged in the large model question-and-answer server. The question receiving and classification preprocessing module is interactively connected to a hot issue intelligent retrieval module. The hot issue intelligent retrieval module is interactively connected to a question-and-answer cache server. The hot issue intelligent retrieval module is also connected to an intelligent temperature client module. Both the hot issue intelligent retrieval module and the question-and-answer cache server are arranged in the large model question-and-answer server.

[0006] Preferably, the hot issue intelligent retrieval module is used for real-time retrieval of the features in hot issues.

[0007] Preferably, the question-and-answer cache server is used for storing the question-and-answer results of corresponding hot issues.

[0008] Preferably, the problem receiving and classification preprocessing module is also interactively connected to a knowledge base retrieval service module and an online search service module, and the knowledge base retrieval service module and the online search service module are arranged in the large model Q&A server.

[0009] Preferably, the problem receiving and classification preprocessing module is also connected to an inference service module, and the inference service module is arranged in the large model Q&A server.

[0010] Preferably, the inference service module is also connected to an output processing module, the output inference module is respectively interactively connected to a Q&A cache server and an intelligent Q&A client module, and the output processing module is arranged in the large model Q&A server.

[0011] Preferably, the intelligent Q&A client module is also connected to a display screen and a speaker, and the display screen and the speaker are arranged in an intelligent terminal device.

[0012] Preferably, the display screen is used to display the text information output by the large model Q&A server, and the speaker is used to play the voice generated by converting the replied text of the large model Q&A server.

[0013] A method for accelerating the output of large model Q&A includes the following steps:

[0014] S1. The user inputs question information through the intelligent Q&A client module and sends the question information to the problem receiving and classification preprocessing module in the large model Q&A server;

[0015] S2. The problem receiving and classification preprocessing module classifies the question and preferentially triggers a hot topic question query, and the hot topic intelligent retrieval module retrieves the hot topic information;

[0016] S3. When a hot topic question is hit, extract the answer corresponding to the question from the Q&A cache server;

[0017] S4. The hot topic intelligent retrieval module directly outputs the corresponding answer to the intelligent Q&A client module;

[0018] S5. The display screen displays the text information output by the large model Q&A;

[0019] S6. The speaker plays the voice generated by converting the replied text of the large model Q&A;

[0020] S7. When the hot topic question is not hit, continue to retrieve in the knowledge base retrieval service module or, as needed, in the online search service module;

[0021] S8. The problem receiving and classification preprocessing module inputs the user's original question and the retrieved information into the inference service module;

[0022] S9. The inference service module converts the vector output from the inference into a user-readable format and sends the information in the user-readable format to the output processing module;

[0023] S10. The output processing module submits the question and the answer to the Q&A cache server for hot issue recognition. After recognition, the processed Q&A is output to the intelligent Q&A client module;

[0024] S11. The display screen shows the text information output from the large model Q&A;

[0025] S12. The speaker plays the voice generated by converting the text replied by the large model.

[0026] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0027] The present invention can cache hot issues and answers, and can quickly retrieve hot issues. Through question matching, the corresponding answers can be accurately extracted from the cache, greatly improving the user Q&A response speed and reducing the consumption of large model inference computing power. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is the overall structural schematic diagram of the system for accelerating the large model Q&A output;

[0029] Figure 2 is the overall structural schematic diagram of the existing system for accelerating the large model Q&A output; DETAILED DESCRIPTION OF THE EMBODIMENTS

[0030] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0031] Please refer to Figure 1 , the present invention provides a technical solution: a system for accelerating the large model Q&A output, including an intelligent terminal device and a large model Q&A server. The intelligent terminal device includes an intelligent Q&A client module. The intelligent Q&A client module is interactively connected with the question receiving and classification preprocessing module through a wireless network. The question receiving and classification preprocessing module is arranged in the large model Q&A server. The question receiving and classification preprocessing module is interactively connected with the hot issue intelligent retrieval module. The hot issue intelligent retrieval module is interactively connected with the Q&A cache server. The hot issue intelligent retrieval module is also connected to the intelligent temperature client module. Both the hot issue intelligent retrieval module and the Q&A cache server are arranged in the large model Q&A server.

[0032] As a specific embodiment of the present application, the hot issue intelligent retrieval module is used to retrieve the features in the hot issues in real time.

[0033] As a specific embodiment of the present application, the Q&A cache server is used to store the Q&A results of the corresponding hot issues.

[0034] As a specific embodiment of the present application, the question receiving and classification preprocessing module is also interactively connected to the knowledge base retrieval service module and the online search service module, and the knowledge base retrieval service module and the online search service module are arranged in the large model Q&A server.

[0035] As a specific embodiment of the present application, the question receiving and classification preprocessing module is also connected to the inference service module, and the inference service module is arranged in the large model Q&A server.

[0036] As a specific embodiment of the present application, the inference service module is also connected to the output processing module, the output inference module is respectively interactively connected to the Q&A cache server and the intelligent Q&A client module, and the output processing module is arranged in the large model Q&A server.

[0037] As a specific embodiment of the present application, the intelligent Q&A client module is also connected to the display screen and the speaker, and the display screen and the speaker are arranged in the intelligent terminal device.

[0038] As a specific embodiment of the present application, the display screen is used to display the text information output by the large model Q&A server, and the speaker is used to play the voice generated by converting the replied text of the large model Q&A server.

[0039] Furthermore, the present application also includes a method for accelerating the output of the large model Q&A, including the following steps:

[0040] Step 1: The user inputs question information through the intelligent Q&A client module and sends the question information to the question receiving and classification preprocessing module in the large model Q&A server;

[0041] Step 2: The question receiving and classification preprocessing module classifies the question and preferentially triggers the query of hot issues, and the hot issue intelligent retrieval module retrieves the hot information;

[0042] Step 3: When a hot issue is hit, extract the answer corresponding to the question from the Q&A cache server;

[0043] Step 4: The hot issue intelligent retrieval module directly outputs the corresponding answer to the intelligent Q&A client module;

[0044] Step 5: The display screen displays the text information output by the large model Q&A;

[0045] Step 6: The speaker plays the voice generated by converting the reply text of the large model.

[0046] Step 7: When the hot issue is not hit, continue to retrieve in the knowledge base retrieval service module or retrieve through the online search service module as needed.

[0047] Step 8: The question receiving and classification preprocessing module inputs the user's original question and retrieval information into the inference service module.

[0048] Step 9: The inference service module converts the vector output by inference into a user-readable format and sends the information in the user-readable format to the output processing module.

[0049] Step 10: The output processing module submits the question and answer to the Q&A cache server for hot issue recognition. After recognition, it processes and outputs the Q&A to the intelligent Q&A client module.

[0050] Step 11: The display screen displays the text information output by the large model Q&A.

[0051] Step 12: The speaker plays the voice generated by converting the reply text of the large model.

[0052] In summary:

[0053] The present invention can cache hot issues and answers, and can quickly retrieve hot issues. Through question matching, the corresponding answers can be accurately extracted from the cache, greatly improving the user Q&A response speed and reducing the consumption of large model inference computing power.

[0054] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A system for accelerating the Q&A output of large models, comprising an intelligent terminal device and a large model Q&A server, characterized in that: The intelligent terminal device includes an intelligent Q&A client module, which is interconnected with a question receiving and classification preprocessing module through a wireless network. The question receiving and classification preprocessing module is arranged in the large model Q&A server. The question receiving and classification preprocessing module is interconnected with a hot issue intelligent retrieval module. The hot issue intelligent retrieval module is interconnected with a Q&A cache server. The hot issue intelligent retrieval module is also connected to an intelligent temperature client module. Both the hot issue intelligent retrieval module and the Q&A cache server are arranged in the large model Q&A server.

2. The system for accelerating the question-answering output of a large model according to claim 1, wherein The hot issue intelligent retrieval module is used for real-time retrieval of features in hot issues.

3. The system for accelerating the Q&A output of a large model according to claim 1, wherein: The Q&A cache server is used for storing the Q&A results of corresponding hot issues.

4. A system for accelerating the Q&A output of a large model according to claim 1, characterized in that, The question receiving and classification preprocessing module is also interconnected with a knowledge base retrieval service module and an online search service module. Both the knowledge base retrieval service module and the online search service module are arranged in the large model Q&A server.

5. The system for accelerating the answer output of the large model according to claim 4, wherein, The question receiving and classification preprocessing module is also connected to an inference service module. The inference service module is arranged in the large model Q&A server.

6. The system for accelerating the Q&A output of a large model according to claim 5, wherein The inference service module is also connected to an output processing module. The output inference module is respectively interconnected with the Q&A cache server and the intelligent Q&A client module. The output processing module is arranged in the large model Q&A server.

7. The system for accelerating the Q&A output of a large model according to claim 6, characterized in that, The intelligent Q&A client module is also connected to a display screen and a speaker. The display screen and the speaker are arranged in the intelligent terminal device.

8. The system for accelerating the Q&A output of a large model according to claim 2, wherein The display screen is used for displaying the text information output by the large model Q&A server. The speaker is used for playing the voice generated by converting the text of the large model's reply.

9. A method for accelerating the Q&A output of a large model according to any one of claims 1-8, characterized in that, It includes the following steps: S1. The user inputs question information through the intelligent Q&A client module and sends the question information to the question receiving and classification preprocessing module in the large model Q&A server. S2. The question receiving and classification preprocessing module classifies the question and preferentially triggers a hot issue query. The hot issue intelligent retrieval module retrieves the hot information. S3. When a hot issue is hit, extract the corresponding answer from the Q&A cache server. S4. The hot issue intelligent retrieval module directly outputs the corresponding answer to the intelligent Q&A client module. S5. The display screen displays the text information output by the large model Q&A. S6. The speaker plays the voice generated by converting the text of the large model's reply. S7. When the hot issue is not hit, continue to retrieve in the knowledge base retrieval service module or, as needed, through the online search service module. S8. The question receiving and classification preprocessing module inputs the user's original question and the retrieved information into the inference service module. S9. The inference service module converts the vector output by the inference into a user-readable format and sends the information in the user-readable format to the output processing module. S10. The output processing module submits the question and the answer to the Q&A cache server for hot issue identification. After identification, it processes and outputs the Q&A to the intelligent Q&A client module. S11. The display screen displays the text information output by the large model Q&A. S12. The speaker plays the voice generated by converting the reply text of the large model.