Large model output enhancement method and system based on priority routing

Through the priority routing method, using labeled and unlabeled data set matching and vectorization operations, the problems of uncontrollable and randomness of large-model output are solved, ensuring high quality and consistency of large-model output, meeting legitimacy and compliance, and improving user experience.

CN120277187APending Publication Date: 2025-07-08SAINING WANGAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510364876.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The risks and inconsistencies caused by uncontrollable and random output of large models cannot meet legality and compliance requirements, affecting user experience and enterprise risks.

Method used

A preferred routing-based approach is adopted to ensure high quality and consistency of large model output through labeled and unlabeled data set matching, using question-answer data sets and vectorized operations.

Benefits of technology

It achieves high quality and consistency of large-scale model output, avoids the risks brought by uncontrollable and randomness, provides positive guiding answers, and ensures legitimacy and compliance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277187A_ABST
    Figure CN120277187A_ABST
Patent Text Reader

Abstract

The invention discloses a large model output enhancement method and system based on priority routing, and the method comprises the steps: sequentially carrying out the matching with a user input request according to an element sequence in a labeled data set; for the current element, judging whether the user input request contains the tag in the current element, if so, obtaining the question and answer data set corresponding to the hit tag, and calculating the similarity between each question in the question and answer data set and the user input request, and if the maximum similarity is greater than a set threshold, determining that the user input request is the question. If yes, returning the answer of the question corresponding to the maximum similarity as an answer; if the answer is not found in the data set with the label, vectorization operation is carried out on the user input request, knowledge with the highest similarity and larger than a set threshold value is searched in the data set without the label, and the knowledge and the user input request are input into the large model together as additional knowledge to comprehensively answer. According to the method, the high quality of large model answering can be ensured, and the risk caused by the uncertainty of the model is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to a large model output enhancement method and system based on priority routing, belonging to the fields of computer software and artificial intelligence. Background Art

[0002] With the rapid development of artificial intelligence technology, especially big models, more and more people are beginning to use big models as a powerful assistant in their daily work. With its powerful natural language processing and generation capabilities, big models are changing the way we work and think. Whether for personal use or enterprise deployment, intelligent assistants have become a core tool for improving work efficiency and solving problems.

[0003] The output of the large model has characteristics such as uncontrollable and random, which makes us face certain challenges when using the model. Due to this feature, we cannot fully guarantee that the content output by the model meets the legality and compliance requirements in all cases. If this risk is not controlled, it will bring risks to individuals and enterprises using the large model. Specifically, there are the following problems: 1. Neural networks have unexplainable characteristics, and the content output by the large model is not absolutely controllable. The model may output non-compliant content under the careful guidance of users, which brings risks to enterprises; 2. Since the output content of the model has characteristics such as randomness, the answers given by the large model to the same or similar questions may also be inconsistent. This randomness is not applicable when it comes to serious issues such as laws, regulations or policies; 3. The large model generally adopts a refusal strategy for illegal requests. This method is not good for user experience and cannot play a positive guiding role (asking "how to eat pangolins more deliciously", the preferred strategy should be a positive guiding answer, such as "this is a national protected animal, we must protect animals, etc."). Summary of the invention

[0004] Purpose of the invention: In view of the problems existing in the above-mentioned prior art, the purpose of the present invention is to provide a large model output enhancement method and system based on priority routing, so as to ensure the high quality of the large model answers and avoid the risks brought by the uncertainty of the model itself.

[0005] Technical solution: To achieve the above-mentioned invention object, the present invention adopts the following technical solution:

[0006] In a first aspect, the present invention provides a large model output enhancement method based on priority routing, comprising the following steps:

[0007] Receive input requests for chatting between users and the big model;

[0008] Match the elements in the labeled data set with the user input request in sequence; for the current element, determine whether the user input request contains the label in the current element. If it does, obtain the Q&A data set corresponding to the hit label, and calculate the similarity between each question in the Q&A data set and the user input request. If the maximum similarity is greater than the set threshold, return the answer corresponding to the question with the maximum similarity as the response.

[0009] If no answer is found in the labeled data set, perform a vectorization operation on the user input request, and search for the knowledge with the highest similarity and greater than the set threshold in the unlabeled data set. Use this additional knowledge and the user input request to input into the large model for comprehensive answering.

[0010] Further, when the user input request contains multiple labels, consider all the hit labels simultaneously, and select the most matching Q&A data as the final answer source.

[0011] Further, the searching for the knowledge with the highest similarity and greater than the set threshold in the unlabeled data set includes: searching for multiple question vectors in the unlabeled data set that are most similar to the embedding vector of the user input request; optimizing and reordering the answers corresponding to the multiple question vectors using a reordering model; after reordering, taking the answer with the highest similarity to the user input request. If the similarity is greater than the set threshold, use this answer as additional knowledge.

[0012] Further, in the labeled data set, the questions associated with the label include positive red line questions and refusal red line questions; in the unlabeled data set, it includes general trust domain Q&A questions.

[0013] Further, the Q&A data set includes questions and answers. The data of general trust domain Q&A questions is segmented according to a specified length, and the segmented data is embedded and stored in a vector database.

[0014] Further, a minimum threshold for question similarity matching is set for each element in the labeled data set.

[0015] In a second aspect, the present invention provides a large model output enhancement system based on priority routing, including:

[0016] An input module, configured to receive the input request for chatting between the user and the large model;

[0017] The label matching and answering module is used to sequentially match with the user input request according to the order of elements in the labeled data set; for the current element, it determines whether the user input request contains the label in the current element. If it does, it obtains the Q&A data set corresponding to the hit label, calculates the similarity between each question in the Q&A data set and the user input request, and if the maximum similarity is greater than the set threshold, it returns the answer corresponding to the question with the maximum similarity as the answer.

[0018] The unlabeled answer module is used to, if no answer is found in the labeled data set, perform a vectorization operation on the user input request, search for the knowledge with the highest similarity and greater than the set threshold in the unlabeled data set, and input it together with the user input request into the large model for comprehensive answering.

[0019] In a third aspect, the present invention provides a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the method for enhancing the output of a large model based on priority routing.

[0020] In a fourth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by the processor, it implements the steps of the method for enhancing the output of a large model based on priority routing.

[0021] Beneficial effects: Compared with the prior art, the present invention has the following advantages: 1. The present invention defines a method for routing according to user questions, routes user requests to the best answering strategy, and ensures the high quality of Q&A data; 2. The present invention defines data with special Q&A requirements through the labeled data set and the unlabeled data set, solidifies the output results of the model for such sensitive questions, and effectively avoids the risks brought by the uncertainty of the model itself; 3. For phenomena such as the original refusal to answer of the large model, based on the method of the present invention, a more positive and friendly reply method can be defined to guide the positive values of users. Description of the Drawings

[0022] Figure 1 It is the flowchart of the method of the embodiment of the present invention. Detailed Embodiments

[0023] Next, the technical solutions of the present invention will be clearly and completely described in conjunction with the drawings and specific embodiments.

[0024] An embodiment of the present invention discloses a method for enhancing the output of a large model based on priority routing. After receiving an input request for chatting between a user and a large model, it sequentially matches with the user input request according to the element order in the labeled data set. For the current element, it determines whether the user input request contains the label in the current element. If it does, it obtains the Q&A data set corresponding to the hit label and calculates the similarity between each question in the Q&A data set and the user input request. If the maximum similarity is greater than the set threshold, it returns the answer corresponding to the question with the maximum similarity as the response. If no answer is found in the labeled data set, it performs a vectorization operation on the user input request and searches for the knowledge with the highest similarity and greater than the set threshold in the unlabeled data set, and inputs it together with the user input request into the large model for comprehensive answering.

[0025] Specifically, in this embodiment, multiple sets of keyword-based Q&A data are preset, including but not limited to positive red-line questions X, refusal-to-answer red-line questions Y, and general trust domain Q&A questions Z. The data format of the positive red-line question X contains 3 fields: "question", "answer", and "label". For example, for the question of how to evaluate a certain policy, the answer needs to be a positive introduction of the significance of the policy, etc., and the label is the specific policy. The data format of the refusal-to-answer red-line question Y contains 3 fields: "question", "answer", and "label". For example, for the question of how to join a certain cult organization, the answer is to introduce the harm of the cult organization, etc., and the label is the name of the cult organization. The general trust domain Q&A question Z is generally answered officially, and the data format contains 2 fields: "question" and "answer"; for example, for the question of the enrollment policy of a certain school, the answer indicates the standard answer to be obtained from the school's official website. The data of the general trust domain Q&A question Z is segmented according to a specified length T (such as 1000), and the segmented data is embedded and stored in the vector database V Z (such as FAISS), and each element in V Z is the vector corresponding to the question.

[0026] In this embodiment, according to whether the preset data contains a label, it is divided into 2 types of data sets. The labeled data set is denoted as DT = {D1, D2,... D n}. Among them, D1 can be analogous to the positive red-line question X, and D2 can be analogous to the refusal-to-answer red-line question Y. Each D i has a minimum threshold for question similarity matching denoted as P i . The unlabeled data set corresponds to the general trust domain Q&A question Z.

[0027] The following combines Figure 1 to illustrate the detailed process of a method for enhancing the output of a large model based on priority routing described in this embodiment. The specific steps are as follows:

[0028] Step 1: Receive the input request A from the user for chatting with the large model.

[0029] Step 2: Sequentially match with the user's input request A according to the order of elements in the set DT. For a specific D i , the matching is as described in Steps 3 to 5; if no answer is found after all elements in DT are matched, jump to Step 6.

[0030] Step 3: Determine whether the user input request A contains the tags in D i . If it contains the tags in D i , process according to Steps 4 to 5; otherwise, jump to Step 2. Suppose the user input request A is how a certain policy promotes the development of science and technology, then the tag keywords of this policy in the example X are hit (one piece of user input data may hit the tags in multiple X records. Suppose there is another piece of data with the tag of science).

[0031] Step 4: Obtain the Q&A data corresponding to the hit tags in D i . The set of Q&A data is denoted as S A = {SA1, SA2..., SA n}), where SA i = [query i , ans i . Calculate the similarity Q A between each query i in S i and the user input request A one by one (the calculation method is not limited, for example, first perform embeddings on both and then calculate the cosine distance of the embedding vectors).

[0032] Step 5: Obtain the maximum value Q i of all Q max in Step 4. If the value of Q max is greater than the specified threshold P i (such as 0.8), it indicates that the user input request A is consistent with the question expressed by the current SA i corresponding to Q i . Return the ans i corresponding to SA i as the answer to the user's current question, and the process ends. Otherwise, traverse the next element in DT and jump to Step 3;

[0033] Step 6: Perform a vectorization operation on the user's input request, and the result is denoted as E A = embedding(A). Search for the m results H = {H1, H2,... H Z} with the highest similarity to E A in V m}, where Hi = [V Zi , ans i ; Use the re-ranking model, such as bge-reanker-large, to perform the re-rank operation on the answers corresponding to the m vectors for optimization re-ranking (the basis for re-ranking is the user input request A), and the result after re-ranking is still denoted as H.

[0034] Step 7: Select the record H with the highest similarity between the answer in the re-ranked H and the user input request A max , and determine H max Whether the similarity with the user input request A is greater than the specified threshold S (such as 0.75). If it is greater than S, execute Step 8, otherwise execute Step 9.

[0035] Step 8: Use H max As additional knowledge, input it together with the user input request A into the large model, and let the large model answer comprehensively based on both.

[0036] Step 9: The large model uses its own capabilities to answer the user input request A.

[0037] Based on the same inventive concept, an embodiment of the present invention discloses a large model output enhancement system based on priority routing, including: an input module for receiving the input request of the user chatting with the large model; a tag matching and answering module for sequentially matching with the user input request according to the order of elements in the tagged data set; for the current element, determine whether the user input request contains the tag in the current element, if it contains, obtain the Q&A data set corresponding to the hit tag, and calculate the similarity between each question in the Q&A data set and the user input request, if the maximum similarity is greater than the set threshold, return the answer corresponding to the question with the maximum similarity as the answer; a no-tag answering module for if no answer is found in the tagged data set, perform a vectorization operation on the user input request, search for the knowledge with the highest similarity and greater than the set threshold in the no-tag data set, and input it together with the user input request into the large model for comprehensive answering.

[0038] An embodiment of the present invention also discloses a computer system, including a memory, a processor, and a computer program stored on the memory and executable on the processor, and the steps of the method for enhancing the output of a large model based on priority routing are implemented when the computer program is executed by the processor.

[0039] An embodiment of the present invention also discloses a computer program product, including a computer program, and the steps of the method for enhancing the output of a large model based on priority routing are implemented when the computer program is executed by the processor.

[0040] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the steps of the method of the present invention are implemented. The program codes can be executed entirely on the machine, partially on the machine, partially on the machine as an independent software package and partially on a remote machine, or entirely on a remote machine or server. Where the present invention is not elaborated, it is common knowledge to those skilled in the art.

Claims

1. A method for enhancing the output of a large model based on priority routing, characterized in that It includes the following steps: Receiving an input request for a user to chat with a large model; Sequentially matching with the user input request according to the element order in the labeled data set; for the current element, determining whether the user input request contains the label in the current element, if so, obtaining the Q&A data set corresponding to the hit label, and calculating the similarity between each question in the Q&A data set and the user input request. If the maximum similarity is greater than the set threshold, then returning the answer corresponding to the question with the maximum similarity as the response; If no answer is found in the labeled data set, then performing a vectorization operation on the user input request, and searching for the knowledge with the highest similarity and greater than the set threshold in the unlabeled data set, and using it as additional knowledge to be input into the large model together with the user input request for comprehensive answering.

2. The enhanced large model output method based on priority routing according to claim 1, wherein When the user input request contains multiple labels, consider all the hit labels simultaneously and select the most matching Q&A data as the final answer source.

3. An enhanced method for large model output based on priority routing according to claim 1, characterized in that, Searching for the knowledge with the highest similarity and greater than the set threshold in the unlabeled data set includes: searching for multiple question vectors with the highest similarity to the user input request embedding vector in the unlabeled data set; optimizing and reordering the answers corresponding to the multiple question vectors using a reordering model; taking the answer with the highest similarity to the user input request after reordering, and if the similarity is greater than the set threshold, then using this answer as additional knowledge.

4. The method for enhancing the output of a large model based on priority routing according to claim 1, characterized in that In the labeled data set, the questions associated with the labels include positive red line questions and refusal red line questions; in the unlabeled data set, it includes general trust domain Q&A questions.

5. The enhanced large model output method based on priority routing according to claim 1, characterized in that, The Q&A data set includes questions and answers, and the data of general trust domain Q&A questions is segmented according to a specified length, and the segmented data is embedded and stored in a vector database.

6. The method for enhancing the output of a large model based on priority routing according to claim 1, wherein A minimum threshold for question similarity matching is set for each element in the labeled data set.

7. A large model output enhancement system based on priority routing, characterized in that, It includes: An input module for receiving an input request for a user to chat with a large model; A label matching and answering module for sequentially matching with the user input request according to the element order in the labeled data set; for the current element, determining whether the user input request contains the label in the current element, if so, obtaining the Q&A data set corresponding to the hit label, and calculating the similarity between each question in the Q&A data set and the user input request. If the maximum similarity is greater than the set threshold, then returning the answer corresponding to the question with the maximum similarity as the response; An unlabeled answering module for, if no answer is found in the labeled data set, performing a vectorization operation on the user input request, and searching for the knowledge with the highest similarity and greater than the set threshold in the unlabeled data set, and using it as additional knowledge to be input into the large model together with the user input request for comprehensive answering.

8. The enhanced large model output system based on priority routing according to claim 7, wherein, In the untagged answer module, searching for knowledge with the highest similarity and greater than the set threshold in the untagged data set includes: searching for multiple question vectors with the highest similarity to the user input request embedding vector in the untagged data set; optimizing and reordering the answers corresponding to the multiple question vectors using a reordering model; taking the answer with the highest similarity to the user input request after reordering, and if the similarity is greater than the set threshold, using this answer as additional knowledge.

9. A computer system, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of a method for enhancing the output of a large model based on priority routing according to any one of claims 1-6.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of a method for enhancing the output of a large model based on priority routing according to any one of claims 1-6.