Model reasoning method and device suitable for question and answer scene, equipment and medium

By constructing new vocabulary and mapping relationships in large language models and using probability transfer models, the problems of slow inference speed and high computational cost in real-time dialogue scenarios are solved, and faster inference speed and higher accuracy in domain proprietary vocabulary generation are achieved.

CN120387522AActive Publication Date: 2025-07-29INSPUR SUZHOU INTELLIGENT TECH CO LTD

Patent Information

Application Number
CN202510873626.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-29
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

Existing large language models (LLMs) are slow inference and high computational costs in real-time conversation scenarios, and existing acceleration methods have problems with accuracy loss or additional computational costs.

Method used

By constructing a mapping relationship between a new vocabulary list and a pretrained model, combining the probability transfer model, quickly determine the output results of the pretrained model, and use the prior knowledge in the domain question and answer to reduce the pressure of computing resource and improve the accuracy of domain proprietary vocabulary generation.

Benefits of technology

It improves the speed of model inference and the accuracy of domain proprietary vocabulary generation, reduces the demand for computing resources, and is suitable for vertical domain question-and-answer scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120387522A_ABST
    Figure CN120387522A_ABST
Patent Text Reader

Abstract

The invention discloses a model reasoning method and device suitable for a question and answer scene, equipment and a medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: building a mapping relation between a new word list composed of generated target long words and an original word list through a preset word segmentation algorithm by using priori knowledge in target corpus data; when the pre-training model generates the specified content through reasoning, the final output result of the pre-training model is rapidly determined through the probability transfer model, the problems that an existing model acceleration method is low in precision and high in model training cost and calculation cost are solved, and only one probability transfer model is additionally added in the whole reasoning process, so that the calculation cost is reduced. The computing resource pressure is effectively reduced, and the generation speed of the whole system is improved. And in addition, priori knowledge in domain questions and answers is fully utilized, and the accuracy rate of domain proprietary vocabulary generation can also be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model reasoning method, apparatus, device, and medium suitable for question-answering scenarios. Background Art

[0002] With the rapid development of artificial intelligence (AI) technology, large language models (LLMs) have rapidly penetrated industrial and consumer scenarios from laboratory research. They have been deployed in a wide range of industries, including intelligent customer service, contract review, financial report generation, virtual assistants, medical consultations, and code generation. The speed and efficiency of model inference have become key factors influencing the implementation of this technology. In specific application scenarios, LLMs are generally deployed in cloud computing or edge devices (such as mobile phones). For scenarios with high real-time requirements, such as conversations, higher requirements are placed on inference latency and energy efficiency. However, the inference latency of models with hundreds of billions of parameters on conventional hardware can reach seconds to minutes. Therefore, how to quickly and efficiently perform model inference has become a pressing technical challenge. Summary of the Invention

[0003] The present application provides a model reasoning method, apparatus, device and medium suitable for question-answering scenarios, so as to at least solve the problems of low model acceleration accuracy, high model training cost and high computing cost in related technologies.

[0004] This application provides a model reasoning method suitable for question-answering scenarios, including: Obtain the generated content of the pre-trained model; the pre-trained model is the model applied to the target question-answering scenario; Matching the generated content with a pre-generated mapping table to obtain a matching result; wherein the mapping table is a table generated by constructing a mapping relationship between a new vocabulary and the original vocabulary of the pre-trained model; the new vocabulary is a vocabulary generated using target long words, and the target long words are words whose length meets the preset long word determination criteria after word segmentation processing of the target corpus data in the target question-answering scenario based on the original vocabulary using a preset word segmentation algorithm; When the matching result representation generates content that matches the mapping table, the pre-trained probability transfer model is called to determine the transfer probability of the target long word, and the output result of the pre-trained model is determined based on the transfer probability.

[0005] This application also provides a model reasoning device suitable for question-answering scenarios, including: A generated content acquisition module is used to obtain the generated content of the pre-trained model; the pre-trained model is the model applied to the target question-answering scenario; A generated content matching module is used to match the generated content with a pre-generated mapping table to obtain a matching result. The mapping table is a table generated by constructing a mapping relationship between a new word table and the original word table of a pre-trained model. The new word table is a word table generated using target long words. The target long words are words generated by performing word segmentation on target corpus data in a target Q&A scenario based on the original word table and using a preset word segmentation algorithm, and the word lengths of the generated words meet the preset long word determination conditions. An output result determination module is used to, when the matching result indicates that the generated content matches the mapping table, call a pre-trained probability transition model to determine the transition probability of the target long words, and determine the output result of the pre-trained model according to the transition probability.

[0006] This application also provides an electronic device, including: a memory for storing a computer program; a processor for implementing the steps of any of the above model inference methods applicable to the Q&A scenario when executing the computer program.

[0007] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above model inference methods applicable to the Q&A scenario are implemented.

[0008] Through this application, since most of the landing application projects of the pre-trained model are vertical domain Q&A, the domain data itself is a kind of prior knowledge. Therefore, the prior knowledge in domain Q&A can be fully utilized, and a mapping relationship is constructed between the new word table composed of the generated target long words and the original word table through a preset word segmentation algorithm. When the pre-trained model infers and generates the specified content, the final output result of the pre-trained model is quickly determined through the probability transition model. It can be seen that during the entire inference process, only an additional probability transition model is added. Since the probability transition model is a discriminative model and the parameter scale is much smaller than that of the language model, the computational resource pressure can be effectively reduced and the generation speed of the entire system can be improved. Moreover, making full use of the prior knowledge in domain Q&A can also improve the accuracy of generating domain-specific vocabulary. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the embodiments of this application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0010] Figure 1 It is a flowchart of a model inference method applicable to the Q&A scenario disclosed in this application; Figure 2 It is a schematic diagram of a mapping relationship disclosed in this application; Figure 3 Schematic diagram of the construction process of the mapping relationship between the new vocabulary and the original vocabulary disclosed in this application; Figure 4 Schematic diagram of the construction process of the training set of the probability transition judgment model disclosed in this application; Figure 5 Schematic diagram of the training of the probability transition model disclosed in this application; Figure 6 Flow chart of the accelerated inference of the LLM disclosed in this application; Figure 7 Schematic diagram of the structure of the model inference device applicable to the question-and-answer scenario disclosed in this application. Detailed implementation manners

[0011] Next, the technical solutions in the embodiments of this application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.

[0012] It should be noted that in the description of this application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusions, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent to such a process, method, article or device. The terms "first", "second", etc. in this application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0013] In the current context of the development of artificial intelligence technology, LLM models are widely used in various fields. However, with the improvement of the capabilities of LLM models, the number of their parameters has increased exponentially. Taking the GPT (Generative Pre-trained Transformer, a pre-trained language model) series as an example, the model scale has grown from 117 million parameters of GPT-1 (2018) to approximately 1.8 trillion parameters of GPT-4 (2023). Consequently, the computational complexity (FLOPs) and memory occupancy of LLM increase superlinearly, and the inference cost also rises steeply. For example, the single training cost of GPT-3 exceeds $4.6 million, and the single inference energy consumption of GPT-3 is equivalent to the charging amount of a mobile phone for 20 hours.

[0014] As LLM models become increasingly complex, the speed and efficiency of model reasoning have become key factors affecting the implementation of technology applications. Especially for scenarios with high real-time requirements such as conversations, how to quickly and efficiently perform model reasoning has become a technical challenge that needs to be solved urgently.

[0015] Currently, many inference acceleration technologies exist, including quantization, pruning, and distillation. Quantization reduces model weights and activations from FP32 / FP32 to INT8 / INT4, reducing memory usage and computational complexity. However, this method can result in a significant loss of accuracy. Pruning reduces the computational complexity of forward inference by trimming the model's inherent structure and removing redundant weights. This method also results in reduced accuracy and weakened model generalization. Knowledge distillation involves training a small model to fit the output distribution of a large model, and then using the small model instead of the large model for inference. However, this method is limited by the number of parameters in the small model, resulting in limited improvement.

[0016] Another method for accelerating inference is speculative sampling (Speculative Decoding). Its core idea is to use a small model to generate multiple tokens (basic units in text) and use a large model for verification and modification. This can reduce the number of calls to the large model and thus increase generation speed. In this method, the large model's computational workload is concentrated on long sequence verification rather than token-by-token generation. The final output is completely determined by the probability distribution of the large model, consistent with the results generated directly using the large model, without sacrificing generation quality. However, speculative sampling schemes have the following disadvantages: (1) If the accuracy of the candidate generated by the small model is low, the rejection rate of the large model will increase, and it will be necessary to frequently fall back to the large model generation, resulting in a decrease in the acceleration effect or even negative optimization. Therefore, in order to ensure that the distribution probability of the two models is close, the current optimization solution is to train the two models on the same batch of data, but this will lead to additional training costs. (2) Since the small model is also a generative model, too few parameters will directly affect the generation accuracy. Therefore, the parameter scale of the small model is only smaller than that of the large model. In fact, it still requires more than 6B parameters. The small model inference itself introduces additional computing costs, which may offset some of the acceleration benefits. In addition, the small model and the large model need to be loaded simultaneously during inference, which increases the memory usage and requires more computing resources.

[0017] To this end, this application provides a model reasoning solution suitable for question-answering scenarios, which can effectively reduce computing resource pressure, increase the generation speed of the entire system, and improve the accuracy of domain-specific vocabulary generation. To enable those skilled in the art to better understand this application solution, the application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0018] In combination with a specific application environment architecture or a specific hardware architecture on which the execution of the model inference method applicable to the question-and-answer scenario depends, the specific application environment architecture or the specific hardware architecture is described herein.

[0019] Embodiments of the present application provide a model inference method applicable to the question-and-answer scenario. Referring to the execution process of the model inference method applicable to the question-and-answer scenario, Figure 1 as shown, the method is described in detail. The method includes: Step S11: Obtain the generated content of the pre-trained model; the pre-trained model is a model applied to the target question-and-answer scenario.

[0020] In the embodiments of the present application, the pre-trained model is illustrated by taking a large language model (LLM) as an example. In real LLM implementation projects, most are vertical domain question-and-answer. Therefore, the target question-and-answer scenario is a question-and-answer scenario for a vertical domain, that is, a question-and-answer scenario for a specific industry, professional field, or sub-scenario, usually focusing on a specific knowledge category.

[0021] Step S12: Match the generated content with a pre-generated mapping table to obtain a matching result; wherein, the mapping table is a table generated by constructing a mapping relationship between a new word table and the original word table of the pre-trained model; the new word table is a word table generated using target long words, and the target long words are words whose word lengths meet the preset long word determination condition after the target corpus data in the target question-and-answer scenario is segmented using a preset word segmentation algorithm based on the original word table.

[0022] In the embodiments of the present application, a mapping table is pre-constructed, and the mapping table is a table generated by constructing a mapping relationship between a new word table and the original word table of the pre-trained model.

[0023] The original word table refers to the vocabulary set used to represent the smallest units in the corpus in natural language processing (NLP) tasks. It usually contains all possible words or characters, which are the basis for constructing word vectors, word tables, and subsequent model training. The original word table of the LLM comes from the text corpus used in its pre-training stage, and these corpora usually come from large-scale unlabeled data such as the Internet, books, and articles.

[0024] The new vocabulary is a vocabulary generated using target long words, wherein the target long words are words whose lengths meet the preset long word determination conditions after the target corpus data in the target question-answering scenario is segmented based on the original vocabulary and using a preset word segmentation algorithm. Since the target question-answering scenario is a question-answering scenario for a vertical field, in the embodiment of the present application, the target corpus data is usually a vertical field corpus, that is, in the current question-answering scenario, a set of text data for a specific industry, professional field or segmented scenario is a priori knowledge. This type of corpus usually has a high degree of domain expertise, semantic concentration and application pertinence. By using a word segmentation algorithm to generate target long words based on the original vocabulary according to domain data, the construction of a new vocabulary is achieved.

[0025] Currently, mainstream word segmentation algorithms include BPE (Byte Pair Encoding) and BBPE (Byte-level Byte Pair Encoding). BPE is essentially a data compression method. It first requires a predefined character set and a set vocabulary size. New tokens are then generated by continuously merging the most frequent consecutive token pairs (the basic unit of text) in the corpus until the set vocabulary size is reached. BBPE, on the other hand, constructs sentences as sequences of UTF-8-encoded bytes (each byte is 8 bits) rather than character sequences. It uses a 256-byte set as its vocabulary and employs the same strategy as BPE, generating new tokens by continuously merging the most frequent consecutive token pairs in the corpus until the set vocabulary size is reached.

[0026] In the embodiment of the present application, BPE is used as an example to generate target long words. The implementation of other word segmentation algorithms is similar to this and will not be described in detail later.

[0027] In the embodiment of the present application, a newly constructed long word vocabulary is used to construct a mapping relationship between the tokens of the original vocabulary and the long words generated in the new vocabulary. Specifically, the words in the original vocabulary are used as keys, and the list of target long words prefixed with the keys in the new vocabulary is used as values to construct a mapping relationship between the new vocabulary and the original vocabulary, so as to obtain a mapping table from the words in the original vocabulary to the target long words in the new vocabulary. Its format is: key: list <value>, where the key is the token in the original vocabulary, and the value is the list of generated long words in the new vocabulary, all of which are prefixed with the key.

[0028] It should be noted that the token (generated long word) in the new vocabulary must start with the token in the original vocabulary, and the long words in the new vocabulary must not exist in the original vocabulary. For example, if a certain token in the new vocabulary is "shadow bank", then there must be any one of "shadow", "shadow bank" in the original vocabulary, otherwise the long word will be discarded. As Figure 2 As shown is a simple mapping relationship, where "cross" is the token in the original vocabulary, and "cross-currency basis" etc. are the words in the new vocabulary.

[0029] Step S13: When the generated content of the matching result representation matches the mapping table, call the pre-trained probability transition model to determine the transition probability of the target long word, and determine the output result of the pre-trained model according to the transition probability.

[0030] In the embodiment of the present application, the fact that the generated content of the matching result representation matches the mapping table means that the generated content of the LLM model matches the key in the mapping table. Further, it is necessary to judge through the probability transition model whether a certain value can be matched. That is, call the pre-trained probability transition model to determine the transition probability of the target long word. Then compare the transition probability with the classification decision threshold of the probability transition model. This classification decision threshold is used for the classification judgment of the model. If it is greater than this threshold, it means that the possibility of belonging to the positive sample is high enough, so it can be judged as a positive sample.

[0031] Further, if the transition probability is greater than the classification decision threshold, it means that the model matches successfully, and it is considered that the model outputs the entire target long word at one time. Then directly generate the target long word and output the target long word as the output result of the pre-trained model, and then continue the subsequent reasoning; if the transition probability is not greater than the classification decision threshold, continue to reason through the pre-trained model, and when the generated content of the pre-trained model matches the key in the mapping table, jump to the step of calling the pre-trained probability transition model to determine the transition probability of the target long word.

[0032] In a feasible implementation, multi-modal data can be utilized to enhance the depth of understanding of domain knowledge. For example, modal information other than text (such as charts, formulas, and symbols in a vertical domain) can be incorporated into the probability transfer model. For instance, in the medical field, examination images and test report data in a case can be combined as auxiliary features to strengthen the judgment basis of the model for long word generation decisions, avoiding misjudgments caused by relying solely on text context. In this way, it can be applied to complex scenarios that require comprehensive information judgment and fill the blind spots of single text. In another feasible implementation, an adversarial training mechanism can also be adopted, introducing a generative adversarial network (GAN) or adversarial samples to enhance the robustness of the probability transfer model. For example, generating adversarial samples (such as deliberately confused domain terms) forces the model to learn to distinguish real long words from interference items, improving the judgment accuracy under noisy inputs. In this way, the adaptability of the model to complex inputs can be enhanced, and the risk of inference errors caused by data bias or malicious inputs can be reduced. In addition, a human-machine collaborative interface can be designed to allow domain experts to manually annotate and adjust the decision logic of the long word list or probability transfer model. By combining expert knowledge to make up for the limitations of data or algorithms, the interpretability and customization capabilities of the system can be improved.

[0033] Through this application, since most of the landing application projects of the pre-trained model are vertical domain Q&A, the domain data itself is a kind of prior knowledge. Therefore, the prior knowledge in domain Q&A can be fully utilized. By presetting a word segmentation algorithm, a mapping relationship is constructed between the new word list composed of the generated target long words and the original word list. When the pre-trained model infers and generates the specified content, the final output result of the pre-trained model is quickly determined through the probability transfer model. It can be seen that during the entire inference process, only a probability transfer model is additionally added. Since the probability transfer model is a discriminative model and the parameter scale is much smaller than that of the language model, the computational resource pressure can be effectively reduced, and the generation speed of the entire system can be improved. Moreover, by fully utilizing the prior knowledge in domain Q&A, the accuracy of generating domain-specific vocabulary can also be improved.

[0034] Based on the above embodiments, this embodiment will specifically elaborate on the generation process of the target long words in the above embodiments. For the process of constructing the token->target long word mapping of the domain data, combined with Figure 3 as shown, the following steps are included: Based on the original word list of the pre-trained model and using the preset word segmentation algorithm, the target corpus data is segmented to obtain different segmentation results; On the target corpus data, the frequency of each unit pair composed of adjacent segmentation results is counted, and the unit pair with the highest frequency is iteratively merged to generate new long words; The new long words are screened according to the preset screening rules, and after determining that the number of new long words reaches the preset new word list size based on the number of new long words, the target long words are determined.

[0035] When training an LLM, the text is converted into tokens by a tokenizer and then encoded into vectors for subsequent calculations. The tokenizer mainly consists of two parts: a vocabulary (the mapping between token IDs and tokens) and a tokenization algorithm. In the embodiments of this application, the preset tokenization algorithm is illustrated by taking the BPE algorithm as an example.

[0036] First, input the target corpus data, the vocabulary size, and the original vocabulary. Among them, the vocabulary size is the preset initial vocabulary size. Then, based on the original vocabulary, the target corpus data is split into the smallest units to obtain different tokenization results. Here, the "smallest unit" refers to splitting the target corpus data into characters / sub-words that cannot be further split according to the tokens in the original vocabulary.

[0037] Furthermore, according to the implementation logic of the BPE algorithm, the frequency of adjacent unit pairs within words is counted on the corpus, and the token pair with the highest frequency is selected for merging to generate a new long word. Taking the medical field as an example, assume that a medical text segment of the target corpus data is: Immunotherapy can improve diabetes. The original vocabulary is the basic vocabulary in the medical field, such as {"cell", "immune", "therapy", "diabetes", "disease"}. Then the corpus tokenization result is {"immune", "cell", "therapy", "diabetes", "disease", "condition"}, and the adjacent unit pairs include {"immune", "cell"}, {"cell", "therapy"}, {"therapy", "can"}, {"can", "improve"}, {"improve", "diabetes"}, {"diabetes", "disease"}, {"disease", "condition"}. If the frequency of {"immune", "cell"} is the highest, then {"immune", "cell"} is merged to generate "immunocyte", and the subsequent tokenization result becomes {"immunocyte", "therapy", "diabetes", "disease", "condition"}.

[0038] After merging to generate new long words, the merged tokens are screened, and the screened tokens are used as the tokens in the new vocabulary. The screening rules are as follows: Judge in sequence whether the new long word appears in the original vocabulary to filter out the new long words that are not in the original vocabulary; that is, the newly generated tokens must not be in the original vocabulary; Judge in sequence whether the unit pair contains punctuation marks to filter out the new long words that do not contain punctuation marks in the unit pair; that is, the newly generated tokens cannot cross sentences and cannot contain punctuation marks such as (.,?!;).

[0039] Repeat the above process of iterative merging and screening until the size of the current vocabulary reaches the pre-set vocabulary size. It should be noted that the new vocabulary is merged based on the original vocabulary, and based on the new vocabulary, the mapping relationship between the original vocabulary and the new vocabulary can be established.

[0040] In the embodiments of the present application, the number of new vocabularies is now determined in advance and is strongly correlated with the data in the vertical domain. After the data scale and data quality in the given vertical domain are determined, if the number of new vocabularies is set too large, the number of generated words will be relatively large, and the probability of matching the generated words during LLM inference will increase. However, the length of the generated words will be relatively short, and the acceleration effect during model inference may not be obvious. For example, when constructing the vocabulary, some common two-word combinations or three-word combinations are likely to enter the vocabulary because their frequencies are not low. In this way, even if there are long words, they may be composed of several short parts spliced together. High-frequency short words and some common combinations will be preferentially selected, making it difficult for long words to become longer. However, if the vocabulary is set too small, the range of choices for the model when generating words will be limited. In order to express various semantics, it can only form long words by combining more small words, which may lead to overly long words, reducing the probability of matching long words during LLM inference and having a negative impact on the inference speed.

[0041] Therefore, based on the above considerations, for the setting of the size of the new vocabulary, it is necessary to comprehensively consider the inference speed and the matching success rate. Therefore, the recommended reference value is: , that is, 1 / 2 of the number of new vocabularies after the second iterative screening.

[0042] Specifically, perform two iterative operations of the preset word segmentation algorithm, and record the number of words included in the new vocabulary when generating the new vocabulary after the second iteration; set half of the number of words as the size of the new vocabulary, and continue to perform the iterative operation of the preset word segmentation algorithm. When it is determined that the generated new vocabulary reaches the size of the new vocabulary based on the number of new long words, determine the target long words.

[0043] In the embodiments of the present application, after setting the initial value size of the new vocabulary, perform another screening after the first screening. The number of words remaining after the second screening is the number of new vocabularies after the second iterative screening. The size of the vocabulary at this time is the recommended final vocabulary size.

[0044] It can be seen that in this embodiment, the prior knowledge in vertical domain Q&A can be fully utilized. Through word segmentation algorithms such as BPE, a mapping relationship (the mapping between tokens and long words, phrases, and proper nouns) can be established between the domain data and the original vocabulary of the LLM. In this way, when the LLM infers and generates a specified key subsequently, it can quickly determine whether the match is successful through the probability transfer judgment model, and multiple tokens can be generated at one time, improving the generation speed of the entire system.

[0045] Based on the above embodiments, in a feasible implementation manner, the process of training the probability transition model may include the following steps: Traverse the target corpus data based on the mapping table to extract samples, so as to obtain the context where the key and value appear in the target corpus data; Replace the position where the key appears in the context with the target long word in the value to generate positive samples, and keep the position where the key appears in the context unchanged to generate negative samples; Count the number of positive samples and negative samples generated under each mapping relationship in the mapping table, and determine the proportionality coefficient of the negative samples based on the number of positive samples and negative samples; If the proportionality coefficient is lower than the preset threshold, extract data containing the same key from the general domain corpus to supplement the negative samples until the proportionality coefficient is not lower than the preset threshold, and then determine the current negative samples; the general domain corpus is a collection of text data for different domains in different Q&A scenarios; Use the current negative samples and positive samples to construct a binary classification data set, and train the probability transition model based on the binary classification data set.

[0046] In the embodiments of the present application, the training data is first constructed, and the construction process of the entire training set is as Figure 4 shown. Based on the long word table mapping, samples are extracted from the vertical corpus; that is, the context of the mapping relationship key and value is searched. Taking {"shadow": ["shadow bank"]} as an example, from the target corpus data, find all the contexts containing "shadow" and "shadow bank", and the specific examples are as follows: The darker area formed because an object blocks the propagation of light and cannot pass through an opaque object is what we usually call a shadow; All in all, all deposit and loan-like businesses outside the bank supervision system can be called shadow banks.

[0047] By replacing the key with the value, a binary classification data set can be constructed. The example is shown in Table 1: Table 1 Binary classification data set

[0048] It is understandable that when constructing a binary classification dataset, the negative samples are the original samples without replacing the key (i.e., keeping the key and not replacing it with the value). For highly domain-specific data, the key is usually a domain-specific term, and the probability of the occurrence of such specific terms may be too high (e.g., "cell"), resulting in insufficiently rich negative samples. For example, in the corpus, there are only descriptions of shadow banks, without related descriptions of "shadow" or "image". Or if the key is "cell" and the value is "immune cell", in medical data, the context of "cell" may frequently appear with domain terms such as "immune" and "treatment", resulting in a very low natural occurrence frequency of negative samples (i.e., the case where "cell" is not replaced) in the domain data (because domain expressions tend to use the complete term "immune cell"). Therefore, to improve the richness of negative samples, additional negative samples need to be sampled from general data. The vocabulary distribution of general domain corpora is more balanced, covering multiple domains and supporting general scenarios (such as search engines and social media).

[0049] It should be noted that for any mapping relationship, the ratio of positive and negative samples should be preserved and used as the threshold for the subsequent discriminant model to make a judgment. That is to say, if the ratio of the number of negative samples to the number of positive samples is 4:1, then the probability y output by the model should use 4 / (4 + 1) as the threshold, that is: ; Example of the threshold mapping storage format: {shadow bank: 0.8}. By traversing all mapping relationships, a binary classification dataset can be constructed; in addition, the length of all samples should be within 510 tokens.

[0050] Furthermore, considering both accuracy and computational efficiency, in this embodiment, the Bert-base structure is selected for fine-tuning the probability transfer model. However, this is not the only implementation method of the present invention, and the implementation of other models should also be within the protection scope of the present invention. The training process is as shown in the appendix Figure 5 As shown, a [CLS] prefix is added to the front of the sample, and a [SEP] suffix is added to the back. [CLS] is the abbreviation of "classification", which usually represents the beginning of a sentence or a document. In BERT, [CLS] corresponds to the first word in the input text; [SEP] is the abbreviation of "separator", which usually represents the end of a sentence or a document. In BERT, [SEP] corresponds to the word vector of the last word in the input text, and its function is to separate different sentences. For example, when processing sentence pairs in BERT, a [SEP] is usually inserted between two sentences to indicate their boundary point. The final output takes the output corresponding to [CLS], and after passing through the pooler layer and softmax, it is the final probability. Using the threshold saved when constructing the training set, the category can be judged.

[0051] Specifically, calling the probability transfer model to determine the transfer probability of the target long word includes: determining the input sequence of the probability transfer model according to the target long word, and according to the token symbols in the input sequence, passing the input sequence through the pooling layer and the activation function in turn to determine the transfer probability of the target long word.

[0052] Further, the entire process of LLM accelerated generation is as shown in the appendix Figure 6 As shown, when performing vertical domain Q&A, when a certain token generated by the LLM matches the key of the mapping table, the above text and the corresponding words are thrown to the probability judgment model for judgment. If the probability is greater than the specified threshold, it is considered that the model outputs the entire target long word at one time, otherwise it is thrown to the LLM for continued reasoning.

[0053] It can be seen that in this embodiment, the training data of the probability transfer judgment model is fully constructed based on the domain data, so that the model can quickly judge whether the next token can be the target long word based on the context of the above text. And because the texts of the domain data and the general data are significantly different, the accuracy of the judgment model is very high, which greatly improves the accuracy of the LLM in domain Q&A. During the entire reasoning process, only an additional probability transfer model is added. Since the probability transfer model is a discriminative model and the parameter scale is much smaller than that of the language model, it can effectively reduce the pressure of computing resources. It can be used in scenarios such as question recommendation, customer service Q&A, article writing, and information extraction.

[0054] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation method.

[0055] The embodiment of the present application also provides a model inference device applicable to the Q&A scenario. Refer to Figure 7 As shown, the device includes: A generated content acquisition module 11, configured to acquire the generated content of the pre-trained model; the pre-trained model is a model applied to the target Q&A scenario; A generated content matching module 12, configured to match the generated content with a pre-generated mapping table to obtain a matching result; wherein, the mapping table is a table generated by constructing a mapping relationship between a new word table and the original word table of the pre-trained model; the new word table is a word table generated by using the target long word, and the target long word is a word generated by performing word segmentation processing on the target corpus data in the target Q&A scenario based on the original word table and using a preset word segmentation algorithm, and the word length meets the preset long word determination condition; An output result determination module 13, configured to, when the generated content in the matching result representation matches the mapping table, call a pre-trained probability transition model to determine the transition probability of the target long word, and determine the output result of the pre-trained model according to the transition probability.

[0056] For the description of the features in the corresponding embodiments of the model inference device applicable to the question-and-answer scenario, reference can be made to the relevant descriptions in the corresponding embodiments of the model inference method applicable to the question-and-answer scenario, which will not be elaborated here one by one.

[0057] In a feasible implementation manner, the model inference device applicable to the question-and-answer scenario further includes: A target long word generation module, configured to: Perform word segmentation on the target corpus data based on the original vocabulary of the pre-trained model and using a preset word segmentation algorithm to obtain different word segmentation results; Count the frequency of each unit pair composed of adjacent word segmentation results in the target corpus data, and iteratively merge the unit pairs with the highest frequency to generate new long words; Screen the new long words according to a preset screening rule, and determine the target long word after determining that the generated new vocabulary reaches a preset new vocabulary size based on the number of new long words.

[0058] In a feasible implementation manner, the target long word generation module further includes: A rule screening unit, specifically configured to: sequentially determine whether the new long word appears in the original vocabulary to screen out the new long words that do not appear in the original vocabulary; sequentially determine whether the unit pair contains a sentence-breaking symbol to screen out the new long words that do not contain a sentence-breaking symbol in the unit pair.

[0059] A long word generation unit, specifically configured to: perform two iterative operations of the preset word segmentation algorithm and record the number of words included in the new vocabulary when the new vocabulary is generated after the second iteration; set half of the number of words as the new vocabulary size, and continue to perform the iterative operation of the preset word segmentation algorithm. When it is determined that the generated new vocabulary reaches the new vocabulary size based on the number of new long words, determine the target long word.

[0060] In a feasible implementation manner, the model inference device applicable to the question-and-answer scenario further includes: A mapping relationship construction module, configured to: use the words in the original vocabulary as keys, and use the list of target long words with the key as the prefix in the new vocabulary as values to construct the mapping relationship between the new vocabulary and the original vocabulary, so as to obtain a mapping table that maps the words in the original vocabulary to the target long words in the new vocabulary.

[0061] A model training module, configured to: Traverse the target corpus data based on the mapping table for sample extraction to obtain the context of the key and value that appear in the target corpus data; Replace the positions of the keys that appear in the above text with the target long words in the values to generate positive samples, and keep the positions of the keys that appear in the above text unchanged to generate negative samples; Count the number of positive samples and negative samples generated under each mapping relationship in the mapping table, and determine the proportionality coefficient of the negative samples based on the number of positive samples and negative samples; If the proportionality coefficient is lower than the preset threshold, extract the data containing the same keys from the general domain corpus to supplement the negative samples until the proportionality coefficient is not lower than the preset threshold, and then determine the current negative samples; the general domain corpus is a set of text data for different domains in different question-and-answer scenarios; Construct a binary classification dataset using the current negative samples and positive samples, and train a probability transition model based on the binary classification dataset; Correspondingly, the output result determination module 13 is specifically configured to: Determine the input sequence of the probability transition model according to the target long word, and determine the transition probability of the target long word after sequentially passing the input sequence through the pooling layer and the activation function calculation according to the token symbols in the input sequence.

[0062] In a feasible implementation manner, the output result determination module 13 is specifically configured to: When the generated content in the matching result representation matches the key in the mapping table, call the pre-trained probability transition model to determine the transition probability of the target long word; Determine the classification decision threshold of the probability transition model, and compare the transition probability with the classification decision threshold; If the transition probability is greater than the classification decision threshold, output the target long word as the output result of the pre-training model; If the transition probability is not greater than the classification decision threshold, continue to infer through the pre-training model, and when the generated content of the pre-training model matches the key in the mapping table, jump to the step of calling the pre-trained probability transition model to determine the transition probability of the target long word.

[0063] An embodiment of the present application further provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above embodiments of the model inference method applicable to the question-and-answer scenario.

[0064] An embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. The computer program is configured to execute the steps in any one of the above embodiments of the model inference method applicable to the question-and-answer scenario when running.

[0065] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media that can store computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), mobile hard disks, magnetic disks, or optical discs.

[0066] The embodiments of the present application also provide a computer program product. The above computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the model inference method applicable to the question-and-answer scenario.

[0067] The embodiments of the present application also provide another computer program product, including a non-volatile computer-readable storage medium. The non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps in any of the above embodiments of the model inference method applicable to the question-and-answer scenario.

[0068] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0069] The above has introduced in detail a model inference method, device, equipment, and medium applicable to the question-and-answer scenario provided by the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.< / value>

Claims

1. A model inference method applicable to the Q&A scenario, characterized in that, Including: Obtain the generated content of the pre-trained model; The pre-trained model is a model applied to the target question-answering scenario; Match the generated content with a pre-generated mapping table to obtain a matching result; wherein, the mapping table is a table generated by constructing a mapping relationship between a new word list and the original word list of the pre-trained model; the new word list is a word list generated using a target long word, and the target long word is a word generated by performing word segmentation on the target corpus data in the target question-answering scenario based on the original word list and using a preset word segmentation algorithm, and the word length of the generated word meets the preset long word determination condition; When the matching result indicates that the generated content matches the mapping table, call a pre-trained probability transition model to determine the transition probability of the target long word, and determine the output result of the pre-trained model according to the transition probability.

2. The model inference method applicable to the Q&A scenario according to claim 1, wherein The process of generating the target long word includes: Perform word segmentation on the target corpus data based on the original word list of the pre-trained model and using a preset word segmentation algorithm to obtain different word segmentation results; Statistically calculate the frequency of occurrence of each unit pair composed of adjacent word segmentation results on the target corpus data, and iteratively merge the unit pairs with the highest frequency of occurrence to generate new long words; Screen the new long words according to a preset screening rule, and determine the target long word when it is determined that the generated new word list reaches a preset new word list size based on the number of the new long words.

3. The model inference method applicable to the Q&A scenario according to claim 2, wherein The screening of the new long words according to the preset screening rule includes: Sequentially determine whether the new long words appear in the original word list to screen out the new long words that do not appear in the original word list; Sequentially determine whether the unit pairs contain sentence-breaking symbols to screen out the new long words that do not contain the sentence-breaking symbols in the unit pairs.

4. The model inference method applicable to the Q&A scenario according to claim 2, wherein The determination of the target long word when it is determined that the generated new word list reaches a preset new word list size based on the number of the new long words includes: Execute two iterative operations of the preset word segmentation algorithm, and record the number of words included in the new word list when the new word list is generated after the second iteration; Set half of the number of words as the new word list size, and continue to execute the iterative operation of the preset word segmentation algorithm. When it is determined that the generated new word list reaches the new word list size based on the number of the new long words, determine the target long word.

5. The model inference method applicable to the Q&A scenario according to claim 1, wherein The process of constructing the mapping table includes: Use the words in the original word list as keys, and use the list of target long words prefixed with the key in the new word list as values to construct the mapping relationship between the new word list and the original word list, so as to obtain a mapping table that maps the words in the original word list to the target long words in the new word list.

6. The model inference method applicable to the question-and-answer scenario according to claim 5, characterized in that, The process of training the probability transition model includes: Traverse the target corpus data based on the mapping table for sample extraction to obtain the context in which the key and the value appear in the target corpus data; Replace the position where the key appears in the context with the target long word in the value to generate a positive sample, and keep the position where the key appears in the context unchanged to generate a negative sample; Count the number of positive samples and negative samples generated under each mapping relationship in the mapping table, and determine the proportionality coefficient of the negative samples based on the number of positive samples and the number of negative samples; If the proportionality coefficient is lower than a preset threshold, extract data containing the same key from the general domain corpus to supplement the negative samples until the proportionality coefficient is not lower than the preset threshold, and then determine the current negative samples; the general domain corpus is a collection of text data for different domains in different Q&A scenarios; Construct a binary classification dataset using the current negative samples and the positive samples, and train the probability transition model based on the binary classification dataset; Correspondingly, call the probability transition model to determine the transition probability of the target long word, including: Determine the input sequence of the probability transition model according to the target long word, and based on the token symbols in the input sequence, determine the transition probability of the target long word after passing the input sequence through the pooling layer and activation function calculations in sequence.

7. The model inference method applicable to the Q&A scenario according to any one of claims 1 to 6, characterized in that, When the matching result indicates that the generated content matches the mapping table, call the pre-trained probability transition model to determine the transition probability of the target long word, and determine the output result of the pre-trained model according to the transition probability, including: When the matching result indicates that the generated content matches the key in the mapping table, call the pre-trained probability transition model to determine the transition probability of the target long word; Determine the classification decision threshold of the probability transition model, and compare the transition probability with the classification decision threshold; If the transition probability is greater than the classification decision threshold, output the target long word as the output result of the pre-trained model; If the transition probability is not greater than the classification decision threshold, continue to reason through the pre-trained model, and when the generated content of the pre-trained model matches the key in the mapping table, jump to the step of calling the pre-trained probability transition model to determine the transition probability of the target long word.

8. A model inference device applicable to a question-and-answer scenario, characterized in that, Including: A generated content acquisition module for acquiring the generated content of the pre-trained model; The pre-trained model is a model applied to the target Q&A scenario; A generated content matching module for matching the generated content with a pre-generated mapping table to obtain a matching result; wherein, the mapping table is a table generated by constructing a mapping relationship between a new word table and the original word table of the pre-trained model; the new word table is a word table generated using the target long word, and the target long word is a word whose word length meets the preset long word determination condition after performing word segmentation processing on the target corpus data in the target Q&A scenario based on the original word table and using a preset word segmentation algorithm; An output result determination module for calling the pre-trained probability transition model to determine the transition probability of the target long word and determining the output result of the pre-trained model according to the transition probability when the matching result indicates that the generated content matches the mapping table.

9. An electronic device, characterized in that, Including: A memory for storing computer programs; A processor, configured to implement the steps of the model inference method applicable to the question-and-answer scenario according to any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the model inference method applicable to the question-and-answer scenario according to any one of claims 1 to 7 when executed by a processor.

Citation Information

Patent Citations

  • Knowledge graph multi-hop question and answer method based on reinforcement learning path reasoning

    CN115640410A

  • Word and sentence generation method and related equipment

    CN116306612A

  • Word list conversion method and device, equipment and storage medium

    CN117350279A

  • Word list construction method and device, storage medium and electronic equipment

    CN118133818A

  • Question and answer task processing model training method and device, equipment and storage medium

    CN119493849A

Cited By

  • Large language model multi-user high-concurrency high-throughput reasoning method and system

    CN120598065A

  • Methods and systems for multi-user, high-concurrency, and high-throughput inference of large language models

    CN120598065B