Problem processing method and device based on artificial intelligence, computer equipment and medium

By receiving user input in a large language model, retrieval and splicing processing, generating target prompts, and adjusting the probability of token generation, the problem of insufficient controllability of generating specific vocabulary in a specific scenario is solved, and the accuracy and security of the generated reply content is achieved.

CN120011540APending Publication Date: 2025-05-16SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510086705.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The existing large language model has insufficient controllability to generate specific vocabulary in specific scenarios, making it difficult to ensure the accuracy and security of the generated text.

Method used

By receiving questions entered by users, searching and obtaining reference text, and splicing the questions, reference text and preset prompt templates to generate target prompts. Then the target prompt is tokenized, the big model is called and the bias vector is created, and the token generation probability is adjusted to ensure that the generated reply content is accurate and controllable.

Benefits of technology

It effectively improves the controllability, accuracy and security of the large language model in the text generation process, ensuring that the generated reply content meets the needs of specific scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011540A_ABST
    Figure CN120011540A_ABST
Patent Text Reader

Abstract

The invention provides a problem processing method and device based on artificial intelligence, computer equipment and a storage medium. The method comprises the steps of receiving a question input by a user, and performing retrieval processing on the question to obtain a corresponding reference text; splicing the question, the reference text and a preset prompt template to obtain a corresponding target prompt; performing token processing on the target prompt to obtain a corresponding target token set; calling a preset large model, and creating a bias vector corresponding to a model vocabulary of the large model; for each token in the target token set, setting a positive value bias at a corresponding index position in a bias vector to obtain a corresponding target bias vector; reasoning the target prompt and the target offset vector based on a large model to obtain a corresponding reply; and performing question feedback processing corresponding to the user based on the reply. According to the method and the device, the accuracy and controllability of the generated reply content can be effectively ensured, and the controllability, accuracy and safety of the reply are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to an artificial intelligence-based problem processing method, device, computer equipment and storage medium. Background Art

[0002] With the rapid development of information technology today, large language models (LLMs), as an important technology in the field of artificial intelligence, have shown great potential in various application scenarios with their powerful text understanding and generation capabilities. This type of model trains massive text data through deep learning algorithms, and can generate coherent and logical text content, greatly improving the efficiency and effectiveness of natural language processing. However, although large language models have made remarkable achievements in text generation, there are still many challenges in their controllability during actual application. Especially in some application scenarios that have strict requirements on text content, such as invitations, customer service, etc., the performance of large language models is often difficult to fully meet actual needs.

[0003] Specifically, in specific scenarios such as invitations and customer service, the controllability of text generation is extremely high. This includes but is not limited to: ensuring that the generated text conforms to a specific language style, accurately conveys the required information, avoids generating bias or inappropriate remarks, etc. However, existing large language models often find it difficult to achieve the expected accuracy and controllability when generating specific vocabulary in specific scenarios. This is mainly because the model focuses more on the overall coherence and logic of the text during training, while neglecting the fine control of vocabulary selection in specific scenarios. In addition, large language models may also face security issues such as biased remarks when generating text. Since the model is trained based on massive text data, these data may contain various biases and inappropriate remarks. Therefore, when generating text, the model may unconsciously copy or amplify these biases, thereby causing security issues.

[0004] In summary, although the existing large language models are widely used in dialogue scenarios such as question answering, they are still insufficient in terms of the controllability of generating specific vocabulary in specific scenarios. How to improve the controllability of large language models in the text generation process and ensure that the generated text is both accurate and safe has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] The main purpose of the present invention is to provide an artificial intelligence-based problem processing method, apparatus, computer equipment and storage medium, aiming to solve the technical problem that although large language models are widely used in dialogue scenarios such as question answering, there are still deficiencies in the controllability of generating specific vocabulary in specific scenarios, and the accuracy and security of the generated text cannot be guaranteed.

[0006] To achieve the above object, the present invention provides a problem-solving method based on artificial intelligence, the method comprising:

[0007] Receiving a question input by a user, and performing a search process on the question to obtain a corresponding reference text;

[0008] The question, the reference text and the preset prompt template are spliced ​​to obtain a corresponding target prompt;

[0009] Tokenizing the target prompt to obtain a corresponding target token set;

[0010] Calling a preset large model and creating a bias vector corresponding to a model vocabulary of the large model;

[0011] For each token in the target token set, a positive bias is set at the corresponding index position in the bias vector to obtain a corresponding target bias vector;

[0012] Reasoning the target prompt and the target bias vector based on the large model to obtain a corresponding response;

[0013] A question feedback process corresponding to the user is performed based on the reply.

[0014] Optionally, the searching and processing the question to obtain a corresponding reference text includes:

[0015] Call the preset database;

[0016] Performing a search process on the database based on the question to obtain a corresponding first text;

[0017] Rearranging the first text to obtain a corresponding second text;

[0018] The second text is used as the reference text.

[0019] Optionally, the step of combining the question, the reference text and a preset prompt template to obtain a corresponding target prompt includes:

[0020] Get the preset splicing strategy and verification strategy;

[0021] Based on the splicing strategy, the question, the reference text and the prompt template are spliced ​​to obtain a corresponding splicing result;

[0022] Verifying the splicing result based on the verification strategy;

[0023] If the splicing result passes the verification, the splicing result is used as the target prompt.

[0024] Optionally, the tokenizing the target prompt to obtain a corresponding target token set includes:

[0025] Call the preset language processing tool;

[0026] Tokenizing the target prompt based on the language processing tool to obtain a corresponding first token sequence;

[0027] Performing a duplicate data removal process on the first token sequence to obtain a corresponding first token set;

[0028] Perform irrelevant data filtering on the first token set to obtain a corresponding second token set;

[0029] The second token set is used as the target token set.

[0030] Optionally, the reasoning the target prompt and the target bias vector based on the large model to obtain a corresponding response includes:

[0031] By using the large model, sampling is performed in the target token set according to the target prompt and the target bias vector to gradually generate preliminary tokens of the reply;

[0032] Collect and process the preliminary tokens to obtain a corresponding second token sequence;

[0033] Decoding the second token sequence to obtain a corresponding decoded text;

[0034] Performing a preset sorting process on the decoded text to obtain a corresponding sorted text;

[0035] The collated text is used as the reply.

[0036] Optionally, the performing question feedback processing corresponding to the user based on the reply includes:

[0037] Get the preset reply optimization strategy;

[0038] Optimizing the response based on the response optimization strategy to obtain a corresponding target response;

[0039] The targeted reply is returned to the user.

[0040] Optionally, the optimizing the reply based on the reply optimization strategy to obtain a corresponding target reply includes:

[0041] Cleaning the reply based on a preset text cleaning tool to obtain a corresponding first reply;

[0042] Filter the first reply for sensitive content based on a preset filtering rule to obtain a corresponding second reply;

[0043] Performing format adjustment on the second reply to obtain a corresponding third reply;

[0044] The third reply is used as the target reply.

[0045] In addition, to achieve the above-mentioned purpose, the present invention also provides a problem-solving device based on artificial intelligence, and the problem-solving device based on artificial intelligence includes:

[0046] A search module is used to receive questions input by users and perform search processing on the questions to obtain corresponding reference texts;

[0047] A splicing module, used for splicing the question, the reference text and the preset prompt template to obtain a corresponding target prompt;

[0048] A processing module, used to tokenize the target prompt to obtain a corresponding target token set;

[0049] A creation module, used for calling a preset large model and creating a bias vector corresponding to a model vocabulary of the large model;

[0050] A setting module, used for setting a positive bias at a corresponding index position in the bias vector for each token in the target token set, to obtain a corresponding target bias vector;

[0051] A reasoning module, used for reasoning the target prompt and the target bias vector based on the large model to obtain a corresponding response;

[0052] An execution module is used to execute question feedback processing corresponding to the user based on the reply.

[0053] In order to solve the above technical problems, the embodiment of the present application further provides a computer device, which adopts the following technical solution:

[0054] The computer device includes a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the artificial intelligence-based problem-solving methods proposed in the embodiments of the present application are implemented.

[0055] In order to solve the above technical problems, the embodiment of the present application further provides a computer-readable storage medium, which adopts the following technical solution:

[0056] The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of any one of the artificial intelligence-based problem-solving methods proposed in the embodiments of the present application are implemented.

[0057] Compared with the prior art, the embodiments of the present application have the following beneficial effects:

[0058] The present invention provides a method, device, computer equipment and storage medium for problem processing based on artificial intelligence, the method comprising: first receiving a question input by a user, and performing a search process on the question to obtain a corresponding reference text; then performing a splicing process on the question, the reference text and a preset prompt template to obtain a corresponding target prompt; then performing a tokenization process on the target prompt to obtain a corresponding target token set; subsequently calling a preset large model, and creating a bias vector corresponding to the model vocabulary of the large model; and for each token in the target token set, setting a positive bias at the corresponding index position in the bias vector to obtain a corresponding target bias vector; further performing an inference on the target prompt and the target bias vector based on the large model to obtain a corresponding reply; and finally performing a problem feedback process corresponding to the user based on the reply. After receiving a question input by a user, the present invention retrieves the question to obtain a corresponding reference text, then concatenates the question, the reference text and a preset prompt template to obtain a corresponding target prompt, and tokenizes the target prompt to obtain a corresponding target token set, then calls a preset big model, and creates a bias vector corresponding to the model vocabulary of the big model, and for each token in the target token set, sets a positive bias at the corresponding index position in the bias vector to obtain a corresponding target bias vector, then infers the target prompt and the target bias vector based on the big model to obtain a corresponding reply, and finally executes question feedback processing corresponding to the user based on the reply. In this way, the present application adopts a problem processing method of only sampling tokens in the target token set when generating a reply based on the big model, which can effectively ensure that the generated reply content is accurate and controllable, thereby improving the controllability, accuracy and security of the reply. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the scheme in the present application, a brief introduction is given below to the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0060] Figure 1 is an exemplary system architecture diagram to which the present application may be applied;

[0061] Figure 2 is a flow chart of a problem-solving method based on artificial intelligence provided by an embodiment of the present invention;

[0062] Figure 3 is a structural schematic diagram of an embodiment of an artificial intelligence-based problem processing device according to the present application;

[0063] Figure 4 This is a basic structural block diagram of the computer device in this embodiment. DETAILED DESCRIPTION

[0064] The problem-solving method based on artificial intelligence provided by the embodiment of the present invention is applied to the problem-solving device based on artificial intelligence. Unless otherwise defined, all technical and scientific terms used in this document have the same meaning as those commonly understood by technicians in the technical field of this application; the terms used in the specification of the application herein are only for the purpose of describing specific embodiments and are not intended to limit this application; the terms "including" and "having" in the specification and claims of this application and the above-mentioned figure descriptions and any variations thereof are intended to cover non-exclusive inclusions. The terms "first", "second", etc. in the specification and claims of this application or the above-mentioned figures are used to distinguish different objects, not to describe a specific order.

[0065] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0066] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings.

[0067] like Figure 1As shown, the system architecture 100 may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0068] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social online platform software, etc.

[0069] Terminal devices 101, 102, 103 can be various electronic devices with display screens and supporting web browsing, including but not limited to smart phones, tablet computers, e-book readers, MP3 players (Moving Picture Experts Group Audio Layer III), MP4 (Moving Picture Experts Group Audio Layer IV) players, laptop computers, desktop computers, etc.

[0070] The server 105 may be a server that provides various services, such as a background server that provides support for web pages displayed on the terminal devices 101 , 102 , and 103 .

[0071] It should be noted that the artificial intelligence-based problem handling method provided in the embodiment of the present application is generally executed by a server / terminal device, and accordingly, the artificial intelligence-based problem handling device is generally set in the server / terminal device.

[0072] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0073] With the rapid development of information technology today, large language models (LLMs), as an important technology in the field of artificial intelligence, have shown great potential in various application scenarios with their powerful text understanding and generation capabilities. This type of model trains massive text data through deep learning algorithms, and can generate coherent and logical text content, greatly improving the efficiency and effectiveness of natural language processing. However, although large language models have made remarkable achievements in text generation, there are still many challenges in their controllability during actual application. Especially in some application scenarios that have strict requirements on text content, such as invitations, customer service, etc., the performance of large language models is often difficult to fully meet actual needs.

[0074] Specifically, in specific scenarios such as invitations and customer service, the controllability of text generation is extremely high. This includes but is not limited to: ensuring that the generated text conforms to a specific language style, accurately conveys the required information, avoids generating bias or inappropriate remarks, etc. However, existing large language models often find it difficult to achieve the expected accuracy and controllability when generating specific vocabulary in specific scenarios. This is mainly because the model focuses more on the overall coherence and logic of the text during training, while neglecting the fine control of vocabulary selection in specific scenarios. In addition, large language models may also face security issues such as biased remarks when generating text. Since the model is trained based on massive text data, these data may contain various biases and inappropriate remarks. Therefore, when generating text, the model may unconsciously copy or amplify these biases, thereby causing security issues.

[0075] In summary, although the existing large language models are widely used in dialogue scenarios such as question answering, they are still insufficient in terms of the controllability of generating specific vocabulary in specific scenarios. How to improve the controllability of large language models in the text generation process and ensure that the generated text is both accurate and safe has become a technical problem that needs to be solved urgently.

[0076] Continue to refer Figure 2 , shows a flow chart of an embodiment of the problem processing method based on artificial intelligence proposed in this application. The embodiment of this application can acquire and process relevant data based on artificial intelligence technology.

[0077] The problem solving method based on artificial intelligence provided by the embodiment of the present invention comprises the following steps:

[0078] S210, receiving a question input by a user, and performing a search process on the question to obtain a corresponding reference text.

[0079] In this step, the executor of the present invention may specifically be a question processing system, which may be referred to as the system for short. The present invention is applicable to scenarios that require high controllability of text generation and low diversity, such as summary generation, reading comprehension, intelligent question and answer, intelligent outbound calling, intelligent customer service, intelligent judgment, etc. Compared with other solutions, the present invention is particularly suitable for scenarios with small base model scale and poor basic capabilities. Among them, users can input questions through the system page, and the system will receive and store the questions input by the user. In addition, the specific implementation process of the above-mentioned retrieval process for the question to obtain the corresponding reference text will be further described in detail in the subsequent specific embodiments of the present invention, and will not be elaborated on here.

[0080] S220, combining the question, the reference text and a preset prompt template to obtain a corresponding target prompt.

[0081] In this step, the above-mentioned prompt template is a system prompt template (system prompt) used to describe the reply-related requirements according to actual business needs, which may specifically include model role settings, execution task flow, safe reply requirements, reply style restrictions, etc. Among them, the model role setting may include: clarifying the role played by the big model in the conversation, such as corporate customer service. This determines the language style, professional terms and possible scope of action of the reply. The execution task flow may include: outlining the basic steps that the big model should perform after receiving user input, such as understanding the question, retrieving information, generating a reply, etc. The safe reply requirements may include: listing the safety guidelines that must be followed, such as avoiding the use of impolite, offensive or discriminatory language, ensuring that the reply content is legal and compliant, and does not contain sensitive or non-compliant information. The reply style restrictions may include: determining the language style of the reply according to the role setting, such as formal, friendly or professional. This helps to maintain the consistency of the conversation and the user experience. In addition, the above-mentioned splicing of the question, the reference text and the preset prompt template to obtain the corresponding target prompt is further described in detail in the subsequent specific embodiments of the present invention, and will not be elaborated on here.

[0082] S230, tokenizing the target prompt to obtain a corresponding target token set.

[0083] In this step, the specific implementation process of tokenizing the target prompt to obtain the corresponding target token set will be further described in detail in the subsequent specific embodiments of the present invention, and will not be elaborated on here. Among them, only the reference text can be tokenized to obtain the target token set.

[0084] S240, calling a preset large model, and creating a bias vector corresponding to the model vocabulary of the large model.

[0085] In this step, there is no specific limitation on the selection of the above large model, which can be determined according to actual business needs, for example, large language models such as GPT, BERT, T5, etc. can be used. The above large model can be loaded and the optional token set (target token set) can be converted into a format acceptable to the large model, such as an ID list or a specific vocabulary index.

[0086] Alternatively, you can obtain the model vocabulary corresponding to the large model above and create a bias vector of the same size as the model vocabulary, initialized to zero.

[0087] S250, for each token in the target token set, set a positive bias at the corresponding index position in the bias vector to obtain a corresponding target bias vector.

[0088] In this step, the uncontrollability of text generation by the large model is manifested in the fact that the token combination of the large model generation result is inconsistent with expectations. In addition, since the potential combination space of tokens replied by the model is huge, it is labor-intensive and difficult to guarantee the effect to improve the controllability of the final result through post-processing. If the potential combination space of tokens replied by the model can be greatly reduced, then the possibility of the model generating content that does not meet expectations can be greatly reduced, which greatly improves the controllability.

[0089] Specifically, for each token in the target token set, a large positive bias (e.g., 1000 or higher) is set at the corresponding index position in the bias vector to increase the generation probability of these tokens, so that the target token set can be used as a sampling space in the decoding process of the subsequent large model to generate responses. If it is necessary to reduce the generation probability of certain tokens (e.g., stop words or undesirable words), a large negative value can be set at the corresponding index position in the bias vector.

[0090] S260, inferring the target prompt and the target bias vector based on the large model to obtain a corresponding response.

[0091] In this step, the large model reasoning and decoding is to input the spliced ​​target prompts into the large language model, and control the process of generating replies through specific strategies. In the process of reply generation, the use of an optional token set (target token set) for sampling can ensure that the generated reply content is correct and controllable. By setting a large logit bias for all tokens in the target token set, the generation probability values ​​of these tokens can be changed, thereby guiding the large model to give priority to these tokens when generating replies. Among them, the above-mentioned specific implementation process of reasoning the target prompt and the target bias vector based on the large model to obtain the corresponding reply will be further described in detail in the subsequent specific embodiments of the present invention, and will not be elaborated on here.

[0092] S270, executing question feedback processing corresponding to the user based on the reply.

[0093] In this step, the specific implementation process of executing the problem feedback processing corresponding to the user based on the reply will be further described in detail in the subsequent specific embodiments of the present invention, and will not be elaborated on here. In addition, after executing the problem feedback processing corresponding to the user based on the above reply, this round of dialogue ends. It is subsequently determined whether it is necessary to continue the dialogue. If it is necessary to continue the dialogue, the above problem processing steps (S210-S270) are re-executed. If not, the process ends.

[0094] In an embodiment of the present invention, a question input by a user is first received, and the question is retrieved to obtain a corresponding reference text; then the question, the reference text and a preset prompt template are spliced ​​to obtain a corresponding target prompt; then the target prompt is tokenized to obtain a corresponding target token set; subsequently, a preset large model is called, and a bias vector corresponding to the model vocabulary of the large model is created; and for each token in the target token set, a positive bias is set at the corresponding index position in the bias vector to obtain a corresponding target bias vector; further based on the large model, the target prompt and the target bias vector are inferred to obtain a corresponding reply; finally, based on the reply, question feedback processing corresponding to the user is performed. After receiving a question input by a user, the present invention retrieves the question to obtain a corresponding reference text, then concatenates the question, the reference text and a preset prompt template to obtain a corresponding target prompt, and tokenizes the target prompt to obtain a corresponding target token set, then calls a preset big model, and creates a bias vector corresponding to the model vocabulary of the big model, and for each token in the target token set, sets a positive bias at the corresponding index position in the bias vector to obtain a corresponding target bias vector, then infers the target prompt and the target bias vector based on the big model to obtain a corresponding reply, and finally executes question feedback processing corresponding to the user based on the reply. In this way, the present application adopts a problem processing method of only sampling tokens in the target token set when generating a reply based on the big model, which can effectively ensure that the generated reply content is accurate and controllable, thereby improving the controllability, accuracy and security of the reply.

[0095] Optionally, the searching and processing the question to obtain a corresponding reference text includes:

[0096] Call the preset database.

[0097] In this step, the database is a pre-built database storing relevant texts corresponding to different questions.

[0098] The database is searched based on the question to obtain a corresponding first text.

[0099] In this step, the retrieval process may refer to a recall process, that is, retrieving texts related to the question from the database to obtain the first text.

[0100] The first text is rearranged to obtain a corresponding second text.

[0101] In this step, the retrieved first texts may be sorted in descending order of relevance according to the relevance between the first texts and the above question, so as to obtain corresponding second texts.

[0102] The second text is used as the reference text.

[0103] In an embodiment of the present invention, a preset database is called; then the database is searched and processed based on the question to obtain a corresponding first text; then the first text is rearranged to obtain a corresponding second text; and the second text is subsequently used as the reference text. The present application calls a preset database, then searches and processes the database based on the question to obtain a corresponding first text, and rearranges the first text to efficiently and accurately retrieve reference texts related to the question, effectively improving the retrieval efficiency of the reference text and ensuring the data accuracy of the reference text obtained.

[0104] Optionally, the step of combining the question, the reference text and a preset prompt template to obtain a corresponding target prompt includes:

[0105] Get the preset splicing strategy and verification strategy.

[0106] In this step, the strategy content of the above-mentioned splicing strategy may include splicing the question, reference text, prompt template and historical conversation (if any) together in a certain format. When splicing, it is necessary to pay attention to maintaining the coherence and logic of the information to ensure that the large model can accurately understand the information in the prompt. Preferably, the specific splicing method can follow the following steps: First, use the prompt template as the basic framework. Then, insert the user question into the specified position in the prompt template. Next, add the retrieved reference text as additional information to the prompt. If there are multiple reference texts, they can be sorted according to relevance or importance. Finally, if there is a historical conversation, it is also integrated into the prompt, usually placed after the user question and reference text.

[0107] The above verification strategy is a strategy for verifying and adjusting the splicing results to ensure their accuracy and completeness. The verification strategy includes: checking whether the information in the prompt is consistent, whether the logic is smooth, whether there is any missing or redundant information, etc. If any problems are found, they need to be adjusted and corrected in time.

[0108] The question, the reference text and the prompt template are spliced ​​based on the splicing strategy to obtain a corresponding splicing result.

[0109] In this step, the question, the reference text and the prompt template may be spliced ​​according to the strategy content of the splicing strategy, so as to obtain a corresponding splicing result.

[0110] The splicing result is verified based on the verification strategy.

[0111] In this step, the splicing result can be verified according to the policy content of the above verification policy, and a corresponding verification result can be obtained. Specifically, if the information in the prompt is detected to be consistent, logically coherent, and there is no missing or redundant information, the splicing result is determined to have passed the verification, otherwise the splicing result is determined to have failed the verification.

[0112] If the splicing result passes the verification, the splicing result is used as the target prompt.

[0113] In an embodiment of the present invention, a preset splicing strategy and a verification strategy are obtained; then the question, the reference text, and the prompt template are spliced ​​based on the splicing strategy to obtain a corresponding splicing result; the splicing result is subsequently verified based on the verification strategy; if the splicing result passes the verification, the splicing result is used as the target prompt. The present invention splices the question, the reference text, and the prompt template based on the use of a splicing strategy to obtain a corresponding splicing result, and then verifies the splicing result based on the use of a verification strategy, and when it is detected that the splicing result passes the verification, the splicing result is used as the target prompt, effectively ensuring the compliance and accuracy of the splicing result.

[0114] Optionally, the tokenizing the target prompt to obtain a corresponding target token set includes:

[0115] Call the preset language processing tool.

[0116] In this step, the language processing tool is a tool or library that can perform tokenization, and specifically may be a natural language processing (NLP) library, such as NLTK, SpaCy, Hugging Face's Transformers library, etc., which provide the function of segmenting text into words, subwords (such as BPE, WordPiece), characters or other units.

[0117] The target prompt is tokenized based on the language processing tool to obtain a corresponding first token sequence.

[0118] In this step, the target prompt is tokenized by using the selected language processing tool to obtain the corresponding first token sequence. The tokenization process divides the text into a sequence of one or more tokens, each of which may be a word, phrase, subword or character sequence.

[0119] Perform deduplication processing on the first token sequence to obtain the corresponding first token set.

[0120] In this step, after obtaining the first token sequence, a unique token set needs to be extracted from the first token sequence, that is, the duplicate tokens in the first token sequence are removed, and only non-duplicate words are retained, so as to obtain the corresponding first token set.

[0121] Perform irrelevant data filtering processing on the first token set to obtain the corresponding second token set.

[0122] In this step, according to the task requirements, the extracted first token set can be further filtered to remove stop words (such as "de", "le", etc.) or other unnecessary tokens, so as to obtain the required second token set.

[0123] Use the second token set as the target token set.

[0124] In this step, the target token set will be used as the optional vocabulary space when the large model generates a response. Among them, the processed target token set can be prepared in a format acceptable to the large model. Specifically, it involves mapping the target token set to a specific ID, or saving the target token set as a file for use during large model inference.

[0125] In the embodiment of the present invention, by calling a preset language processing tool; then performing tokenization processing on the target prompt based on the language processing tool to obtain the corresponding first token sequence; then performing deduplication processing on the first token sequence to obtain the corresponding first token set; subsequently performing irrelevant data filtering processing on the first token set to obtain the corresponding second token set; finally using the second token set as the target token set. In this application, tokenization processing is performed on the target prompt based on the use of the language processing tool to obtain the corresponding first token sequence, and then the first token sequence will be automatically and intelligently subjected to deduplication processing and irrelevant data filtering processing, so as to accurately obtain the required target token set, effectively ensuring the conciseness and accuracy of the obtained target token set, and further facilitating improving the controllability and accuracy of the subsequent response generated by the large model.

[0126] Optionally, the inference of the target prompt and the target bias vector based on the large model to obtain the corresponding response includes:

[0127] Through the large model, sampling is performed in the target token set according to the target prompt and the target bias vector to gradually generate preliminary tokens of the reply.

[0128] In this step, after passing the above target hint and target bias vector as input to the large model, you can further set sampling parameters such as temperature, top-k, top-p, etc. according to your needs. These parameters will affect the way of sampling from the optional token set (target token set). The large model will then sample from the target token set according to the target hint and sampling parameters, and gradually generate preliminary tokens for the reply. This involves multiple iterations, and each iteration generates one or more tokens.

[0129] The preliminary tokens are collected and processed to obtain a corresponding second token sequence.

[0130] In this step, all tokens generated by the large model can be collected to form a complete token sequence, that is, the second token sequence mentioned above.

[0131] The second token sequence is decoded to obtain a corresponding decoded text.

[0132] In this step, the generated second token sequence can be decoded back into human-readable text to obtain the corresponding decoded text. The decoding process includes converting the token ID back into the corresponding vocabulary. The target token set is converted into a format acceptable to the large model (such as a token ID list) in advance.

[0133] The decoded text is subjected to a preset collation process to obtain a corresponding collated text.

[0134] In this step, the above-mentioned sorting process includes removing special marks, processing punctuation marks, adjusting the format, etc. to ensure the accuracy and fluency of the reply.

[0135] The collated text is used as the reply.

[0136] In an embodiment of the present invention, through the large model, according to the target prompt and the target bias vector, sampling is performed in the target token set to gradually generate preliminary tokens of the reply; then the preliminary tokens are collected and processed to obtain the corresponding second token sequence; then the second token sequence is decoded to obtain the corresponding decoded text; subsequently the decoded text is subjected to a preset arrangement process to obtain the corresponding arranged text; finally, the arranged text is used as the reply. The present invention uses a large model to sample in the target token set according to the target prompt and the target bias vector to gradually generate preliminary tokens of the reply, then the preliminary tokens are collected and processed to obtain the corresponding second token sequence, and the second token sequence is decoded to obtain the corresponding decoded text, and then the decoded text is subjected to a preset arrangement process, so that the model inference process can be automatically and accurately completed and the corresponding reply is generated, effectively ensuring the accuracy and fluency of the generated reply.

[0137] Optionally, the performing question feedback processing corresponding to the user based on the reply includes:

[0138] Gets a preset reply optimization strategy.

[0139] In this step, the reply optimization strategy may refer to a strategy for performing text cleaning, content filtering, and format adjustment on the reply to further improve the quality and accuracy of the reply generated by the large model.

[0140] The reply is optimized based on the reply optimization strategy to obtain a corresponding target reply.

[0141] In this step, the specific implementation process of optimizing the reply based on the reply optimization strategy to obtain the corresponding target reply will be described in further detail in subsequent specific embodiments of the present invention and will not be elaborated on here.

[0142] The targeted reply is returned to the user.

[0143] In this step, the optimized target response may be presented to the user through a user interface.

[0144] In an embodiment of the present invention, a preset reply optimization strategy is obtained; then the reply is optimized based on the reply optimization strategy to obtain a corresponding target reply; and the target reply is subsequently returned to the user. After the present invention infers the target prompt and the target bias vector based on the large model to obtain the corresponding reply, it will also intelligently optimize the reply based on the use of the reply optimization strategy to obtain the corresponding target reply, thereby further improving the quality and accuracy of the reply generated by the large model, so that the optimized target reply is subsequently returned to the user, which can be beneficial to improving the user experience and improving user satisfaction.

[0145] Optionally, the optimizing the reply based on the reply optimization strategy to obtain a corresponding target reply includes:

[0146] The reply is cleaned up based on a preset text cleaning tool to obtain a corresponding first reply.

[0147] In this step, the text cleaning tool is a tool with an automated cleaning function. The cleaning process may include removing irrelevant characters, punctuation marks or extra spaces in the generated reply, and correcting grammatical errors.

[0148] The first reply is filtered for sensitive content based on preset filtering rules to obtain a corresponding second reply.

[0149] In this step, filtering rules are pre-set according to actual business needs to perform keyword matching or semantic analysis on the generated replies to identify and filter out sensitive content that does not meet the requirements or is inappropriate.

[0150] The second reply is formatted to obtain a corresponding third reply.

[0151] In this step, the format adjustment process includes splicing and formatting the generated second reply to make the entire reply more fluent and consistent.

[0152] The third reply is used as the target reply.

[0153] In an embodiment of the present invention, the reply is cleaned based on a preset text cleaning tool to obtain a corresponding first reply; then the first reply is filtered for sensitive content based on a preset filtering rule to obtain a corresponding second reply; then the second reply is formatted to obtain a corresponding third reply; and the third reply is subsequently used as the target reply. The present invention cleans the reply based on the use of a text cleaning tool, filters the reply for sensitive content based on the use of filtering rules, and formats the reply, thereby automatically and intelligently completing the optimization processing of the reply, which is beneficial to further improving the quality and accuracy of the reply generated by the large model.

[0154] In some optional implementations, the user information obtained is subject to the user's consent and complies with relevant laws and policies.

[0155] In addition, the present invention converts the large language model text generation task into a reading comprehension task. The answer to the reading comprehension task should be found in the reference text. Therefore, the reference text of each question is found as a prompt to input into the large language model. When the large language model generates a reply, it only samples the tokens in the prompt, which can ensure that the generated content is correct while greatly improving the controllability of the generated content. The innovative points of the present invention include: creatively combining prompt learning, decoding intervention, and retrieval enhancement, and using the tokens in the prompt as the sampling space for decoding.

[0156] Further references Figure 3 , as a response to the above Figure 2 In order to realize the method shown in the figure, the present application provides an embodiment of a problem processing device 300 based on artificial intelligence, and the embodiment of the device is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0157] An embodiment of the present invention provides an artificial intelligence-based problem processing device 300, and the artificial intelligence-based problem processing device 300 includes:

[0158] The search module 310 is used to receive a question input by a user and perform a search process on the question to obtain a corresponding reference text;

[0159] A splicing module 320 is used to splice the question, the reference text and the preset prompt template to obtain a corresponding target prompt;

[0160] A processing module 330 is used to tokenize the target prompt to obtain a corresponding target token set;

[0161] A creation module 340, for calling a preset large model and creating a bias vector corresponding to a model vocabulary of the large model;

[0162] A setting module 350, configured to set a positive bias at a corresponding index position in the bias vector for each token in the target token set, to obtain a corresponding target bias vector;

[0163] A reasoning module 360, configured to reason the target prompt and the target bias vector based on the large model to obtain a corresponding response;

[0164] The execution module 370 is used to execute the question feedback processing corresponding to the user based on the reply.

[0165] Optionally, the retrieval module 310 includes:

[0166] The first calling submodule is used to call a preset database;

[0167] A retrieval submodule, configured to perform a retrieval process on the database based on the question to obtain a corresponding first text;

[0168] A rearrangement submodule, used for rearranging the first text to obtain a corresponding second text;

[0169] The first determining submodule is configured to use the second text as the reference text.

[0170] Optionally, the splicing module 320 includes:

[0171] The first acquisition submodule is used to acquire a preset splicing strategy and a verification strategy;

[0172] A splicing submodule, used for splicing the question, the reference text and the prompt template based on the splicing strategy to obtain a corresponding splicing result;

[0173] A verification submodule, used for verifying the splicing result based on the verification strategy;

[0174] The second determination submodule is used to use the splicing result as the target prompt if the splicing result passes the verification.

[0175] Optionally, the processing module 330 includes:

[0176] The second calling submodule is used to call a preset language processing tool;

[0177] A first processing submodule, configured to tokenize the target prompt based on the language processing tool to obtain a corresponding first token sequence;

[0178] A second processing submodule, configured to perform a duplicate data removal process on the first token sequence to obtain a corresponding first token set;

[0179] A third processing submodule, configured to perform irrelevant data filtering processing on the first token set to obtain a corresponding second token set;

[0180] The third determination submodule is used to use the second token set as the target token set.

[0181] Optionally, the reasoning module 360 ​​includes:

[0182] A generation submodule, for sampling in the target token set to gradually generate preliminary tokens of the reply according to the target prompt and the target bias vector through the large model;

[0183] A collection submodule, used to collect and process the preliminary tokens to obtain a corresponding second token sequence;

[0184] A decoding submodule, used to decode the second token sequence to obtain a corresponding decoded text;

[0185] The arranging submodule is used to perform a preset arranging process on the decoded text to obtain a corresponding arranging text;

[0186] The fourth determining submodule is used to use the sorted text as the reply.

[0187] Optionally, the execution module 370 includes:

[0188] The second acquisition submodule is used to acquire a preset reply optimization strategy;

[0189] An optimization submodule, used to optimize the reply based on the reply optimization strategy to obtain a corresponding target reply;

[0190] The return submodule is used to return the target reply to the user.

[0191] Optionally, the optimization submodule includes:

[0192] A cleaning unit, configured to clean the reply based on a preset text cleaning tool to obtain a corresponding first reply;

[0193] A filtering unit, configured to filter the first reply for sensitive content based on a preset filtering rule to obtain a corresponding second reply;

[0194] an adjusting unit, configured to perform format adjustment processing on the second reply to obtain a corresponding third reply;

[0195] A determination unit is configured to use the third reply as the target reply.

[0196] To solve the above technical problems, the present application also provides a computer device. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.

[0197] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected through a system bus. It should be noted that the figure only shows a computer device 4 with components 41-43, but it should be understood that it is not required to implement all the components shown, and more or fewer components can be implemented instead. Among them, those skilled in the art can understand that the computer device here is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application specific integrated circuits (Application Specific Integrated Circuit, ASIC), programmable gate arrays (Field-Programmable Gate Array, FPGA), digital processors (Digital Signal Processor, DSP), embedded devices, etc.

[0198] The computer device may be a computing device such as a desktop computer, a notebook, a PDA, a cloud server, etc. The computer device may interact with a user through a keyboard, a mouse, a remote controller, a touch pad, or a voice control device.

[0199] The memory 41 includes at least one type of readable storage medium, and the readable storage medium includes flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, disk, optical disk, etc. In some embodiments, the memory 41 can be an internal storage unit of the computer device 4, such as a hard disk or memory of the computer device 4. In other embodiments, the memory 41 can also be an external storage device of the computer device 4, such as a plug-in hard disk equipped on the computer device 4, a smart memory card (Smart Media Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (FlashCard), etc. Of course, the memory 41 can also include both the internal storage unit of the computer device 4 and its external storage device. In this embodiment, the memory 41 is generally used to store the operating system and various application software installed on the computer device 4, such as the program code of the problem processing method based on artificial intelligence, etc. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or are to be output.

[0200] The processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips in some embodiments. The processor 42 is generally used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to run the program code stored in the memory 41 or process data, such as running the program code of the problem-solving method based on artificial intelligence.

[0201] The network interface 43 may include a wireless network interface or a wired network interface. The network interface 43 is generally used to establish a communication connection between the computer device 4 and other electronic devices.

[0202] The present application also provides another implementation, namely, providing a computer-readable storage medium, wherein the computer-readable storage medium stores the application crash processing program, and the application crash processing program can be executed by at least one processor to enable the at least one processor to perform the steps of the above-mentioned artificial intelligence-based problem handling method.

[0203] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware online platform, and of course, by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present application.

[0204] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0205] Obviously, the embodiments described above are only some embodiments of the present application, rather than all embodiments. The preferred embodiments of the present application are given in the accompanying drawings, but they do not limit the patent scope of the present application. The present application can be implemented in many different forms. On the contrary, the purpose of providing these embodiments is to make the understanding of the disclosure of the present application more thorough and comprehensive. Although the present application is described in detail with reference to the aforementioned embodiments, for those skilled in the art, it is still possible to modify the technical solutions recorded in the aforementioned specific implementation methods, or to perform equivalent replacement of some of the technical features therein. Any equivalent structure made using the contents of the specification and drawings of this application, directly or indirectly used in other related technical fields, is similarly within the scope of patent protection of this application.

Claims

1. A problem-solving method based on artificial intelligence, characterized in that: include: Receiving a question input by a user, and performing a search process on the question to obtain a corresponding reference text; The question, the reference text and the preset prompt template are spliced ​​to obtain a corresponding target prompt; Tokenizing the target prompt to obtain a corresponding target token set; Calling a preset large model and creating a bias vector corresponding to a model vocabulary of the large model; For each token in the target token set, a positive bias is set at the corresponding index position in the bias vector to obtain a corresponding target bias vector; Reasoning the target prompt and the target bias vector based on the large model to obtain a corresponding response; A question feedback process corresponding to the user is performed based on the reply.

2. The method according to claim 1, characterized in that The step of performing a search process on the question to obtain a corresponding reference text includes: Call the preset database; Performing a search process on the database based on the question to obtain a corresponding first text; Rearranging the first text to obtain a corresponding second text; The second text is used as the reference text.

3. The method according to claim 1, characterized in that The step of combining the question, the reference text and the preset prompt template to obtain a corresponding target prompt includes: Get the preset splicing strategy and verification strategy; Based on the splicing strategy, the question, the reference text and the prompt template are spliced ​​to obtain a corresponding splicing result; Verifying the splicing result based on the verification strategy; If the splicing result passes the verification, the splicing result is used as the target prompt.

4. The method according to claim 1, characterized in that: The target prompt is tokenized to obtain a corresponding target token set, including: Call the preset language processing tool; Tokenizing the target prompt based on the language processing tool to obtain a corresponding first token sequence; Performing a duplicate data removal process on the first token sequence to obtain a corresponding first token set; Perform irrelevant data filtering on the first token set to obtain a corresponding second token set; The second token set is used as the target token set.

5. The method according to claim 1, characterized in that The reasoning of the target prompt and the target bias vector based on the large model to obtain a corresponding response includes: By using the large model, sampling is performed in the target token set according to the target prompt and the target bias vector to gradually generate preliminary tokens of the reply; Collect and process the preliminary tokens to obtain a corresponding second token sequence; Decoding the second token sequence to obtain a corresponding decoded text; Performing a preset sorting process on the decoded text to obtain a corresponding sorted text; The collated text is used as the reply.

6. The method according to claim 1, characterized in that The performing of question feedback processing corresponding to the user based on the reply includes: Get the preset reply optimization strategy; Optimizing the response based on the response optimization strategy to obtain a corresponding target response; The targeted reply is returned to the user.

7. The method according to claim 6, characterized in that The optimizing the reply based on the reply optimization strategy to obtain a corresponding target reply includes: Cleaning the reply based on a preset text cleaning tool to obtain a corresponding first reply; Filter the first reply for sensitive content based on a preset filtering rule to obtain a corresponding second reply; Performing format adjustment on the second reply to obtain a corresponding third reply; The third reply is used as the target reply.

8. A problem-solving device based on artificial intelligence, characterized in that: include: A search module is used to receive questions input by users and perform search processing on the questions to obtain corresponding reference texts; A splicing module, used for splicing the question, the reference text and the preset prompt template to obtain a corresponding target prompt; A processing module, used to tokenize the target prompt to obtain a corresponding target token set; A creation module, used for calling a preset large model and creating a bias vector corresponding to a model vocabulary of the large model; A setting module, used for setting a positive bias at a corresponding index position in the bias vector for each token in the target token set, to obtain a corresponding target bias vector; A reasoning module, used for reasoning the target prompt and the target bias vector based on the large model to obtain a corresponding response; An execution module is used to execute question feedback processing corresponding to the user based on the reply.

9. A computer device, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of the problem processing method based on artificial intelligence as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the artificial intelligence-based problem-solving method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Large model reasoning consistency evaluation method, device, equipment, medium and product

    CN122285460A