Question and answer method and device based on large model, equipment and medium

Through the big model-based Q&A method, the pre-trained model is used to convert the questions into embedded vectors and perform similarity calculations with the standard question bank, which solves the problem that traditional Q&A systems are difficult to deal with complex and illegal problems, and realizes a more efficient, accurate and safe Q&A system.

CN119988562APending Publication Date: 2025-05-13SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510132095.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-06
Publication Date
2025-05-13

Smart Images

  • Figure CN119988562A_ABST
    Figure CN119988562A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large models, in particular to a question answering method and device based on a large model, equipment and a medium, and the method comprises the steps: carrying out the text conversion according to a to-be-processed question through a pre-training model, and obtaining an embedded vector corresponding to the to-be-processed question; calculating the similarity between the to-be-processed problem and the standard problem according to the embedded vector and the standard embedded vector; when a first preset similarity threshold value corresponding to the standard problem which is an illegal problem is smaller than a preset similarity threshold value corresponding to the standard problem which is a legal problem, according to the similarity of the standard problem and the preset similarity threshold value corresponding to the standard problem, judging whether the standard problem is an illegal problem or not; determining a target standard question corresponding to the to-be-processed question and a target answer corresponding to the target standard question; if the target standard question is an illegal question, generating a prompt statement, and outputting the prompt statement; and if the target standard question is not the illegal question, outputting a target answer. The question and answer data security can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large models, and in particular to a question-answering method, device, equipment and medium based on large models. Background Art

[0002] With the rapid development of Internet technology, online interaction has become an indispensable part of daily life and work. As an important channel for information acquisition and communication, the efficiency and accuracy of online question-answering systems directly affect user experience and the harmony of the network environment. However, in the increasingly complex network environment, question-answering systems face many challenges, especially the proliferation of illegal questions and tricky questions, which pose a serious threat to the security of the network ecosystem.

[0003] Most of the current traditional question-answering systems are based on keyword matching, template matching or simple rule reasoning. These methods are good for dealing with routine questions, but they are powerless when facing complex and changeable network problems with unclear intentions. Traditional question-answering systems often find it difficult to effectively deal with illegal questions that are cleverly designed and have unclear intentions, which poses a severe challenge to the harmony and security of the network environment.

[0004] Therefore, how to improve the data security of question and answer has become a technical problem that needs to be solved urgently in this field. Summary of the invention

[0005] The purpose of the present invention is to provide a question-answering method, device, equipment and medium based on a large model, which can improve the data security of the question-answering.

[0006] In a first aspect, a question-answering method based on a large model is provided, comprising:

[0007] Get pending questions input by the user;

[0008] According to the problem to be processed, use the pre-trained model to perform text conversion to obtain an embedding vector corresponding to the problem to be processed;

[0009] Calculate the similarity between the problem to be processed and the standard problem according to the embedding vector corresponding to the problem to be processed and the standard embedding vector corresponding to the standard problem in the standard problem library; the standard problem library includes a plurality of standard embedding vectors corresponding to the standard problems and a plurality of standard answers corresponding to the standard problems;

[0010] Determining a preset similarity threshold corresponding to the standard question;

[0011] When the first preset similarity threshold corresponding to the illegal question of the standard question is less than the preset similarity threshold corresponding to the legal question of the standard question, determining the target standard question corresponding to the question to be processed and the target answer corresponding to the target standard question according to the similarity of the standard question and the preset similarity threshold corresponding to the standard question;

[0012] If the target standard question is an illegal question, a prompt statement is generated and output; if the target standard question is not an illegal question, the target answer is output.

[0013] In a preferred example, the present invention can be further configured as follows: according to the similarity of the standard question and the preset similarity threshold corresponding to the standard question, determining the target standard question corresponding to the problem to be processed and the target answer corresponding to the target standard question, including:

[0014] Determining whether the similarity of the standard question is greater than a preset similarity threshold corresponding to the standard question;

[0015] The standard question with a similarity greater than a preset similarity threshold corresponding to the standard question is used as the first standard question;

[0016] According to the similarity, the target standard question is determined from the first standard questions; and the target answer corresponding to the target standard question is retrieved from the standard question library.

[0017] In a preferred example, the present invention can be further configured as follows: determining the target standard question from the first standard questions according to the similarity, including:

[0018] If the first standard question includes a plurality of first standard questions of different categories, determining whether the first standard question includes an illegal question;

[0019] If the first standard question of the illegal question is included, then the first standard question with the greatest similarity is selected from the first standard questions of the illegal question as the target standard question corresponding to the question to be processed;

[0020] If the first standard question of the illegal question is not included, then the first standard question with the greatest similarity is selected from the first standard questions as the target standard question corresponding to the question to be processed.

[0021] In a preferred example, the present invention can be further configured as follows: before calculating the similarity between the problem to be processed and the standard problem according to the embedding vector corresponding to the problem to be processed and the standard embedding vector corresponding to the standard problem in the standard problem library, the method further includes:

[0022] Obtaining a plurality of the standard questions and standard answers corresponding to the plurality of the standard questions;

[0023] Performing text conversion on the plurality of standard questions using the pre-trained model to obtain embedding vectors corresponding to the standard questions;

[0024] The embedding vectors corresponding to each of the plurality of standard questions and the standard answers corresponding to each of the plurality of standard questions are stored in the standard question library.

[0025] In a preferred example, the present invention can be further configured to: store the embedding vectors corresponding to the plurality of standard questions and the standard answers corresponding to the plurality of standard questions in the standard question library, including:

[0026] Extracting data features corresponding to each of the plurality of standard questions;

[0027] Normalizing the data features corresponding to each of the plurality of standard questions to obtain normalized features corresponding to each of the plurality of standard questions;

[0028] Based on the normalized features corresponding to each of the plurality of standard questions, density clustering is performed to obtain a plurality of categories of data sets;

[0029] The standard question of selecting the target number from each data set is input into the classification algorithm to obtain the corresponding category of each data set;

[0030] The category with the largest number of occurrences in the category corresponding to each data set is taken as the target category of the data set;

[0031] According to the categories corresponding to each of the multiple standard questions, the embedding vectors corresponding to each of the multiple standard questions, and the standard answers corresponding to each of the multiple standard questions, they are stored in the standard question library.

[0032] In a preferred example, the present invention can be further configured as follows:

[0033] According to the user's question feedback information, the answers to the standard questions corresponding to the question feedback information in the standard question library are updated;

[0034] and / or;

[0035] Regularly adding embedding vectors corresponding to new standard questions and answers corresponding to new standard questions in the standard question library;

[0036] and / or;

[0037] After getting the pending questions from the user, it also includes:

[0038] Matching the to-be-processed question with a historical search question in a preset question library; the preset question library includes a plurality of historical search questions and answers corresponding to the plurality of historical search questions;

[0039] If there is a successfully matched target history retrieval question, determining the answer corresponding to the target history retrieval question from the preset question library, and outputting the answer corresponding to the target history retrieval question;

[0040] If there is no successfully matched target historical retrieval question, a step is performed to convert text using a pre-trained model according to the problem to be processed to obtain an embedding vector corresponding to the problem to be processed.

[0041] In a preferred example, the present invention can be further configured as follows:

[0042] Obtaining an initial training model and a training sample set, wherein the training sample set includes a plurality of training problem samples and training embedding vectors corresponding to each of the plurality of training problem samples;

[0043] Using the training sample set, training the initial training model to obtain a pre-training model;

[0044] According to a first total amount of illegal questions and a second total amount of legal questions of the training question samples in the training sample set, a preset similarity threshold corresponding to the illegal questions and a preset similarity threshold corresponding to the legal questions are determined.

[0045] In a second aspect, a question-answering device based on a large model is provided, comprising:

[0046] The acquisition module is used to obtain pending issues input by the user;

[0047] A text conversion module, used to perform text conversion using a pre-trained model according to the problem to be processed, to obtain an embedding vector corresponding to the problem to be processed;

[0048] A similarity calculation module, configured to calculate the similarity between the problem to be processed and the standard problem according to the embedding vector corresponding to the problem to be processed and the standard embedding vector corresponding to the standard problem in the standard problem library; the standard problem library includes the standard embedding vectors corresponding to the plurality of standard problems and the standard answers corresponding to the plurality of standard problems;

[0049] A threshold determination module, used to determine a preset similarity threshold corresponding to the standard question;

[0050] A matching module, configured to determine a target standard question corresponding to the problem to be processed and a target answer corresponding to the target standard question according to the similarity of the standard question and the preset similarity threshold corresponding to the standard question, when a first preset similarity threshold corresponding to the standard question being an illegal question is less than a preset similarity threshold corresponding to the standard question being a legal question;

[0051] The reply module is used to generate a prompt statement and output the prompt statement if the target standard question is an illegal question; if the target standard question is not an illegal question, output the target answer.

[0052] According to a third aspect, an electronic device is provided. The electronic device includes a memory and a processor. A computer program is stored in the memory. When the processor runs the computer program, the method described in any one of the first aspects is executed.

[0053] In a fourth aspect, a computer-readable storage medium is provided, wherein at least one program code is stored in the computer-readable storage medium, and the program code is loaded and executed by a processor to implement any method as described in the first aspect.

[0054] In a fifth aspect, a computer program product is provided, comprising a computer program or instructions, wherein when the computer program or instructions are executed by a processor, the method described in any one of the first aspects is implemented.

[0055] In summary, the question-answering method based on a large model provided by the present invention includes the following beneficial technical effects:

[0056] In this solution, the problem to be processed is obtained from user input; the pre-trained model is used for text conversion to convert the problem to be processed into an embedding vector, so as to convert the natural language problem into a point in a high-dimensional space to obtain an embedding vector, which can capture the semantic information of the problem and understand the user's intention; by calculating the similarity between the embedding vector of the problem to be processed and the standard embedding vector of the standard problem in the standard question library, an accurate match of the problem to be processed is achieved; at the same time, different preset similarity thresholds are set to distinguish between illegal questions and legal questions, so that the preset similarity threshold of illegal questions is lower, and potential illegal questions can be identified more sensitively. Then, after determining the target standard question of the problem to be processed and its corresponding target answer, they are processed separately according to whether the target standard question is an illegal question, which greatly improves the efficiency and accuracy of the question-answering system and the data security of the question-answering.

[0057] In addition, the present invention also provides a question-and-answer device, equipment and medium based on a large model, all of which have the above-mentioned beneficial technical effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] In order to more clearly illustrate the embodiments of the present invention or the technical solutions of the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0059] Figure 1 It is a flowchart of a question-answering method based on a large model provided by an embodiment of the present invention;

[0060] Figure 2 is a flowchart of another large model-based question-answering method provided by an embodiment of the present invention;

[0061] Figure 3 is a structural schematic diagram of a large model-based question-answering device provided by an embodiment of the present invention;

[0062] Figure 4 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0063] This specific embodiment is merely an explanation of the present invention and is not a limitation of the present invention. After reading this specification, those skilled in the art may make non-creative modifications to the present embodiment as needed, but such modifications are protected by patent law as long as they are within the scope of the present invention.

[0064] It should be noted that in the optional embodiments of the present invention, the object information and other related data involved, when the embodiments of the present invention are applied to specific products or technologies, need to obtain the permission or consent of the object, and the collection, use and processing of the relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions. In other words, if the embodiments of the present invention involve data related to the object, it needs to be obtained with the authorization and consent of the object, the authorization and consent of the relevant department, and in compliance with the relevant laws, regulations and standards of the country and region. If personal information is involved in the embodiments, the acquisition of all personal information needs to obtain the consent of the individual. If sensitive information is involved, the separate consent of the information subject needs to be obtained. The embodiments also need to be implemented with the authorization and consent of the object.

[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] In addition, the term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article, unless otherwise specified, generally means that the associated objects before and after are in an "or" relationship.

[0067] In the current field of network security, we also focus on improving the ability to respond intelligently to complex, changeable and potentially illegal network interaction problems. In today's increasingly complex and changeable Internet environment, traditional question-and-answer systems often find it difficult to effectively handle those cleverly designed, unclearly intended illegal questions or tricky questions, which poses a severe challenge to the harmony and security of the network environment. Illegal questions are often carefully designed to circumvent keyword filtering, while tricky questions may involve knowledge in multiple fields, requiring the system to have higher semantic understanding and reasoning capabilities.

[0068] With the continuous maturity of big data technology and the widespread application of deep learning algorithms, it is possible to upgrade the network question-answering system to be intelligent. Big data technology enables the system to process massive amounts of network interaction data and explore potential patterns and rules; deep learning models, especially pre-trained models in the field of natural language processing, such as BERT and GPT, have shown unprecedented performance improvements through huge parameter scales and massive training data. They can handle more complex and diverse tasks and have significantly improved the system's semantic understanding and generation capabilities. Although big data and deep learning technologies have shown great potential in network question-answering systems, there are still many shortcomings in existing research and applications.

[0069] In view of the above background, in order to deal with the difficulties of illegal questions, tricky questions and answers in today's complex and changeable online interactive questions and answers, question banks based on big data and deep learning models have become a research hotspot. The present invention proposes a network interactive question and answer processing technology based on big data analysis and deep learning models, aiming to achieve accurate identification and effective interception of illegal network questions by building an efficient and intelligent question and answer system; by collecting and processing massive amounts of network interaction data, and deep learning models can learn and extract features from these data to improve the accuracy of question identification; the invention relates to the field of network security and characteristic answers to specific questions, especially answering technology through large-scale data processing and deep learning model application. Specifically, through the pre-training model, the question text can be mapped to points in a high-dimensional space, which can capture the semantic information of the question and obtain an embedded vector. The text in the question bank is encoded into a vector form through the same pre-training model and stored in a high-performance vector database (i.e., a standard question bank). By building an index, question vectors can be quickly retrieved and compared, thereby improving the response speed of the question-answering system. Specifically, in the question-answering system, by calculating the cosine similarity between the question vector and the vector in the question library and comparing it with the preset threshold, it can be determined whether the question is illegal or tricky, and illegal questions and tricky questions can be intelligently identified and blocked, effectively responding to various types of network illegal question threats, thereby achieving intelligent identification and accurate problem finding.

[0070] Furthermore, building an efficient and intelligent question-answering system requires not only the dual-wheel drive of big data analysis and deep learning technology, but also the continuous optimization of algorithms and models to achieve accurate identification and effective interception of illegal network issues, while ensuring accurate answers to specific questions.

[0071] Through the above methods, the performance of the online question-and-answer system can be significantly improved, network security can be enhanced, and users can be provided with safer, more accurate, and more efficient question-and-answer services.

[0072] Specifically, the embodiment of the present invention provides a question-answering method based on a large model, such as Figure 1 As shown, the method provided in the embodiment of the present invention can be executed by an electronic device, and the electronic device is a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services. The terminal device can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc., but is not limited to this. The terminal device and the electronic device can be directly or indirectly connected through wired or wireless communication, and the embodiment of the present invention is not limited here. The method includes:

[0073] S101, obtaining a pending question input by a user;

[0074] The user may input text or voice in an input box on the display interface, so as to obtain the pending question input by the user. The pending question is a question that the user hopes to be able to answer.

[0075] In some possible situations, in order to make the question-answering effect more accurate, after obtaining the question to be processed, the question to be processed input by the user can be segmented, and the text of the question to be processed can be split into independent words; stop words, that is, those words that frequently appear in the text but have no practical meaning for understanding the question to be processed, can be removed; the text of the question to be processed can be cleaned and standardized, including removing irrelevant characters, unifying uppercase and lowercase, etc., to ensure the cleanliness of the question data and facilitate subsequent vectorization.

[0076] Furthermore, a model can be used to convert text into a numerical vector, where the model can be TF-IDF (Term Frequency-Inverse Document Frequency), Word Embedding, or BERT (Bidirectional Encoder Representations from Transformers), etc. It can also be a model that is fine-tuned for a specific field or specialized in question parsing or vector tasks.

[0077] S102, according to the problem to be processed, use the pre-trained model to perform text conversion to obtain an embedding vector corresponding to the problem to be processed;

[0078] The embodiment of the present invention does not limit the pre-trained model, and the user can choose according to actual needs. The pre-trained model is a machine learning model that has been trained on a large-scale data set, and the model can learn the potential representation or features of the text. Exemplarily, the pre-trained model can be a BGE model. The embedding vector of the problem to be processed refers to the conversion of the problem to be processed into a vector representation. The embedding vector is a high-dimensional numerical representation that can capture the semantic features of the problem to be processed, so that similar texts have similar representations in the vector space.

[0079] It can be understood that the pre-trained model has learned a lot of text features and semantic information, so it can understand the meaning of the text and convert it into an embedding vector; specifically, the problem to be processed is loaded with the pre-trained base model BGE, and the embedding vector corresponding to the problem to be processed is obtained while ensuring that the model can accurately convert the text into an embedding vector.

[0080] S103, calculating the similarity between the problem to be processed and the standard problem according to the embedding vector corresponding to the problem to be processed and the standard embedding vector corresponding to the standard problem in the standard problem library;

[0081] The standard question library includes standard embedding vectors corresponding to multiple standard questions and standard answers corresponding to multiple standard questions;

[0082] S104, determining a preset similarity threshold corresponding to the standard question;

[0083] S105, when the first preset similarity threshold corresponding to the standard question being an illegal question is less than the preset similarity threshold corresponding to the standard question being a legal question, determining a target standard question corresponding to the question to be processed and a target answer corresponding to the target standard question according to the similarity of the standard question and the preset similarity threshold corresponding to the standard question;

[0084] Among them, the similarity calculation method can use cosine similarity and Euclidean distance to quantify the similarity between the input question text vector and the vector stored in the question library.

[0085] For each embedded vector of the input question text to be processed, the similarity between the question text and the question database is evaluated by calculating its similarity with all vectors in the question database.

[0086] Determine the similarity threshold corresponding to the standard question. This similarity threshold can find a relatively balanced point between the question to be processed and the standard question in the standard question library. The similarity threshold can be set statically, which can more accurately find the standard question in the question library that is more similar to the question to be processed. The static setting but not fixed threshold, the user can adjust the similarity threshold by himself to adapt to different queries and question library updates. It is understandable that a preset similarity threshold can be set for each standard question, or a preset similarity threshold can be set for each type of standard question, and the user can set it according to the actual situation.

[0087] When the similarity between the user's pending question and a standard question is higher than the threshold set for the question, the match is considered successful, and the question with similarity higher than the threshold is used as a candidate answer. The similarity is sorted from high to low, and the target standard question with the highest similarity and its answer are selected as the reply.

[0088] It should be noted that among the multiple standard questions included in the standard database, there are legal questions and illegal questions. The specific illegal questions can be customized by the user and are not limited in the embodiment of the present invention.

[0089] In one possible scenario, when the preset similarity threshold corresponding to the standard question being an illegal question is lower than the preset similarity threshold corresponding to the standard question being a legal question, that is, the preset similarity threshold for illegal questions is set low in advance, at this time, the electronic device can more sensitively identify potential illegal questions. When only one question greater than the corresponding preset similarity threshold is determined based on the standard question and the corresponding preset similarity threshold, the question is used as the target standard question; if at least one question greater than the corresponding preset similarity threshold is determined, it is determined whether there is an illegal question in the at least one question, and if so, the question with the maximum similarity among the illegal questions is used as the target standard question.

[0090] At this time, through the similarity threshold of illegal questions, the electronic device can more sensitively identify potential illegal questions, thereby effectively preventing illegal content from being obtained and improving data security.

[0091] S106. If the target standard question is an illegal question, a prompt statement is generated and output; if the target standard question is not an illegal question, a target answer is output.

[0092] If the target standard question is an illegal question, a corresponding prompt statement is generated according to the illegal question and output through the user interface; if the target standard question is not an illegal question, the target answer is output.

[0093] It can be seen that in the embodiment of the present invention, the problem to be processed is obtained from the user input; the pre-trained model is used to perform text conversion, and the problem to be processed is converted into an embedding vector, so as to convert the natural language problem into a point in a high-dimensional space to obtain an embedding vector, which can capture the semantic information of the problem and understand the user's intention; by calculating the similarity between the embedding vector of the problem to be processed and the standard embedding vector of the standard problem in the standard question library, an accurate match of the problem to be processed is achieved; at the same time, different preset similarity thresholds are set to distinguish between illegal questions and legal questions, so that the preset similarity threshold of illegal questions is lower, and potential illegal questions can be identified more sensitively. Then, after determining the target standard question of the problem to be processed and its corresponding target answer, they are processed separately according to whether the target standard question is an illegal question, which greatly improves the efficiency and accuracy of the question and answer system, as well as the data security of the question and answer.

[0094] A possible implementation of the embodiment of the present invention is to determine the target standard question corresponding to the problem to be processed and the target answer corresponding to the target standard question according to the similarity of the standard question and the preset similarity threshold corresponding to the standard question, including:

[0095] Determine whether the similarity of the standard question is greater than a preset similarity threshold corresponding to the standard question;

[0096] The standard question with a similarity greater than a preset similarity threshold corresponding to the standard question is used as the first standard question;

[0097] According to the similarity, a target standard question is determined from the first standard question; and a target answer corresponding to the target standard question is retrieved from the standard question library.

[0098] In the embodiment of the present invention, a first standard question is obtained by pre-screening through a preset similarity threshold, and the first standard question is a candidate question. If there is one first standard question, the first standard question is used as the target standard question;

[0099] If the first standard questions are at least 2, then if the first standard questions include a plurality of first standard questions of different categories, determining whether the first standard questions include illegal questions;

[0100] If the first standard question of the illegal question is included, then the first standard question with the greatest similarity is selected from the first standard questions of the illegal question as the target standard question corresponding to the question to be processed;

[0101] If the first standard question of the illegal question is not included, then the first standard question with the greatest similarity is selected from the first standard questions according to the similarity to determine the target standard question.

[0102] In a possible implementation of the embodiment of the present invention, before calculating the similarity between the problem to be processed and the standard problem according to the embedding vector corresponding to the problem to be processed and the standard embedding vector corresponding to the standard problem in the standard problem library, the method further includes:

[0103] Obtain multiple standard questions and standard answers corresponding to the multiple standard questions;

[0104] Use the pre-trained model to convert multiple standard questions into text and obtain the embedding vectors corresponding to the standard questions;

[0105] The embedding vectors corresponding to the multiple standard questions and the standard answers corresponding to the multiple standard questions are stored in the standard question library.

[0106] In an embodiment of the present invention, data collection is first performed to collect and organize common problems in the industry as standard questions and their standard answers, and common problems in the industry are widely collected through various channels such as industry experts, technical literature, user feedback, etc.

[0107] Specifically, for each standard question collected, experts or technicians in related fields are organized to write accurate and standardized standard answers. Standard questions and standard answers are classified and sorted to ensure that the structure of the standard question library is clear and easy to search, and to ensure that the collected data is based on specific business needs and laws and regulations.

[0108] Before inputting the text of the sorted standard questions into the pre-trained model for embedding, data cleaning is required, including removing punctuation marks, stop words, text normalization (such as converting to lowercase, removing special characters, etc.), stem extraction or word form restoration, etc. Preprocessing helps to improve the quality of the embedding vector; then the standard questions and standard answers are stored in the database in the form of structured data to facilitate subsequent retrieval and matching.

[0109] Furthermore, before vector data is entered into the vector library, a strict quality review process is implemented to prevent low-quality data from negatively affecting system performance. According to business needs, the vector representation of illegal text is continuously updated to ensure that the system can effectively capture and block the latest illegal content.

[0110] Furthermore, during storage, classified storage can be performed, divided into legal categories and illegal categories. At the same time, the legal category can be further subdivided into multiple categories, illustratively, legal category 1, legal category 2, and legal category 3. The classification of categories can be performed by pre-defining labels for each standard question and classifying according to the labels, or by using a matching algorithm.

[0111] In one feasible method, a classification algorithm and a training data set are used to train a model to obtain a classification model, wherein the classification algorithm is such as Naive Bayes, Support Vector Machine (SVM), Decision Tree, Random Forest Text Classification Algorithm, and the training data set includes labeled questions and corresponding categories. Then, the trained classification algorithm is used to classify each standard question.

[0112] It is understandable that, since the number of standard questions in the standard database is large, it may take a long time to use the classification algorithm in sequence, so the data features corresponding to each of the multiple standard questions can be extracted;

[0113] Normalizing the data features corresponding to the multiple standard questions to obtain normalized features corresponding to the multiple standard questions;

[0114] Based on multiple standard questions, each corresponding to the normalized features is density clustered to obtain data sets of multiple categories;

[0115] The standard question of selecting the target number from each data set is input into the classification algorithm to obtain the corresponding category of each data set;

[0116] The category with the largest number of occurrences in the category corresponding to each data set is taken as the target category of the data set;

[0117] According to the categories corresponding to each of the multiple standard questions, the embedding vectors corresponding to each of the multiple standard questions, and the standard answers corresponding to each of the multiple standard questions, they are stored in a standard question library.

[0118] A possible implementation of the embodiment of the present invention further includes:

[0119] According to the user's question feedback information, update the answers to the standard questions corresponding to the question feedback information in the standard question library;

[0120] Regularly add embedding vectors corresponding to new standard questions and answers corresponding to new standard questions in the standard question library;

[0121] In an embodiment of the present invention, a user feedback mechanism is designed to allow users to evaluate automatically generated answers and obtain question feedback information. The question library, threshold or algorithm is adjusted according to user feedback to improve the accuracy of the question-answering system and user satisfaction. By providing a user feedback interface, the interface is integrated into the question-answering system, allowing users to evaluate the automatically generated answers in real time after receiving them. Users can rate the answers provided by the system with stars, select satisfaction, and provide answers corresponding to the questions. Users can express their satisfaction with the answers through star ratings and satisfaction selections, and these quantitative data provide an important basis for system optimization. Users have the right to submit answers that they think are correct, which not only helps to enrich the diversity of answers in the question library, but also provides valuable real data for the system.

[0122] For the user's question feedback information, if there are more than a first preset number of feedback information with revised answers for the same question and the similarity between them reaches a preset threshold, the answer to the standard question is adjusted.

[0123] At the same time, if during the actual question-and-answer process, a question asked by a user does not appear in the standard data, and the number of questions asked for the same question exceeds a second preset number, the technician can be notified to add the question and its corresponding answer.

[0124] A possible implementation of the present invention is shown in Figure 2 , after getting the pending questions input by the user, it also includes:

[0125] Matching the to-be-processed question with the historical search questions in the preset question library; the preset question library includes multiple historical search questions and answers corresponding to the multiple historical search questions;

[0126] If there is a successfully matched target history retrieval question, the answer corresponding to the target history retrieval question is determined from the preset question library, and the answer corresponding to the target history retrieval question is output;

[0127] If there is no successfully matched target historical retrieval question, the step of performing text conversion using a pre-trained model according to the question to be processed is executed to obtain an embedding vector corresponding to the question to be processed.

[0128] It is understandable that the answers fed back by calling the large model are relatively slow compared to directly querying the database. In order to avoid this problem, the user's questions are matched with the standard question library. If the match is successful, the question and its corresponding answer are stored in the database; when the same question appears again, the stored answer is directly retrieved and extracted from the database, thereby achieving a quick response, which can significantly improve the interaction efficiency and user experience of the question-answering system.

[0129] It can be seen that in the embodiment of the present invention, the verified user feedback question and answer pairs are stored in the database to ensure that the next time the user asks the same or similar question, the system can directly provide a verified answer. By continuously collecting and integrating user feedback, the question library of the present invention can achieve dynamic optimization and improve the accuracy of the answer and user satisfaction.

[0130] A possible implementation of the embodiment of the present invention further includes:

[0131] Obtain an initial training model and a training sample set, where the training sample set includes a plurality of training problem samples and training embedding vectors corresponding to the plurality of training problem samples;

[0132] Using the training sample set, the initial training model is trained to obtain a pre-training model;

[0133] According to a first total amount of illegal questions and a second total amount of legal questions of the training question samples in the training sample set, a preset similarity threshold corresponding to the illegal questions and a preset similarity threshold corresponding to the legal questions are determined.

[0134] In an embodiment of the present invention, the initial training model can be trained to obtain a pre-trained model. It can be understood that, in one possible case, the initial training model is an untrained model, and in another possible case, the initial training model is a machine learning model that has been trained on a large-scale dataset, which can learn the potential representation or features of the text.

[0135] Furthermore, a regular training plan can be implemented, such as retraining the vector generation model (such as the BGE model) every two or three months or based on the accumulation of new data. In this process, the latest compliant and non-compliant text samples are incorporated to enhance the model's ability to identify different text contents.

[0136] Specifically, the preset similarity threshold corresponding to illegal questions and the preset similarity threshold corresponding to legal questions can be determined based on a first correspondence between a preset range of illegal questions and a preset similarity threshold corresponding to illegal questions, and a second correspondence between a preset range of legal questions and a preset similarity threshold corresponding to legal questions; the above two correspondences are set by technical personnel based on experience, and it can be understood that the larger the number, the larger the threshold.

[0137] Of course, we can also adopt user feedback information, conduct in-depth analysis of misidentification cases, fine-tune the threshold manually or automatically, and continuously refine the decision boundary. According to the specific application scenario and the requirements for false positive and false negative rates, we can flexibly adjust the threshold to achieve a more accurate recognition and rejection effect.

[0138] A possible implementation of the embodiment of the present invention is to determine a preset similarity threshold corresponding to the standard question, including:

[0139] According to the user's question-answering demand strategy, determine the preset similarity threshold corresponding to the standard question;

[0140] When the second preset similarity threshold corresponding to the standard question being an illegal question is greater than the first preset similarity threshold, the target standard question and the target answer corresponding to the target standard question are determined according to the similarities of the various standard questions and the preset similarity thresholds corresponding to the various standard questions; and the target answer is output.

[0141] In the embodiment of the present invention, two demand strategies may be set for the user to choose, one is a security priority strategy, and the other is a user experience priority strategy.

[0142] For the security priority strategy, the preset similarity threshold corresponding to the standard question being an illegal question is lower than the preset similarity threshold corresponding to the standard question being a legal question, and the threshold corresponding to the strategy is directly selected; then steps S105 and S106 are executed; by lowering the similarity threshold, the system can more sensitively identify potential illegal questions, so as to maximize the identification of illegal questions and effectively prevent illegal content from being obtained.

[0143] For the user experience priority strategy, the preset similarity threshold of illegal questions is increased. Increasing the threshold can reduce the situation where legitimate questions are mistakenly judged as illegal, thereby reducing false alarms and improving answering efficiency.

[0144] Based on any of the above embodiments, the embodiment of the present invention provides a specific question-answering method, including:

[0145] Question text embedding: Use pre-trained models to convert questions that require access to large models into embedding vectors.

[0146] The collected question library is also embedded in the vector library: the standard question content related to each field is sorted and stored in a spreadsheet, encoded into a vector form through the same pre-trained model, and the embedding.pkl file is generated. It is then imported into a high-performance vector database management system to build an index.

[0147] First, build the vector library corresponding to the question and answer library. When building the question and answer vector library, you need to pay attention to the fact that these questions may change over time, so you need to update the embedding.pkl file regularly. This can be achieved through an automated process. You can write a scheduled task to regularly read new questions from a spreadsheet or other data source, and then use the pre-trained model to encode and update the vector library. Write the prepared question content into a spreadsheet containing the question and answer content (such as question.xlsx). This step is to organize all the questions (at that time, you can also divide them according to different fields. If you design separate spreadsheets for the health field, safety field, etc., it is convenient for daily maintenance). Use the same pre-trained model (BGE model) to encode each content in question.xlsx, convert it into vector form, and generate the embedding.pkl file, which encapsulates the semantic features of all questions and answers. In order to achieve fast and efficient vector retrieval, we can use ChromaDB, a high-performance vector database management system. First, we import the embedding.pkl file into ChromaDB. This file contains the embedding vectors we previously generated through the BGE model. Next, we will build an index for these vectors. This index will help ChromaDB quickly find the vectors most similar to the query vector.

[0148] ChromaDB uses optimized algorithms to ensure query efficiency and provide real-time responses even with large amounts of data. This means that no matter how large our data set is, ChromaDB can instantly find the most relevant results when a user asks a query. This enables fast retrieval of large amounts of data. Its optimized algorithms ensure query efficiency while supporting real-time responses, allowing the system to maintain high performance even with massive amounts of data, which greatly improves the user experience of our system, so we can provide fast and accurate search results.

[0149] Calculate the cosine similarity and compare it with the threshold: For each question content vector, calculate its cosine similarity with all vectors in the question library, and based on the set threshold, take the questions with similarity higher than the threshold as the answer to the reply.

[0150] For example, the user asks: What day of the week is today? The threshold is set to 0.8. The question is matched with the standard question library. According to the threshold matching, the matching result one is: Monday, and the matching result two is: Today is Monday and it is cloudy. The matching result with the highest matching degree is output.

[0151] The present invention integrates the advanced algorithms of big data processing capabilities and deep learning models, adopts a pre-trained vector model, converts the question text into a vector in a high-dimensional space, and performs in-depth semantic analysis on user queries, ensuring that the system can accurately capture the user's query intention and the deep semantic information of the question. By calculating the cosine similarity between the question vector and the vector in the question library and comparing it with the preset threshold, the illegal questions can be intelligently identified and effectively intercepted, the threats to the network ecological security can be reduced, and the harmony and order of the network environment can be maintained; through rapid response and accurate answers, the user's satisfaction when using the network question and answer service can be improved, ensuring that the user can obtain timely and useful information. In addition, the algorithm and model can be continuously optimized, the algorithm can be continuously studied and improved, and the deep learning model can be optimized to adapt to the ever-changing network environment and user needs. It can also provide customized answers according to the user's personalized needs and preferences, so that users can enjoy more personalized and flexible question and answer services. At the same time, the technical solution of the present invention is not limited to the field of network security, but is also applicable to multiple industries such as customer service, online education, and medical consultation that require intelligent question and answer functions, and has broad application prospects.

[0152] In summary, the present invention not only optimizes the performance of the question-and-answer system, improves the user experience, but also strengthens network security, and opens up new possibilities for technological innovation and industry applications. At the same time, due to the high accuracy of the present invention, it can reduce the situation of misjudgment and missed judgment, and further reduce the cost of post-processing. It has a positive impact on the performance, user experience, network security, and technological development of the network question-and-answer system, and provides strong support for building a more intelligent, secure, and efficient network question-and-answer environment.

[0153] The following is an introduction to a device provided by an embodiment of the present invention. The device described below and the method described above can be referred to each other. The device 200 of this embodiment is set in an electronic device, refer to, 3, Figure 3 : is a structural block diagram of a device according to one embodiment of the present invention, comprising:

[0154] The acquisition module 210 is used to acquire the pending question input by the user;

[0155] A text conversion module 220 is used to perform text conversion using a pre-trained model according to the problem to be processed to obtain an embedding vector corresponding to the problem to be processed;

[0156] A similarity calculation module 230 is used to calculate the similarity between the problem to be processed and the standard problem based on the embedding vector corresponding to the problem to be processed and the standard embedding vector corresponding to the standard problem in the standard problem library; the standard problem library includes standard embedding vectors corresponding to multiple standard problems and standard answers corresponding to the multiple standard problems;

[0157] A threshold determination module 240, for determining a preset similarity threshold corresponding to the standard question;

[0158] A matching module 250, configured to determine a target standard question corresponding to the problem to be processed and a target answer corresponding to the target standard question according to the similarity of the standard question and the preset similarity threshold corresponding to the standard question when the first preset similarity threshold corresponding to the standard question being an illegal question is less than the preset similarity threshold corresponding to the standard question being a legal question;

[0159] The answering module 260 is used to generate and output a prompt statement if the target standard question is an illegal question; and output a target answer if the target standard question is not an illegal question.

[0160] In a preferred example, the present invention can be further configured as: a matching module 250, for:

[0161] Determine whether the similarity of the standard question is greater than a preset similarity threshold corresponding to the standard question;

[0162] The standard question with a similarity greater than a preset similarity threshold corresponding to the standard question is used as the first standard question;

[0163] According to the similarity, a target standard question is determined from the first standard question; and a target answer corresponding to the target standard question is retrieved from the standard question library.

[0164] In a preferred example, the present invention can be further configured as: a matching module 250, for:

[0165] If the first standard question includes a plurality of first standard questions of different categories, determining whether the first standard question includes an illegal question;

[0166] If the first standard question of the illegal question is included, then the first standard question with the greatest similarity is selected from the first standard questions of the illegal question as the target standard question corresponding to the question to be processed;

[0167] If the first standard question of the illegal question is not included, then the first standard question with the greatest similarity is selected from the first standard questions as the target standard question corresponding to the question to be processed.

[0168] In a preferred example, the present invention can be further configured as follows:

[0169] The standard question library construction module is used to obtain multiple standard questions and the standard answers corresponding to each of the multiple standard questions; use the pre-trained model to convert the multiple standard questions into text to obtain the embedding vectors corresponding to the standard questions; and store the embedding vectors corresponding to each of the multiple standard questions and the standard answers corresponding to each of the multiple standard questions in the standard question library.

[0170] In a preferred example, the present invention can be further configured as: a standard question library construction module, further used for:

[0171] Extract the data features corresponding to multiple standard questions;

[0172] Normalizing the data features corresponding to the multiple standard questions to obtain normalized features corresponding to the multiple standard questions;

[0173] Based on multiple standard questions, each corresponding to the normalized features is density clustered to obtain data sets of multiple categories;

[0174] The standard question of selecting the target number from each data set is input into the classification algorithm to obtain the corresponding category of each data set;

[0175] The category with the largest number of occurrences in the category corresponding to each data set is taken as the target category of the data set;

[0176] According to the categories corresponding to each of the multiple standard questions, the embedding vectors corresponding to each of the multiple standard questions, and the standard answers corresponding to each of the multiple standard questions, they are stored in a standard question library.

[0177] In a preferred example, the present invention can be further configured as follows:

[0178] An updating module, used to update the answers to the standard questions corresponding to the question feedback information in the standard question library according to the user's question feedback information;

[0179] The update module is also used to regularly add embedding vectors corresponding to new standard questions and answers corresponding to new standard questions in the standard question library;

[0180] A quick retrieval module is used to match a question to be processed with a historical retrieval question in a preset question library; the preset question library includes multiple historical retrieval questions and answers corresponding to the multiple historical retrieval questions; if there is a successfully matched target historical retrieval question, the answer corresponding to the target historical retrieval question is determined from the preset question library, and the answer corresponding to the target historical retrieval question is output; if there is no successfully matched target historical retrieval question, the step of performing text conversion based on the question to be processed using a pre-trained model to obtain an embedding vector corresponding to the question to be processed is executed.

[0181] In a preferred example, the present invention can be further configured as follows:

[0182] A model training module is used to obtain an initial training model and a training sample set, wherein the training sample set includes multiple training problem samples and training embedding vectors corresponding to each of the multiple training problem samples; the initial training model is trained using the training sample set to obtain a pre-trained model; and a preset similarity threshold corresponding to the illegal questions and a preset similarity threshold corresponding to the legal questions are determined based on a first total amount of illegal questions and a second total amount of legal questions of the training problem samples in the training sample set.

[0183] An embodiment of the present invention provides an electronic device, such as Figure 4 As shown, Figure 4 The electronic device 300 shown includes: at least one processor 301 ( Figure 4 301 and memory 303. The processor 301 and memory 303 are connected, such as through a bus 302. Optionally, the electronic device 300 may further include a transceiver 304. It should be noted that in actual applications, the transceiver 304 is not limited to one, and the structure of the electronic device 300 does not constitute a limitation on the embodiments of the present invention.

[0184] The processor 301 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It may implement or execute various exemplary logic blocks, modules and circuits described in conjunction with the disclosure of the present invention. The processor 301 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0185] The bus 302 may include a path to transmit information between the above components. The bus 302 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus. The bus 302 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.

[0186] The memory 303 may be a ROM (Read Only Memory) or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory) or other types of dynamic storage devices that can store information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory) or other optical disk storage, optical disk storage (including compressed optical disk, laser disk, optical disk, digital versatile disk, Blu-ray disk, etc.), a magnetic disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.

[0187] The memory 303 is used to store application code for executing the solution of the present invention, and the execution is controlled by the processor 301. The processor 301 is used to execute the application code stored in the memory 303 to implement the content shown in the above method embodiment.

[0188] Figure 4 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present invention.

[0189] An embodiment of the present invention provides a computer-readable storage medium, in which at least one program code is stored. When the computer-readable storage medium is run on a computer, the computer can execute the corresponding content of the aforementioned method embodiment.

[0190] An embodiment of the present invention provides a computer program product, including a computer program or instructions, which implements the corresponding contents of the aforementioned method embodiment when the computer program or instructions are executed by a processor.

[0191] It should be understood that, although the steps in the flowchart of the accompanying drawings are displayed in sequence as indicated by the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least a part of the steps in the flowchart of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a part of the sub-steps or stages of other steps.

[0192] The above are only some embodiments of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. A question-answering method based on a large model, characterized in that: include: Get pending questions input by the user; According to the problem to be processed, use the pre-trained model to perform text conversion to obtain an embedding vector corresponding to the problem to be processed; Calculating the similarity between the problem to be processed and the standard problem according to the embedding vector corresponding to the problem to be processed and the standard embedding vector corresponding to the standard problem in the standard problem library; The standard question library includes a plurality of standard embedding vectors corresponding to the standard questions and a plurality of standard answers corresponding to the standard questions; Determining a preset similarity threshold corresponding to the standard question; When the first preset similarity threshold corresponding to the illegal question of the standard question is less than the preset similarity threshold corresponding to the legal question of the standard question, determining the target standard question corresponding to the question to be processed and the target answer corresponding to the target standard question according to the similarity of the standard question and the preset similarity threshold corresponding to the standard question; If the target standard question is an illegal question, a prompt statement is generated and output; if the target standard question is not an illegal question, the target answer is output.

2. The large model-based question-answering method according to claim 1, characterized in that: Determining a target standard question corresponding to the problem to be processed and a target answer corresponding to the target standard question according to the similarity of the standard question and a preset similarity threshold corresponding to the standard question, including: Determining whether the similarity of the standard question is greater than a preset similarity threshold corresponding to the standard question; The standard question with a similarity greater than a preset similarity threshold corresponding to the standard question is used as the first standard question; According to the similarity, the target standard question is determined from the first standard questions; and the target answer corresponding to the target standard question is retrieved from the standard question library.

3. The large model-based question-answering method according to claim 2, characterized in that: Determining the target standard question from the first standard questions according to the similarity includes: If the first standard question includes a plurality of first standard questions of different categories, determining whether the first standard question includes an illegal question; If the first standard question of the illegal question is included, then the first standard question with the greatest similarity is selected from the first standard questions of the illegal question as the target standard question corresponding to the question to be processed; If the first standard question of the illegal question is not included, then the first standard question with the greatest similarity is selected from the first standard questions as the target standard question corresponding to the question to be processed.

4. The large model-based question-answering method according to claim 1, characterized in that: Before calculating the similarity between the problem to be processed and the standard problem according to the embedding vector corresponding to the problem to be processed and the standard embedding vector corresponding to the standard problem in the standard problem library, the method further includes: Obtaining a plurality of the standard questions and standard answers corresponding to the plurality of the standard questions; Performing text conversion on the plurality of standard questions using the pre-trained model to obtain embedding vectors corresponding to the standard questions; The embedding vectors corresponding to each of the plurality of standard questions and the standard answers corresponding to each of the plurality of standard questions are stored in the standard question library.

5. The large model-based question-answering method according to claim 4, characterized in that: Storing the embedding vectors corresponding to the plurality of standard questions and the standard answers corresponding to the plurality of standard questions in the standard question library includes: Extracting data features corresponding to each of the plurality of standard questions; Normalizing the data features corresponding to each of the plurality of standard questions to obtain normalized features corresponding to each of the plurality of standard questions; Based on the normalized features corresponding to each of the plurality of standard questions, density clustering is performed to obtain a plurality of categories of data sets; The standard question of selecting the target number from each data set is input into the classification algorithm to obtain the corresponding category of each data set; The category with the largest number of occurrences in the category corresponding to each data set is taken as the target category of the data set; According to the categories corresponding to each of the multiple standard questions, the embedding vectors corresponding to each of the multiple standard questions, and the standard answers corresponding to each of the multiple standard questions, they are stored in the standard question library.

6. The large model-based question-answering method according to claim 4, characterized in that: Also includes: According to the user's question feedback information, the answers to the standard questions corresponding to the question feedback information in the standard question library are updated; and / or; Regularly adding embedding vectors corresponding to new standard questions and answers corresponding to new standard questions in the standard question library; and / or; After getting the pending questions from the user, it also includes: Matching the to-be-processed question with a historical search question in a preset question library; the preset question library includes a plurality of historical search questions and answers corresponding to the plurality of historical search questions; If there is a successfully matched target history retrieval question, determining the answer corresponding to the target history retrieval question from the preset question library, and outputting the answer corresponding to the target history retrieval question; If there is no successfully matched target historical retrieval question, a step is performed to convert text using a pre-trained model according to the problem to be processed to obtain an embedding vector corresponding to the problem to be processed.

7. The large model-based question-answering method according to any one of claims 1 to 6, characterized in that: Also includes: Obtaining an initial training model and a training sample set, wherein the training sample set includes a plurality of training problem samples and training embedding vectors corresponding to each of the plurality of training problem samples; Using the training sample set, training the initial training model to obtain a pre-training model; According to a first total amount of illegal questions and a second total amount of legal questions of the training question samples in the training sample set, a preset similarity threshold corresponding to the illegal questions and a preset similarity threshold corresponding to the legal questions are determined.

8. A question-answering device based on a large model, characterized in that: include: The acquisition module is used to obtain pending issues input by the user; A text conversion module, used to perform text conversion using a pre-trained model according to the problem to be processed, to obtain an embedding vector corresponding to the problem to be processed; A similarity calculation module, used to calculate the similarity between the problem to be processed and the standard problem according to the embedding vector corresponding to the problem to be processed and the standard embedding vector corresponding to the standard problem of the standard problem library; The standard question library includes a plurality of standard embedding vectors corresponding to the standard questions and a plurality of standard answers corresponding to the standard questions; A threshold determination module, used to determine a preset similarity threshold corresponding to the standard question; A matching module, configured to determine a target standard question corresponding to the problem to be processed and a target answer corresponding to the target standard question according to the similarity of the standard question and the preset similarity threshold corresponding to the standard question, when a first preset similarity threshold corresponding to the standard question being an illegal question is less than a preset similarity threshold corresponding to the standard question being a legal question; The reply module is used to generate a prompt statement and output the prompt statement if the target standard question is an illegal question; if the target standard question is not an illegal question, output the target answer.

9. An electronic device, characterized in that: The electronic device includes a memory and a processor, wherein a computer program is stored in the memory, and the processor executes the large model-based question-answering method according to any one of claims 1 to 7 when running the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one program code, and the program code is loaded and executed by the processor to implement the large model-based question-answering method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Question and answer pair generation method and device for function codes

    CN120470099A