A reasoning method, device, equipment, readable storage medium and program product
By constructing multi-choice questions and using multi-language large language model for inference analysis, the problem of low inference accuracy among different languages is solved, and higher accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510518324.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-24
AI Technical Summary
The existing multilingual large language model has low accuracy when dealing with inference tasks between different languages, mainly due to grammatical, context and cultural differences between languages.
By using the first language model to obtain the initial inference results of each language based on the target prompt template and the problem to be reasoned, a multi-choice question is constructed, and using the second largest language model for inference analysis, we finally obtain the comprehensive multi-language inference results.
It improves the accuracy of reasoning, reduces the limitations of the single language model, and enhances the robustness and accuracy of cross-language reasoning.
Smart Images

Figure CN120046741B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent reasoning technology, and in particular to a reasoning method, device, equipment, readable storage medium and program product. Background Art
[0002] Currently, the research on multilingual large language models mainly focuses on how to improve the model's understanding and reasoning abilities between different languages. When performing reasoning currently, a large language model is mainly used to reason about the problems to be reasoned in various languages to obtain reasoning results. However, this method usually faces challenges brought by grammar, context, and cultural differences between languages, resulting in the technical problem of low reasoning accuracy when the model processes complex reasoning tasks.
[0003] It can be seen that how to improve the accuracy of reasoning is a problem that needs to be solved by those skilled in the art. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a reasoning method, device, equipment and computer-readable storage medium, which solves the technical problem of low reasoning accuracy in the prior art.
[0005] To solve the above technical problem, the present invention provides a reasoning method, including:
[0006] Using a first large language model based on a target prompt template and a problem to be reasoned to obtain initial reasoning results corresponding to each language; wherein, the target prompt template is a template including the logic of the reasoning problem;
[0007] Based on the initial reasoning results corresponding to each language and a multiple-choice question prompt template, a multiple-choice question is obtained; wherein, the multiple-choice question prompt template is a template including the logic of constructing a multiple-choice question based on the initial reasoning results;
[0008] Performing reasoning analysis on the multiple-choice question using a second large language model to obtain the final reasoning result.
[0009] On the one hand, using a first large language model based on a target prompt template and a problem to be reasoned to obtain initial reasoning results corresponding to each language, including:
[0010] Obtaining the problem to be reasoned and the target language corresponding to the problem to be reasoned;
[0011] Using a target language prompt template based on the problem to be reasoned and the target language to obtain first input information;
[0012] Using a language prompt template corresponding to the remaining languages based on the problem to be reasoned and the remaining languages; the target prompt template includes the target language prompt template and the language prompt template;
[0013] Based on the first input information and the second input information, use the first large language model to obtain the initial inference results corresponding to each language.
[0014] On the one hand, based on the initial inference results corresponding to each language and the multiple-choice question prompt template, obtain multiple-choice questions, including:
[0015] Based on the initial inference results corresponding to each language, obtain the inference process and the initial inference results corresponding to each language;
[0016] Use the multiple-choice question target, and use the inference process and the initial inference results corresponding to each language as options to construct the multiple-choice questions.
[0017] On the other hand, after performing inference analysis on the multiple-choice questions using the second large language model to obtain the final inference results, it further includes:
[0018] Determine the target language corresponding to the question to be inferred;
[0019] Determine whether the final inference results include the inference results corresponding to the target language;
[0020] When the final inference results include the inference results corresponding to the target language, send the inference results corresponding to the target language so that the language of the feedback inference results is consistent with the target language.
[0021] On the other hand, after determining whether the final inference results include the inference results corresponding to the target language, it further includes:
[0022] When the final inference results do not include the inference results corresponding to the target language, use the translation model to translate the final inference results based on the target language to obtain the translated inference results;
[0023] Send the translated inference results.
[0024] On the other hand, after using the translation model to translate the final inference results based on the target language to obtain the translated inference results when the final inference results do not include the inference results corresponding to the target language, it further includes:
[0025] Perform semantic detection on the translated inference results to determine whether the semantics is accurate;
[0026] When it is accurate, send the translated inference results;
[0027] When it is inaccurate, perform reasoning analysis based on the final reasoning result and the translated reasoning result to obtain the corrected translated reasoning result.
[0028] On the one hand, after performing reasoning analysis using the second large language model based on the multiple-choice question to obtain the final reasoning result, it further includes:
[0029] Determine the confidence threshold corresponding to the final reasoning result;
[0030] On the one hand, the training process of the first large language model includes:
[0031] Train the large language model using a multilingual dataset to obtain an initial large language model; wherein, the data in the multilingual dataset includes news, encyclopedia, papers, and social media texts;
[0032] Fine-tune the initial large language model using a task question set to obtain the first large language model.
[0033] On the one hand, the training process of the second large language model includes:
[0034] Obtain a multilingual multiple-choice question dataset; wherein, the multilingual multiple-choice question dataset includes distractors, and the distractors include multimodal interference data, and the multimodal interference data includes mixed unit interference and mixed semantic interference;
[0035] Train the multilingual pre-trained model using the multilingual multiple-choice question dataset to obtain the second large language model.
[0036] On the one hand, the multiple-choice question prompt template includes question information about the options and each option, and each option is composed of the initial reasoning result.
[0037] On the one hand, the question to be reasoned includes at least one of an intelligent education question and a medical diagnosis question.
[0038] An embodiment of the present invention further provides an inference device, including:
[0039] An initial inference module, configured to use the first large language model based on a target prompt template and a question to be reasoned to obtain an initial reasoning result corresponding to each language; wherein, the target prompt template is a template including the logic of the reasoning question;
[0040] A multiple-choice question determination module, configured to obtain a multiple-choice question based on the initial reasoning result corresponding to each language and a multiple-choice question prompt template; wherein, the multiple-choice question prompt template is a template including the logic of constructing a multiple-choice question based on the initial reasoning result;
[0041] A final reasoning module for performing reasoning and analysis on the multiple-choice question using a second large language model to obtain a final reasoning result.
[0042] An embodiment of the present invention further provides a reasoning device, including:
[0043] A memory for storing a computer program;
[0044] A processor for executing the computer program to implement the steps of the above-mentioned reasoning method.
[0045] An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned reasoning method are implemented.
[0046] An embodiment of the present invention further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above-mentioned reasoning method are implemented.
[0047] To solve the above technical problems, an embodiment of the present invention provides a reasoning method, including: obtaining an initial reasoning result corresponding to each language by using a first large language model based on a target prompt template and a question to be reasoned; wherein, the target prompt template is a template including the logic of the reasoning question; obtaining a multiple-choice question based on the initial reasoning result corresponding to each language and a multiple-choice question prompt template; wherein, the multiple-choice question prompt template is a template including the logic of constructing a multiple-choice question based on the initial reasoning result; performing reasoning and analysis on the multiple-choice question by using a second large language model to obtain a final reasoning result.
[0048] It can be seen from the above technical solutions that the beneficial effect of the present invention is that: compared with the current reasoning process that depends on the prompt words of a certain specific language, resulting in low reasoning accuracy, the present invention constructs a multiple-choice question based on the reasoning results corresponding to each language, so that context analysis can be performed based on the reasoning results of multiple languages to obtain a final reasoning result. Since the final reasoning result can integrate the reasoning of multiple languages, the limitations of single-language model reasoning can be reduced, so the accuracy of reasoning can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] To more clearly illustrate the embodiments of the present invention, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0050] Figure 1 It is a flowchart of a reasoning method provided by an embodiment of the present invention;
[0051] Figure 2 A flowchart example of a first large language model training method provided by an embodiment of the present invention;
[0052] Figure 3 A flowchart example of an inference method provided by an embodiment of the present invention;
[0053] Figure 4 A structural schematic diagram of an inference device provided by an embodiment of the present invention;
[0054] Figure 5 A structural schematic diagram of an inference device provided by an embodiment of the present invention. Detailed implementation manners
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] The terms "including" and "having" in the specification of the present invention and the accompanying drawings above, and any variations related to "including" and "having", are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may include unlisted steps or units.
[0057] To enable those skilled in the art of the present technology to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0058] Next, an inference method provided by an embodiment of the present invention will be introduced in detail. Figure 1 A flowchart of an inference method provided by an embodiment of the present invention, the method may include:
[0059] S101, using a first large language model based on a target prompt template and a problem to be inferred, to obtain an initial inference result corresponding to each language; wherein, the target prompt template is a template including the logic of the inference problem.
[0060] The execution subject of this embodiment is an electronic device. This embodiment does not limit the specific electronic device. For example, the electronic device in this embodiment can be a computer, a mobile phone, etc. The problem to be inferred in this embodiment can include at least one of intelligent education problems and medical diagnosis problems, that is, this solution can be applied to the fields of intelligent education technology and medical diagnosis technology. In addition, it can also be applied to problems in other fields, such as the field of mathematics, the field of aerospace, etc. The target prompt template of this embodiment can include: the target to be inferred and the problem to be inferred. For example, it can be "According to the following question, give the correct answer in Chinese: If a cat's speed is 10 kilometers per hour, then how much time does it take for it to run 15 kilometers?" or "Based on the following question, provide the correct answer in English: If a cat's speed is 10 kilometers per hour, then how much time does it take for it to run 15 kilometers?". This embodiment does not limit the specific solution for obtaining the inference results corresponding to each language. For example, this embodiment can use the same target prompt template for each language; or it can use different target prompt templates for each language, where the form of each prompt template can be the inference target and the problem to be inferred in the corresponding language form, or it can be the inference target in the corresponding language form (the inference target is consistent with the current language) and the problem to be inferred in the corresponding language form (referring to translating the problem to be inferred into the same language as the current language). This embodiment does not limit the specific first large language model. For example, the first large language model in this embodiment can be mBERT (Multilingual Model BERT), XLM-R (Cross-lingual Model - RoBERTa), etc. The multilingual dataset used by the first large language model in this embodiment includes but is not limited to large-scale corpora such as news, encyclopedias, and social media texts in various languages (the multilingual dataset is a corpus composed of texts in multiple languages, covering natural language data in different languages (such as English, Chinese, Spanish, etc.). Through the training of multilingual corpora, the model can master the basic grammar structures, common vocabulary, and semantic relationships between languages, providing support for cross-lingual reasoning and problem understanding. As an alternative solution, an existing multilingual open-source pre-trained model can also be directly used. After completing the multilingual pre-training, the present invention further fine-tunes and trains the large language model by using a multilingual downstream task question set (various reasoning questions). Through this fine-tuning method, the present invention enables the model to not only be general in terms of language but also show high accuracy in specific downstream task fields.
[0061] It should be further noted that, based on any of the above embodiments, in order to avoid language bias, obtaining the initial inference results corresponding to each language by using the first large language model based on the target prompt template and the problem to be inferred may include:
[0062] S1011, Obtain the problem to be inferred and the target language corresponding to the problem to be inferred;
[0063] S1012, Based on the problem to be inferred and the target language, use the target language prompt template to obtain the first input information;
[0064] S1013, Based on the problem to be inferred and the remaining languages, use the language prompt templates corresponding to the remaining languages to obtain the second input information; The target prompt templates include the target language prompt template and the language prompt templates.
[0065] S1014, Based on the first input information and the second input information, use the first large language model to obtain the initial inference results corresponding to each language.
[0066] This embodiment is different from the traditional multi-language inference method. In the present invention, the problem to be inferred input by the user is not forced to be converted into a form consistent with each language, but the original language of the problem to be inferred is used, thus avoiding possible translation errors or semantic deviations. For example, first identify the language of the user input problem and pass the problem to the string variable question. For example, the problem input by the user is: "If a cat's speed is 10 kilometers per hour, how much time does it take for it to run 15 kilometers?", then the system analyzes the input content through the language recognition module, identifies that the language of this problem is Chinese, and question = "If a cat's speed is 10 kilometers per hour, how much time does it take for it to run 15 kilometers?". By looping through multiple target languages, the problem is passed to the corresponding target language prompt templates respectively to obtain the complete model input, generate the inference results in the target language version, and save them in the candidate result answers list. For example, the system may select the following target languages (depending on which languages have been used in the pre-training and fine-tuning stages of the model): Chinese, English, and call the prompt templates specific to each language version. For example, the Chinese prompt template is: "According to the following question, give the correct answer in Chinese: {question}"; the English prompt template is: "Based on the following question, provide the correct answer in English: {question}", and the complete inputs obtained are "According to the following question, give the correct answer in Chinese: If a cat's speed is 10 kilometers per hour, how much time does it take for it to run 15 kilometers?" and "Based on the following question, provide the correct answer in English: If a cat's speed is 10 kilometers per hour, how much time does it take for it to run 15 kilometers?".
[0067] It should be further explained that, based on any of the above embodiments, the training process of the above first large language model may include: using a multilingual data set to train the large language model to obtain an initial large language model; wherein the data in the multilingual data set includes news, encyclopedias, papers, and social media texts; using the task problem set to fine-tune the initial large language model to obtain the first large language model. This embodiment will use data from various fields to improve the coverage of knowledge, thereby improving the accuracy and comprehensiveness of the second large language model's reasoning.
[0068] S102, obtaining multiple-choice questions based on the initial reasoning results corresponding to each language and the multiple-choice question prompt template; wherein the multiple-choice question prompt template is a template including a logic for constructing the multiple-choice questions based on the initial reasoning results.
[0069] The multiple-choice prompt template in this embodiment is a template for constructing the logic of multiple-choice questions based on the initial reasoning results. This embodiment does not limit the specific method of constructing multiple-choice questions. For example, this embodiment can take the problem to be inferred as the question, and each initial reasoning result constitutes an option of the question; or this embodiment can combine the original question and the candidate answers in all languages into a multiple-choice question format, and the reasoning process and answer order of each language correspond to each option to construct the final multiple-choice question. The multiple-choice prompt template in this embodiment may include question information about the options and each option, and each option is composed of the initial reasoning results. For example, the prompt word template of the multiple-choice question may be "{question} Please select the correct answer, option A (Chinese): {answers[0]}, option B (English): {answers[1]}". Assume that the candidate result list composed of the initial inference results corresponding to each language is answers=["The time it takes for the cat to run 15 kilometers is 1.5 hours", "The time it takes for the cat to run 15 kilometers is 1.5 hours"], then the constructed multiple-choice question is "If a cat's speed is 10 kilometers per hour, how long does it take to run 15 kilometers? Please choose the correct answer, option A (Chinese): The time it takes for the cat to run 15 kilometers is 1.5 hours, option B (English): The time it takes for the cat to run15 kilometers is 1.5 hours".
[0070] It should be further explained that, in order to improve the accuracy of multiple-choice question construction, the multiple-choice questions obtained based on the initial reasoning results corresponding to each language and the multiple-choice question prompt template may include:
[0071] S1021, obtain the inference processes and initial inference results corresponding to each language based on the initial inference results corresponding to each language;
[0072] S1022, use the multiple-choice question objective, and construct a multiple-choice question with the inference processes and initial inference results corresponding to each language as options.
[0073] In this embodiment, when constructing the multiple-choice question, not only the initial inference results are used, but also the inference process corresponding to each language is used, so that when using the second large language model for inference analysis, the inference process can be combined to improve the accuracy of multi-language context analysis.
[0074] S103, perform inference analysis using the second large language model based on the multiple-choice question to obtain the final inference result.
[0075] The second large language model in this embodiment is a model that can perform inference based on multiple-choice questions. The second large language model in this embodiment can be GPT; or it can also be Doubao, etc. The second large language model in this embodiment can be trained using a multi-language corpus and specifically trained with training data for multiple-choice question inference. When performing inference analysis in this embodiment, each language version can be analyzed independently first to generate the confidence of each option (for example, the Chinese version supports option A with a confidence of 85%). Compare the derivation results of each language, mark the conflict points, and conduct a specific analysis of the conflicting parts to obtain a more accurate answer. For example, the second large language model synthesizes the information of each language version to automatically select the final answer. For example, in the above example, whether it is Chinese or English, the final model will understand that 1.5 hours is the time required for the cat to run 15 kilometers, and the answers of each language version are the same. Therefore, the final answers are answers[0] and answers[1].
[0076] It should be further noted that, based on any of the above embodiments, the training process of the above second large language model may include: obtaining a multilingual multiple-choice question dataset; wherein, the multilingual multiple-choice question dataset includes distraction options, and the distraction options include multimodal distraction data, and the multimodal distraction data includes mixed unit distraction and mixed semantic distraction; using the multilingual multiple-choice question dataset to train the multilingual pre-trained model to obtain the second large language model. This embodiment takes into account that the grammar and word order differences in different languages may affect the accuracy of the model, so it uses distraction options for training. In multimodal machine learning tasks (such as combining data of different modalities such as images, texts, and voices), distraction options refer to artificially introduced or naturally existing distracting inputs, and the purpose is to disrupt the model's ability to understand or associate multimodal information. During the reasoning process, the data in different modalities are inconsistent in units, formats, or dimensions, resulting in difficulties for the model to align or fuse multimodal information. The mixed unit distraction in this embodiment refers to the cross-modal alignment problem caused by inconsistent data formats, units, or dimensions, mainly the interference at the physical level (units, formats, dimensions). The mixed semantic distraction in this embodiment refers to the logical understanding error caused by cross-modal semantic contradictions or irrelevancies. This embodiment trains the second large language model based on multimodal distraction data, enabling the model to reason more accurately.
[0077] It should be further noted that, based on any of the above embodiments, after using the second large language model to perform reasoning and analysis based on multiple-choice questions and obtaining the final reasoning result, it may further include:
[0078] S1: Determine the target language corresponding to the question to be reasoned;
[0079] S2: Determine whether the final reasoning result includes the reasoning result corresponding to the target language;
[0080] S3: When the final reasoning result includes the reasoning result corresponding to the target language, send the reasoning result corresponding to the target language so that the language of the feedback reasoning result is consistent with the target language.
[0081] This embodiment will determine whether the final reasoning result contains the language version in the target language corresponding to the question to be reasoned. If so, the answer of the language version will be returned to the user. Otherwise, the final answer will be translated into the language of the question to be reasoned and then returned to the user. Assuming that the large language model selects the Chinese version and the English version of the answer from the multiple-choice question, since the user's original question is in Chinese, the system will directly return the Chinese version of the answer to the user. If the large language model only selects the English version of the answer from the multiple-choice question, the system will call the translation module at this time, translate the English answer into Chinese, and finally return it to the user. In practical applications, this translation mechanism can ensure that users can receive answers consistent with the language of their questions, while providing high-quality multilingual reasoning capabilities. This embodiment does not limit the specific method of translation. For example, this embodiment can determine the specific technical field of the problem to be reasoned based on the problem to be reasoned, thereby utilizing a dedicated translator in this field to improve the accuracy of the translation.
[0082] It should be further explained that, based on any of the above embodiments, after determining whether the final reasoning result includes the reasoning result corresponding to the target language, it may also include: when the final reasoning result does not include the reasoning result corresponding to the target language, translating the final reasoning result based on the target language using a translation model to obtain a translated reasoning result; and sending the translated reasoning result. This embodiment does not limit the specific translation model. For example, the translation model in this embodiment can be a cross-language reasoning result conversion method based on a logical dependency tree, wherein a logical dependency tree (LDT, Logical Dependency Tree) and a text synchronization conversion mechanism are constructed, and the predicate logical relationship of the original reasoning is retained during the translation process (such as , quantifier, → implied relationship); introduce cognitive symbol encoder (CSE), unify the encoding of text, mathematical symbols, and chart elements, support cross-language conversion of non-textual reasoning elements such as formulas and flowcharts, develop domain-aware routing network (DARN), call professional terminology library in real time, and build in 500+ vertical domain knowledge packages (medicine / law / engineering, etc.). The logic integrity verification module uses a dual-channel verification mechanism: Channel 1: original reasoning result → translation → back translation back to source language → logical equivalence verification; Channel 2: directly compare the similarity of the logical dependency tree structure before and after translation. In this process, a bilingual comparison logic report can be generated to mark the conversion process of key reasoning nodes. This model can improve the accuracy of translation. Alternatively, this embodiment can also construct a translation model corresponding to each field of the problem to be reasoned.
[0083] It should be further noted that, in order to improve the accuracy of the inference result, when the final inference result does not include the inference result corresponding to the target language, the final inference result is translated using a translation model based on the target language. After obtaining the translated inference result, it may further include: performing semantic detection on the translated inference result to determine whether the semantics is accurate; when it is accurate, sending the translated inference result; when it is inaccurate, performing inference analysis based on the final inference result and the translated inference result to obtain a corrected translated inference result. This embodiment performs semantic detection after obtaining the translated inference result, improving the accuracy of translation.
[0084] It should be further noted that, based on any of the above embodiments, after performing inference analysis using the second large language model based on a multiple-choice question to obtain the final inference result, it may further include: determining the confidence threshold corresponding to the final inference result; when the confidence threshold is lower than the set confidence threshold, sending a prompt message. This embodiment determines the confidence threshold corresponding to the final inference result, so that when the confidence threshold is low, the user can participate in manual inference.
[0085] An inference method provided by an embodiment of the present invention may include: S101, using a first large language model based on a target prompt template and a question to be inferred to obtain an initial inference result corresponding to each language; where the target prompt template is a template including the logic of the inference question; S102, obtaining a multiple-choice question based on the initial inference result corresponding to each language and a multiple-choice question prompt template; where the multiple-choice question prompt template is a template including the logic of constructing a multiple-choice question based on the initial inference result; S103, performing inference analysis using a second large language model based on the multiple-choice question to obtain the final inference result. Compared with the current inference process relying on prompt words in a certain specific language, resulting in low inference accuracy, the present invention constructs a multiple-choice question based on the inference results corresponding to each language, enabling context analysis based on the inference results in multiple languages to obtain the final inference result. Since the final inference result can synthesize the inferences in multiple languages, it can reduce the limitations of single language model inference, so the accuracy of inference can be improved.
[0086] For a better understanding of the present invention, please specifically refer to Figure 2 , Figure 2 which is a flowchart example of a training method for a first large language model provided by an embodiment of the present invention, and may specifically include:
[0087] S201. Pre-train a large language model using a multilingual dataset to obtain a pre-trained inference model; where the multilingual dataset is a large-scale corpus including news, encyclopedias, and social media in various languages.
[0088] The pre-trained inference model in this embodiment is a model that can perform inference based on questions to obtain inference results. In this embodiment, the large language model is pre-trained on a multilingual dataset, which includes but is not limited to large-scale corpora such as news, encyclopedias, and social media texts in various languages. Through the training of the multilingual dataset, the model can master the basic grammar structures, common vocabulary, and semantic relationships between languages, providing support for cross-lingual inference and question understanding. As an alternative, existing open-source multilingual pre-trained models can also be directly used, such as directly using existing open-source multilingual pre-trained models like mBERT (Multilingual Model BERT), XLM-R (Cross-lingual Model - RoBERTa), etc.
[0089] S202, fine-tune the pre-trained inference model using multilingual downstream task questions to obtain the first large language model.
[0090] This embodiment uses a dataset composed of multilingual downstream task questions to fine-tune and train the large language model. After completing the pre-training, the present invention further fine-tunes and trains the large language model by using multilingual downstream task questions. Through this fine-tuning method, the present invention enables the model to not only be general in terms of language but also show high accuracy in specific downstream task fields.
[0091] The main purpose of the present invention is to provide a large language model inference method based on multilingual context, which creatively introduces self-generated multilingual context into the large language model.
[0092] Compared with the prior art, the advantage of the present invention is that during the inference process, the inference results of each language are concatenated into a multiple-choice question form, which can integrate the inference advantages of different languages and improve the inference accuracy and reliability in a multilingual environment. This method effectively solves the problems of difficult handling of language differences and inference results depending on a single input language in existing multilingual inference models, thereby improving the inference accuracy and stability of the model.
[0093] For a better understanding of the present invention, please specifically refer to Figure 3 , Figure 3 which is a flow example diagram of an inference method provided by an embodiment of the present invention, and specifically may include:
[0094] S301, obtain the question to be inferred, based on the target prompt template for each language, obtain the input information corresponding to each language, and combine the input information to obtain the complete input.
[0095] S302, perform inference using the first large language model based on the complete input to obtain the inference process and initial inference results for each language.
[0096] In this embodiment, in order to obtain a multilingual context, the system will generate corresponding initial inference results for each language. For each language, candidate answers for the corresponding language are generated through the following steps: 1) For each language (such as Chinese, English, French, etc.), the system substitutes the string of the problem to be inferred in the original input into the target prompt templates of each language to obtain the complete input for the large model; 2) For each language, the model performs inference based on the domain knowledge (such as math problem sets, logical reasoning questions, etc.) trained in the fine-tuning stage to obtain the generation results for each target language, including the inference process and inference results, etc.
[0097] S303, perform inference using the multiple-choice question prompt template based on the inference questions and initial inference results for each language to obtain multiple-choice questions; wherein, the inference process and inference results for each language are used as one option.
[0098] This embodiment splices the generated candidate answers (initial inference results) to obtain a multilingual context: the inference process and initial inference results for each language are used as one option. In the form of multiple-choice questions, the system asks the large language model to select the most correct ones from multiple candidate answers. This selection depends not only on the inference results of a single language, but also comprehensively considers the multilingual inference context.
[0099] S304, based on the multiple-choice questions, use the second large language model to select the most correct inference result from multiple options;
[0100] S305, determine whether the language of the most correct inference result is the same as the language of the problem to be inferred.
[0101] S306, if they are the same, return the most correct inference result as the target inference result to the user.
[0102] S307, if they are not the same, translate the most correct inference result to make it consistent with the language of the problem to be inferred to obtain the target inference result, and return the target inference result to the user.
[0103] The technical solution of the present invention effectively combines the inference results of multiple language models through an innovative multi - language inference framework, thereby providing a cross - language inference solution. The main features of this technical solution include: (1) This technical solution can automatically identify the language of the user's question (the question to be inferred) and generate inference answers (inference results) in multiple language versions. This multi - language support enables users to obtain accurate inference results regardless of the language they use to ask questions. (2) The inference results in different languages are combined in the form of multiple - choice questions, and the large - language model synthesizes the information of each language version to select the final answer. This method can integrate the inference processes of different languages, thereby reducing the limitations of single - language model inference. (3) After selecting the final answer in the multiple - choice question, the system will determine whether the answer is consistent with the language of the user's original question. If not, the final answer will be translated back to the user's language to ensure that the user obtains an answer that matches the language of their question.
[0104] Traditional multi - language inference usually relies on translating the question into a single language for inference. The inference accuracy of this method is easily affected by the single - language ability of the model and cannot fully utilize the advantages of multi - language models. Many existing multi - language models only support basic translation or simple multi - language tasks and lack the integrated use of the cross - language inference ability of the model.
[0105] Compared with existing methods, the present invention has the following advantages:
[0106] 1. By using the inference results in multiple language versions simultaneously, the present invention not only enhances the robustness of inference but also effectively avoids the error propagation problem in single - language inference.
[0107] 2. Using the inference results of multi - language inference to form multiple - choice questions allows the model to synthesize information from different languages and make a final choice. This method significantly improves the inference accuracy compared with the existing technology. By integrating multi - language context information, the system can fully consider the differences of each language in inference and provide more comprehensive and multi - perspective answers. This multi - dimensional inference method has not been applied in the existing technology.
[0108] Next, the inference device provided by the embodiments of the present invention will be introduced. The inference device described below can be correspondingly referred to the inference method described above.
[0109] Figure 4 A structural schematic diagram of an inference device provided by an embodiment of the present invention may include:
[0110] An initial inference module 100, configured to use a first large - language model based on a target prompt template and a question to be inferred to obtain initial inference results corresponding to each language; wherein, the target prompt template is a template including the logic of the inference question;
[0111] The multiple-choice question determination module 200 is used to obtain multiple-choice questions based on the initial inference results corresponding to each language and the multiple-choice question prompt template; wherein, the multiple-choice question prompt template is a template including constructing the logic of multiple-choice questions based on the initial inference results.
[0112] The final inference module 300 is used to perform inference analysis on the multiple-choice questions using the second large language model to obtain the final inference result.
[0113] Furthermore, based on the above embodiments, the initial inference module 100 may include:
[0114] The target language determination module is used to obtain the question to be inferred and the target language corresponding to the question to be inferred.
[0115] The first input information determination unit is used to obtain the first input information based on the question to be inferred and the target language using the target language prompt template.
[0116] The second input information determination unit is used to obtain the second input information based on the question to be inferred and the remaining languages using the language prompt units corresponding to the remaining languages; the target prompt template includes the target language prompt template and the language prompt template.
[0117] The initial inference result determination unit is used to obtain the initial inference results corresponding to each language based on the first input information and the second input information using the first large language model.
[0118] Furthermore, based on any of the above embodiments, the multiple-choice question determination module 200 may include:
[0119] The inference process and result determination unit is used to obtain the inference process and the initial inference results corresponding to each language based on the initial inference results corresponding to each language.
[0120] The multiple-choice question construction unit is used to construct the multiple-choice question by using the multiple-choice question target and taking the inference process and the initial inference results corresponding to each language as options.
[0121] Furthermore, based on the above embodiments, the above inference device may further include:
[0122] The target language determination module is used to determine the target language corresponding to the question to be inferred.
[0123] The judgment module is used to determine whether the final inference result includes the inference result corresponding to the target language.
[0124] Inference result range module, which is used to send the inference result corresponding to the target language when the final inference result includes the inference result corresponding to the target language, so that the language of the feedback inference result is consistent with the target language.
[0125] Further, based on the above embodiment, the above inference device may further include:
[0126] Translation module, which is used to translate the final inference result using a translation model based on the target language to obtain a translated inference result when the final inference result does not include the inference result corresponding to the target language;
[0127] Sending module, which is used to send the translated inference result.
[0128] Further, based on the above embodiment, the above inference device may further include:
[0129] Semantic detection module, which is used to detect the semantics of the translated inference result to determine whether the semantics is accurate;
[0130] Translated inference result sending module, which is used to send the translated inference result when it is accurate;
[0131] Correction module, which is used to perform inference analysis based on the final inference result and the translated inference result to obtain a corrected translated inference result when it is inaccurate.
[0132] Further, based on the above embodiment, the above inference device may further include:
[0133] Confidence threshold determination module, which is used to determine the confidence threshold corresponding to the final inference result;
[0134] Prompt information sending module, which is used to send prompt information when the confidence threshold is lower than the set confidence threshold.
[0135] Further, based on the above embodiment, the above inference device may further include:
[0136] Initial large language model determination module, which is used to train a large language model using a multilingual dataset to obtain an initial large language model; wherein, the data in the multilingual dataset includes news, encyclopedias, papers, and social media texts;
[0137] Second large language model training module, which is used to fine-tune the initial large language model using a task question set to obtain the first large language model.
[0138] Further, based on the above embodiment, the above inference device may further include:
[0139] A multilingual multiple-choice question dataset determination module for obtaining a multilingual multiple-choice question dataset; wherein, the multilingual multiple-choice question dataset includes distractors, and the distractors include multimodal interference data, and the multimodal interference data includes mixed unit interference and mixed semantic interference;
[0140] A second large language model training module for training a multilingual pre-trained model using the multilingual multiple-choice question dataset to obtain the second large language model.
[0141] Further, based on any of the above embodiments, the multiple-choice question prompt template includes question information about the options and each option, and each option is composed of the initial inference result.
[0142] Further, based on any of the above embodiments, the question to be inferred includes at least one of intelligent education questions and medical diagnosis questions.
[0143] It should be noted that the order of the modules and units in the above inference device can be changed before and after without affecting the logic.
[0144] Figure 4 The description of the features in the corresponding embodiments can be referred to Figure 4 the relevant descriptions of the corresponding embodiments, which will not be elaborated here one by one.
[0145] The inference device provided by the embodiment of the present invention may include: an initial inference module 100 for obtaining an initial inference result corresponding to each language by using a first large language model based on a target prompt template and a question to be inferred; wherein, the target prompt template is a template including the logic of the inference question; a multiple-choice question determination module 200 for obtaining a multiple-choice question based on the initial inference result corresponding to each language and a multiple-choice question prompt template; wherein, the multiple-choice question prompt template is a template including the logic of constructing a multiple-choice question based on the initial inference result; a final inference module 300 for performing inference analysis on the multiple-choice question by using a second large language model to obtain a final inference result. Compared with the current inference process relying on the prompt words of a certain specific language, resulting in low inference accuracy, the present invention constructs a multiple-choice question based on the inference results corresponding to each language, so that context analysis can be performed based on the inference results of multiple languages to obtain the final inference result. Since the final inference result can synthesize the inferences of multiple languages, the limitations of single language model inference can be reduced, so the inference accuracy can be improved.
[0146] Next, an inference device provided by an embodiment of the present invention is introduced, and the inference device described below can be mutually referred to the inference method described above.
[0147] Figure 5The following is a schematic structural diagram of an inference device provided by an embodiment of the present invention. As Figure 5 shown, the inference device includes: a memory 60 for storing computer programs;
[0148] a processor 61 for implementing the steps of the inference method in the above embodiment when executing the computer program.
[0149] The inference device provided in this embodiment may include, but is not limited to, a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.
[0150] Among them, the processor 61 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 61 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 61 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 61 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 61 may further include an artificial intelligence (AI) processor, and the AI processor is used to process computing operations related to machine learning.
[0151] The memory 60 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 60 may further include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices and flash storage devices. In this embodiment, the memory 60 is at least used to store the following computer program 601. After the computer program is loaded and executed by the processor 61, it can implement the relevant steps of the inference method disclosed in any of the foregoing embodiments. In addition, the resources stored in the memory 60 may also include an operating system 602 and data 603, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 602 may include Windows, Unix, Linux, etc. The data 603 may include, but is not limited to, data required for the inference method.
[0152] In some embodiments, the inference device may further include a display screen 62, an input / output interface 63, a communication interface 64, a power supply 65, and a communication bus 66.
[0153] Those skilled in the art can understand that Figure 5 the structure shown in does not constitute a limitation on the inference device, and it may include more or fewer components than those shown in the figure.
[0154] It can be understood that if the inference method in the above embodiments is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the current technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and executes all or part of the steps of the methods in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM, magnetic disk, or optical disk, etc., all of which can store program codes.
[0155] Based on this, the embodiments of the present invention further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the inference method as described above.
[0156] The above has introduced in detail an inference method provided by the embodiments of the present invention. The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the device disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0157] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0158] The above has introduced in detail a reasoning method, device, equipment, readable storage medium and program product provided by the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and modifications can still be made to the present invention, and these improvements and modifications also fall within the protection scope of the present invention.
Claims
1. A reasoning method, characterized in that Including: Using a first large language model based on a target prompt template and a question to be inferred, to obtain initial inference results corresponding to each language; wherein, the target prompt template is a template including the logic of the inference question; Based on the initial inference results corresponding to each language and a multiple-choice question prompt template, to obtain a multiple-choice question; wherein, the multiple-choice question prompt template is a template including the logic of constructing a multiple-choice question based on the initial inference results; Based on the multiple-choice question, using a second large language model to perform inference analysis to obtain a final inference result; the process of inference analysis is to first independently analyze the initial inference results of each language, generate the confidence of each option, then compare the initial inference results of each language, mark the conflicting parts, and analyze the conflicting parts; Among them, using a first large language model based on a target prompt template and a question to be inferred, to obtain initial inference results corresponding to each language, including: Obtaining the question to be inferred and the target language corresponding to the question to be inferred; Based on the question to be inferred and the target language, using a target language prompt template to obtain first input information; Based on the question to be inferred and the remaining languages, using the language prompt templates corresponding to the remaining languages to obtain second input information; the target prompt template includes the target language prompt template and the language prompt templates; Under the condition that the question to be inferred always maintains its original language form during the inference process, based on the first input information and the second input information, using the first large language model to obtain the initial inference results corresponding to each language.
2. The inference method according to claim 1, wherein Based on the initial inference results corresponding to each language and a multiple-choice question prompt template, to obtain a multiple-choice question, including: Based on the initial inference results corresponding to each language, obtaining the inference process and the initial inference results corresponding to each language; Using the multiple-choice question target, taking the inference process and the initial inference results corresponding to each language as options to construct the multiple-choice question.
3. The reasoning method according to any one of claims 1 to 2, characterized in that After using a second large language model to perform inference analysis based on the multiple-choice question to obtain a final inference result, it further includes: Determining the target language corresponding to the question to be inferred; Determining whether the final inference result includes an inference result corresponding to the target language; When the final inference result includes an inference result corresponding to the target language, sending the inference result corresponding to the target language so that the language of the feedback inference result is consistent with the target language.
4. The inference method according to claim 3, characterized in that After determining whether the final inference result includes an inference result corresponding to the target language, it further includes: When the final inference result does not include an inference result corresponding to the target language, using a translation model to translate the final inference result based on the target language to obtain a translated inference result; Sending the translated inference result.
5. The reasoning method according to claim 4, characterized in that, After, when the final inference result does not include an inference result corresponding to the target language, using a translation model to translate the final inference result based on the target language to obtain a translated inference result, it further includes: Performing semantic detection on the translated inference result to determine whether the semantics is accurate; When accurate, send the inference result of the translation; When inaccurate, perform inference analysis based on the final inference result and the inference result of the translation to obtain the corrected inference result of the translation.
6. The reasoning method according to claim 1, characterized in that After performing inference analysis using the second large language model based on the multiple-choice question to obtain the final inference result, it further includes: Determine the confidence threshold corresponding to the final inference result; When the confidence threshold is lower than the set confidence threshold, send a prompt message.
7. The reasoning method according to claim 1, wherein The training process of the first large language model includes: Training the large language model using a multilingual dataset to obtain an initial large language model; wherein, the data in the multilingual dataset includes news, encyclopedias, papers, and social media texts; Fine-tuning the initial large language model using a task question set to obtain the first large language model.
8. The reasoning method according to claim 1, characterized in that, The training process of the second large language model includes: Obtain a multilingual multiple-choice question dataset; wherein, the multilingual multiple-choice question dataset includes interference options, and the interference options include multimodal interference data, and the multimodal interference data includes mixed unit interference and mixed semantic interference; Training the multilingual pre-trained model using the multilingual multiple-choice question dataset to obtain the second large language model.
9. The reasoning method according to claim 1, characterized in that, The multiple-choice question prompt template includes question information about the options and each option, and each option is composed of the initial inference result.
10. The reasoning method according to claim 1, wherein The question to be inferred includes at least one of an intelligent education question and a medical diagnosis question.
11. An inference device, characterized in that, It includes: An initial inference module, configured to use the first large language model based on a target prompt template and a question to be inferred to obtain an initial inference result corresponding to each language; wherein, the target prompt template is a template including the logic of the inference question; A multiple-choice question determination module, configured to obtain a multiple-choice question based on the initial inference result corresponding to each language and the multiple-choice question prompt template; wherein, the multiple-choice question prompt template is a template including the logic of constructing a multiple-choice question based on the initial inference result; A final inference module, configured to perform inference analysis using the second large language model based on the multiple-choice question to obtain the final inference result; the process of inference analysis is to first independently analyze the initial inference result of each language, generate the confidence of each option, then compare the initial inference results of each language, mark the conflicting parts, and analyze the conflicting parts; The initial inference module includes: A target language determination module, configured to obtain the question to be inferred and the target language corresponding to the question to be inferred; A first input information determination unit, configured to obtain first input information based on the question to be inferred and the target language using the target language prompt template; A second input information determination unit, configured to obtain second input information based on the question to be inferred and the remaining languages using the language prompt templates corresponding to the remaining languages; the target prompt template includes the target language prompt template and the language prompt template; the language of the question to be inferred in the first input information and the second input information is the same; An initial inference result determination unit, configured to obtain the initial inference result corresponding to each language by using the first large language model based on the first input information and the second input information, under the condition that the problem to be inferred always maintains its original language form during the inference process.
12. An inference device, characterized in that, Comprising: A memory for storing computer programs; A processor for executing the computer program to implement the steps of the inference method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, the steps of the inference method according to any one of claims 1 to 10 are implemented.
14. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps of the inference method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Translated text verification method and related device
CN117252217A
Question and answer method and device based on large language model, medium and equipment
CN118568242A
Multilingual word alignment method based on attention mechanism
CN119808797A