Sample generation method, training method and question answering method based on artificial intelligence

By screening and processing multiple sample reflection results in the big model, detailed target sample reflection results are generated, which solves the problem of insufficient self-reflection ability of the big model, and improves the reflection clarity of the big model and the accuracy of the interactive big model.

CN120494107APending Publication Date: 2025-08-15BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510677293.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-23
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

There are improvement questions in existing big models that have unclear content and are difficult to obtain effective answers in terms of self-reflection ability.

Method used

The sample question and answer pairs are input into multiple first large models respectively. By filtering and processing multiple sample reflection results, the target sample reflection results are generated, and the big model is trained using reflection samples to improve the detailedness and clarity of reflection.

Benefits of technology

Through the training of target sample reflection results of diversity and accuracy, the reflection ability of the big model is improved, and errors can be corrected more clearly and in detail, and the accuracy and robustness of the interactive big model are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120494107A_ABST
    Figure CN120494107A_ABST
Patent Text Reader

Abstract

The invention provides a sample generation method based on artificial intelligence, a large model training method, a question and answer method and device, electronic equipment, a storage medium, a program product and an intelligent agent, and relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like. According to the specific implementation scheme, sample question and answer pairs are input into a plurality of different first large models respectively, a plurality of sample reflection results are obtained, and the sample reflection results comprise a plurality of sub-results used for indicating correlation of the sample question and answer pairs; under the condition that the multiple sample reflection results are different, an intermediate sample reflection result is determined from the multiple sample reflection results, and at least one sub-result of the intermediate sample reflection result meets a preset screening condition; processing other sub-results in the intermediate sample reflection result to obtain a target sample reflection result; and generating an reflection sample based on the sample question-answer pair and the target sample reflection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as computer vision, deep learning, and large models, and specifically to artificial intelligence-based sample generation methods, large model training methods, question-answering methods, devices, electronic devices, storage media, program products, and intelligent agents. Background Art

[0002] In the field of artificial intelligence, significant progress has been made in the development of large models, such as large language models and multimodal models. These models have demonstrated strong capabilities in a variety of tasks, including text-to-image processing, machine translation, and question-answering. However, as models scale and tasks become more complex, improving the self-reflection capabilities of these models has become a key challenge. Summary of the Invention

[0003] The present disclosure provides an artificial intelligence-based sample generation method, a large model training method, a question-answering method, an apparatus, an electronic device, a storage medium, a program product, and an intelligent agent.

[0004] According to one aspect of the present disclosure, a sample generation method based on artificial intelligence is provided, comprising: inputting sample question and answer pairs into different multiple first large models respectively to obtain multiple sample reflection results, wherein the above-mentioned sample reflection results include multiple sub-results for indicating the correlation of the above-mentioned sample question and answer pairs; when the multiple above-mentioned sample reflection results are different, determining an intermediate sample reflection result from the multiple above-mentioned sample reflection results, and at least one sub-result of the above-mentioned intermediate sample reflection result meets a predetermined screening condition; processing other sub-results in the above-mentioned intermediate sample reflection result to obtain a target sample reflection result; and generating a reflection sample based on the above-mentioned sample question and answer pairs and the above-mentioned target sample reflection result.

[0005] According to another aspect of the present disclosure, a large model training method is provided, comprising: training a second large model using reflection samples to obtain a reflection large model; wherein the reflection samples are generated using the sample generation method described above.

[0006] According to another aspect of the present disclosure, a question-answering method is provided, comprising: obtaining a question input through an interactive interface; inputting the question into an interactive big model to obtain an initial answer; inputting the question and the initial answer into a reflection big model to obtain a reflection result; when the reflection result indicates that the question and the initial answer are irrelevant, inputting the question, the initial answer and the reflection result into the interactive big model to obtain a target answer; and displaying the target answer on the interactive interface; wherein the reflection big model is trained using the training method of the big model as described above.

[0007] According to another aspect of the present disclosure, there is provided an artificial intelligence-based sample generation device, comprising: a sample input module for inputting sample question-answer pairs into different multiple first large models respectively to obtain multiple sample reflection results, wherein the above-mentioned sample reflection results include multiple sub-results for indicating the correlation of the above-mentioned sample question-answer pairs; a screening module for determining an intermediate sample reflection result from the multiple sample reflection results when the multiple sample reflection results are different, and at least one sub-result of the above-mentioned intermediate sample reflection result meets a predetermined screening condition; a sample processing module for processing other sub-results in the above-mentioned intermediate sample reflection result to obtain a target sample reflection result; and a sample generation module for generating a reflection sample based on the above-mentioned sample question-answer pairs and the above-mentioned target sample reflection result.

[0008] According to another aspect of the present disclosure, a large model training device is provided, comprising: a training module for training a second large model using reflection samples to obtain a reflection large model; wherein the reflection samples are generated using the sample generation device as described above.

[0009] According to another aspect of the present disclosure, a question-answering device is provided, including: an acquisition module for acquiring questions input through an interactive interface; an interaction module for inputting the above questions into an interactive big model to obtain an initial answer; a reflection module for inputting the above questions and the above initial answers into the reflection big model to obtain a reflection result; a re-interaction module for inputting the above questions, the above initial answers and the above reflection results into the interactive big model to obtain a target answer when the above reflection result indicates that the above questions and the above initial answers are irrelevant; and a display module for displaying the above target answer on the above interactive interface; wherein the above reflection big model is trained using the training device of the big model as described above.

[0010] According to another aspect of the present disclosure, an artificial intelligence agent is provided, comprising: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and obtaining output information by calling the large model to execute the method described above; and an output module for outputting the output information obtained by the processing module.

[0011] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described above.

[0012] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described above.

[0013] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, which implements the method described above when executed by a processor.

[0014] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The accompanying drawings are used to better understand the present invention and do not constitute a limitation of the present invention.

[0016] Figure 1 Schematically illustrates an exemplary system architecture to which an artificial intelligence-based sample generation method, a large model training method, a question-answering method, and an apparatus according to an embodiment of the present disclosure can be applied;

[0017] Figure 2 The flowchart of the sample generation method based on artificial intelligence according to an embodiment of the present disclosure is schematically shown;

[0018] Figure 3 The following schematically illustrates a flow chart of determining a sample reflection result according to an embodiment of the present disclosure;

[0019] Figure 4 The following schematically illustrates a process flow for determining a target sample reflection result according to an embodiment of the present disclosure;

[0020] Figure 5 A schematic diagram schematically illustrates updating a reflection sample set according to an embodiment of the present disclosure;

[0021] Figure 6 The flowchart of the training method of the large model according to the embodiment of the present disclosure is schematically shown;

[0022] Figure 7 The following schematically shows a flow chart of a question-answering method according to an embodiment of the present disclosure;

[0023] Figure 8 The following schematically shows a flow chart of a question-answering method according to another embodiment of the present disclosure;

[0024] Figure 9 A block diagram of an artificial intelligence-based sample generation device according to an embodiment of the present disclosure is schematically shown;

[0025] Figure 10 A block diagram of a large model training device according to an embodiment of the present disclosure is schematically shown;

[0026] Figure 11 Schematically shows a block diagram of a question-answering device according to an embodiment of the present disclosure;

[0027] Figure 12 A block diagram schematically illustrates a structure of an artificial intelligence agent according to an embodiment of the present disclosure; and

[0028] Figure 13 A block diagram of an electronic device suitable for implementing an artificial intelligence-based sample generation method, a large model training method, and a question-answering method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0029] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0030] With the rapid development of artificial intelligence (AI), technologies such as large language models (LLMs), multimodal large language models (MLLMs), and intelligent agents are playing an increasingly important role. Large models represent a major breakthrough in AI in recent years. Based on deep learning, specifically the Transformer (encoder-decoder) architecture, they are pre-trained on large-scale corpus data to achieve in-depth understanding and generation of human language. Multimodal large models are capable of processing and understanding multiple types of information. Unlike large language models, which only process text, multimodal large models can integrate multiple modalities such as text, images, audio, and video, and perform comprehensive understanding and reasoning, ultimately achieving more powerful capabilities. Intelligent agents are systems that can autonomously perceive their environment, make decisions, and execute actions. They possess fundamental characteristics such as autonomy, interactivity, responsiveness, and adaptability, enabling them to independently complete tasks in complex and ever-changing environments. The emergence of intelligent agents marks the advancement of AI from simple rule-matching and computational simulation to higher levels of autonomous intelligence.

[0031] Reflectors are used to evaluate the accuracy of AI-based reasoning results against user-entered queries, helping to identify reasoning errors and improve them. However, existing reflectors still have some issues. For example, the reflection results are unclear and lack detail, making it difficult to effectively improve the answers based on the reflection results.

[0032] In view of this, the present disclosure provides a sample generation method based on artificial intelligence, including: inputting sample question and answer pairs into different multiple first large models respectively to obtain multiple sample reflection results, wherein the sample reflection results include multiple sub-results for indicating the correlation of the sample question and answer pairs; when the multiple sample reflection results are different, determining an intermediate sample reflection result from the multiple sample reflection results, and at least one sub-result of the intermediate sample reflection result meets a predetermined screening condition; processing other sub-results in the intermediate sample reflection result to obtain a target sample reflection result; and generating a reflection sample based on the sample question and answer pair and the target sample reflection result.

[0033] By using the sample generation method provided by the embodiment of the present disclosure, it is possible to use multiple first large models to obtain sample reflection results respectively, thereby ensuring the diversity of the target sample reflection results, and by screening the sub-results and processing other sub-results, the accuracy of the target sample reflection results is guaranteed, and then the reflection samples including the target sample reflection results are used to train the large model, thereby improving the detailedness and clarity of the reflection of the large model after training.

[0034] Figure 1 The exemplary system architecture of the artificial intelligence-based sample generation method, large model training method, question-answering method and device according to the embodiments of the present disclosure is schematically shown.

[0035] It should be noted that Figure 1 The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure, but do not mean that the embodiments of the present disclosure may not be used in other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which an artificial intelligence-based sample generation method, a large model training method, and a question-answering method and apparatus may be applied may include a terminal device, but the terminal device may implement the artificial intelligence-based sample generation method, large model training method, question-answering method, and apparatus provided in the embodiments of the present disclosure without interacting with a server.

[0036] like Figure 1As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used as a medium for providing communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.

[0037] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, and 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).

[0038] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0039] Server 105 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by users using terminal devices 101, 102, and 103. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.

[0040] It should be noted that the artificial intelligence-based sample generation method, large model training method, and question-answering method provided in the embodiments of the present disclosure can generally be executed by the terminal device 101, 102, or 103. Accordingly, the artificial intelligence-based sample generation method, large model training method, and question-answering device provided in the embodiments of the present disclosure can also be set in the terminal device 101, 102, or 103.

[0041] Alternatively, the sample generation method based on artificial intelligence, the training method of the large model, and the question-answering method provided in the embodiments of the present disclosure may generally be executed by the server 105. Accordingly, the sample generation method based on artificial intelligence, the training method of the large model, and the question-answering device provided in the embodiments of the present disclosure may generally be set in the server 105. The sample generation method based on artificial intelligence, the training method of the large model, and the question-answering method provided in the embodiments of the present disclosure may also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Accordingly, the sample generation method based on artificial intelligence, the training method of the large model, and the question-answering device provided in the embodiments of the present disclosure may also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.

[0042] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.

[0043] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.

[0044] In the technical solution disclosed herein, the user's authorization or consent is obtained before obtaining or collecting the user's personal information.

[0045] It should be noted that the sequence numbers of the operations in the following method are only used to indicate the operation for the purpose of description, and should not be regarded as indicating the order in which the operations should be performed. Unless explicitly stated, the method does not need to be performed in the order shown.

[0046] Figure 2 The flowchart of the sample generation method based on artificial intelligence according to an embodiment of the present disclosure is schematically shown.

[0047] like Figure 2 As shown, the method includes operations S210 to S240.

[0048] In operation S210, the sample question-answer pairs are respectively input into a plurality of different first large models to obtain a plurality of sample reflection results.

[0049] In operation S220 , if the plurality of sample reflection results are different, an intermediate sample reflection result is determined from the plurality of sample reflection results, and at least one sub-result of the intermediate sample reflection result meets a predetermined screening condition.

[0050] In operation S230 , other sub-results in the intermediate sample reflection result are processed to obtain a target sample reflection result.

[0051] In operation S240 , a reflection sample is generated based on the sample question-answer pair and the target sample reflection result.

[0052] The plurality of first large models may include a large language model, but is not limited thereto, and may also include a multimodal large model. Any model may be used for reflection to obtain a sample reflection result that characterizes the correlation between the question and the answer in the sample question-answer pair.

[0053] There is no limitation on the type of sample question-answer pairs. For example, they can be text-type question-answer pairs or multimodal question-answer pairs.

[0054] Optionally, the sample reflection result includes multiple sub-results indicating the relevance of the sample question-answer pair. The sub-results are of different types. Relevance can be understood as the correctness of the question and the corresponding answer in the sample question-answer pair, but is not limited to this. Relevance can also indicate, if an error exists, irrelevant content, the reason, and how to correct the answer to make it relevant to the question.

[0055] For example, the sample reflection result may include the following multiple sub-results: whether the answer is wrong relative to the corresponding question, the type of error if there is an error, the reason why the answer is wrong relative to the corresponding question, and prompt information used to guide correct reasoning.

[0056] By using multiple first-class models to reflect on the same sample question-answer pair, multiple sample reflection results can be obtained. If the multiple sample reflection results are the same, the accuracy of the multiple sample reflection results can be guaranteed, and any sample reflection result can be used as the target sample reflection result.

[0057] In the case where the multiple sample reflection results are different, it can be preliminarily determined that at least one sub-result in at least one sample reflection result is erroneous. An intermediate sample reflection result can be determined from the multiple sample reflection results.

[0058] For example, the sample reflection results include three: sample reflection results A, B, and C. Sample reflection result A includes sub-result A1, sub-result A2, and sub-result A3. Sample reflection result B includes sub-result B1, sub-result B2, and sub-result B3. Sample reflection result C includes sub-result C1, sub-result C2, and sub-result C3.

[0059] If the sub-results A1, B1, and C1 of the same type in the sample reflection results A, B, and C are different, it can be determined that the sample reflection results A, B, and C are different. An intermediate sample reflection result can be determined from the sample reflection results A, B, and C.

[0060] For example, if the sample reflection results A and B in the sample reflection results A, B and C both have at least one sub-result that meets the predetermined screening condition, then it is determined that the sample reflection results A and B are both intermediate sample reflection results.

[0061] There is no limitation on the predetermined screening conditions, as long as they are conditions that can characterize that at least one sub-result of the intermediate sample reflection result is correct.

[0062] For example, a model recognition approach can be used to input multiple sample reflection results into a screening model to determine an intermediate sample reflection result that includes sub-results that meet predetermined screening criteria. Alternatively, a rule recognition approach can be used to match multiple sub-results of each of the multiple sample reflection results against a rule to determine an intermediate sample reflection result that includes sub-results that meet predetermined screening criteria.

[0063] If any sub-result in the intermediate sample reflection result meets the predetermined screening conditions, the intermediate sample reflection result can be used as the target sample reflection result. If any sub-result in the intermediate sample reflection result does not meet the predetermined screening conditions, the sub-result that does not meet the predetermined screening conditions can be used as another sub-result, and the other sub-results can be processed, such as updated, to obtain the target sample reflection result.

[0064] By using the sample generation method provided by the present invention, it is possible to use multiple first large models to obtain sample reflection results respectively, thereby ensuring the diversity of the target sample reflection results, and by screening the sub-results and processing other sub-results, the accuracy of the target sample reflection results is guaranteed, and then the reflection samples including the target sample reflection results are used to train the large model, thereby improving the detailedness and clarity of the reflection of the large model after training.

[0065] In related examples, multiple sample reflection results can be directly processed, such as weighted summation or semantic fusion, to obtain the target sample reflection result.

[0066] Compared to directly generating target sample reflection results without filtering, using predetermined filtering conditions to filter multiple sample reflection results and determine intermediate sample reflection results can reduce subsequent processing volume and improve processing efficiency. Furthermore, it is possible to perform detailed identification of multiple sub-results within the intermediate sample reflection results, filtering out sub-results that meet the predetermined filtering conditions and processing only the remaining sub-results, further reducing processing volume and improving processing efficiency.

[0067] Figure 3 The flowchart of determining the sample reflection result according to the embodiment of the present disclosure is schematically shown.

[0068] like Figure 3 As shown, question 310 in a sample question-answer pair may include an image 311 of an hourglass object and a text question 312 including the question "What is this? Who invented it? Explain the principle in plain language." Question 310 may be input into interactive macro model M310 to obtain answer 320 in the sample question-answer pair. Answer 320 may include a detailed explanation of text question 312.

[0069] like Figure 3 As shown, the question 310 and the answer 320 can be combined and input into the first large model M320 as a sample question-answer pair to obtain a sample reflection result 330.

[0070] like Figure 3 As shown, the sample reflection result 330 may include: "state", identification information used to characterize that the answer is wrong relative to the corresponding question, such as "0", the error type "type" such as "image recognition error", the reason for the error "reason" such as "the glass hourglass in the image is a sculpture or decoration and does not have a timing function", and prompt information "hint" to guide correct reasoning such as "carefully observe the structure and use of the objects in the image to determine whether they have actual functionality rather than similarity in appearance".

[0071] The above describes in detail the generation of the sample reflection result and the contents of the multiple sub-results in the sample reflection result through the embodiments. The following further describes how to determine whether the multiple sample reflection results are different.

[0072] According to an embodiment of the present disclosure, when executing Figure 2 In operation S220 , before determining the intermediate sample reflection result from the multiple sample reflection results, the sample generation method may further include an operation of determining that the multiple sample reflection results are different when multiple sub-results of any type in the multiple sample reflection results are different.

[0073] For example, multiple sub-results of the same type may include a sub-result indicating whether an answer is incorrect relative to the corresponding question. If, among sample reflection results A, B, and C, sub-result A1 indicating whether an error exists is "0," sub-result B1 is "0," and sub-result C1 is "1," then the multiple sample reflection results are determined to be different.

[0074] According to the embodiments of the present disclosure, combined with the characteristic that the sample reflection results include multiple sub-results, if there are differences in multiple sub-results of any type, it is determined that the multiple sample reflection results are different. This can refine the recognition granularity of the multiple sample reflection results, thereby ensuring the diversity and clarity of the target sample reflection results while improving the accuracy of the target sample reflection results.

[0075] The following will use multiple embodiments to specifically illustrate how to determine the intermediate sample reflection result when the reflection results of multiple samples are different.

[0076] According to an optional embodiment of the present disclosure, for example Figure 2 Operation S220 shown, in which, when it is determined that the multiple sample reflection results are different, an intermediate sample reflection result is determined from the multiple sample reflection results, may include the following operations.

[0077] For example, multiple first reference sub-results of classification types are determined from the multiple sample reflection results. A target first reference sub-result that meets a predetermined screening condition is determined from the multiple first reference sub-results. The sample reflection results including the first target reference sub-result are used as intermediate sample reflection results.

[0078] The first reference sub-result of the classification type includes at least one of the following: whether the answer in the sample question-answer pair is incorrect relative to the corresponding question, and the error type if an error exists.

[0079] Optionally, the result content of the first reference sub-result of a classification type is relatively simple and fixed. For example, a first reference sub-result indicating whether an error exists may include identification information of "0" to identify an error or identification information of "1" to identify a correct result. In the event of an error, the first reference sub-result of an error type may include: image recognition error, logic error, instruction error, etc.

[0080] According to an embodiment of the present disclosure, multiple first reference sub-results of a classification type are used as reference data for priority screening. This can simplify the difficulty of screening and improve screening efficiency and accuracy by combining the simple and fixed content characteristics of the first reference sub-results of a classification type.

[0081] According to an embodiment of the present disclosure, the predetermined screening condition may include at least one of the following: the effect evaluation value of the first large model is greater than a predetermined evaluation threshold; and the proportion of identical sub-results is greater than a predetermined proportion threshold.

[0082] Taking the example of a classification type where the first reference sub-result includes whether the answer in the sample question-answer pair is incorrect relative to the corresponding question, and the predetermined screening condition includes that the proportion of the same sub-result is greater than a predetermined proportion threshold, the first reference sub-result A1 for identifying whether there is an error in the sample reflection results A, B, and C is "0", the first reference sub-result B1 is "0", and the first reference sub-result C1 is "1". If the proportion of the first reference sub-result identified as "0" is greater than the predetermined proportion threshold, then the sample reflection results A and B are determined to be intermediate sample reflection results. The sample reflection result C is deleted.

[0083] Optionally, the predetermined ratio threshold may be flexibly adjusted according to the number of first reference sub-results, which will not be described in detail here.

[0084] Using the ratio of identical sub-results greater than a predetermined ratio threshold as a predetermined screening condition, combined with the characteristics of the first reference sub-result of the classification type, can improve screening accuracy while reducing screening difficulty.

[0085] Taking the example of the first reference sub-result of the classification type including whether the answer in the sample question-answer pair is wrong relative to the corresponding question, and the predetermined screening conditions including the proportion of the same sub-results being greater than the predetermined proportion threshold and the effect evaluation value of the first large model being greater than the predetermined evaluation threshold, the first reference result A1 for identifying whether there is an error in the sample reflection results A, B and C is "0", the first reference sub-result B1 is "0" and the first reference sub-result C1 is "1". If the proportion of the first reference sub-result identified as "0" is greater than the predetermined proportion threshold, the sample reflection results A and B are determined to be candidate intermediate sample reflection results. The first reference result A2 of the error type in the candidate intermediate sample reflection results A and B is determined to be "image recognition error" and the first reference sub-result B2 is determined to be "logical error". In addition, the effect evaluation value of the first large model A is determined to be effect evaluation value A, the effect evaluation value of the first large model B is determined to be effect evaluation value B, and the effect evaluation value A is greater than the predetermined evaluation threshold, and the effect evaluation value B is less than the predetermined evaluation threshold. Then, the candidate intermediate sample reflection result A is determined to be the intermediate sample reflection result.

[0086] Using multiple predetermined screening conditions to jointly participate in screening can adapt to sample reflection results including multiple sub-results of different types, thereby improving screening accuracy, reducing screening difficulty, and expanding screening capabilities and scope.

[0087] Optionally, the effect evaluation value may include one or more of a precision evaluation value, an accuracy evaluation value, a processing efficiency evaluation value, and the like. Exemplarily, the effect evaluation value may be mapped according to the application scenario or application field, and a mapping relationship may be established for each first large model. For example, for the image recognition scenario, the effect evaluation value of the first large model A is A1; for the tool call scenario, the effect evaluation value of the first large model A is A2; and for the instruction understanding scenario, the effect evaluation value of the first large model A is A3.

[0088] For sample question-answer pairs in different application scenarios, the effect evaluation value of the first model can be determined based on the application scenario of the sample question-answer pairs and the mapping relationship of the first model.

[0089] This improves the effectiveness and pertinence of the effect evaluation value of the first model.

[0090] According to another optional embodiment of the present disclosure, Figure 2 Operation S220 shown, determining an intermediate sample reflection result from a plurality of sample reflection results, may further include the following operations.

[0091] For example, based on the respective effect evaluation values of the plurality of first-large models, a target first-large model is determined from the plurality of first-large models, and a sub-result in the sample reflection result output by the target first-large model is determined to meet a predetermined screening condition to obtain an intermediate sample reflection result.

[0092] The application scenario of the sample question-answer pair can be determined first, and based on the mapping relationship between the application scenario and each of the multiple first-large models, the effect evaluation value of each first-large model can be determined.

[0093] For example, if the effect evaluation values A, B, and C of the first large model A, the first large model B, and the first large model C are respectively greater than the predetermined evaluation threshold, the first large model A is used as the target first large model. The sample reflection result output by the first large model A is used as the intermediate sample reflection result.

[0094] Optionally, the first model is preferentially screened using the effect evaluation value, which can deal with different types of sub-results, such as whether the answer is wrong relative to the corresponding question, the type of error if there is an error, the reason why the answer is wrong relative to the corresponding question, prompt information used to guide correct reasoning, and other sub-results.

[0095] By using domain information to screen the first model, the first model that is suitable for reflecting on the sample question and answer pairs of the application scenario can be determined, avoiding the introduction of the output results of other first models that do not meet the predetermined evaluation threshold as noise, thereby causing the problem of misidentification. In this way, the screening efficiency and quality of the sample reflection results can be improved by combining the characteristics of model applicability.

[0096] According to the embodiments of the present disclosure, Figure 2 Operation S230, in which the other sub-results in the intermediate sample reflection result are processed to obtain a target sample reflection result, may include determining a second reference sub-result of a generative type from the other sub-results, generatively processing the second reference sub-result to obtain a target second reference sub-result, and obtaining a target sample reflection result based on the target second reference sub-result.

[0097] The generative type second reference sub-result includes at least one of the following: the reason why the answer in the sample question-answer pair is incorrect relative to the question in the sample question-answer pair, and prompt information for guiding correct reasoning.

[0098] The content of the second reference sub-result of a generative type varies for different sample question-answer pairs. Furthermore, due to the diversity of textual content, even semantically identical expressions may differ. Therefore, multiple second reference sub-results of the same generative type in multiple intermediate sample reflection results may differ from each other.

[0099] Therefore, the second reference sub-result can be subjected to generative processing, such as semantic fusion processing, but is not limited to this. It can also include content expansion processing, polishing processing, etc. to obtain the target second reference sub-result. Any target second reference sub-result that can be obtained using artificial intelligence generated content (AIGC) technology will be sufficient.

[0100] The target second reference sub-result and the sub-results meeting the predetermined screening conditions may be combined to obtain the target sample reflection result.

[0101] Optimizing the generative type's second reference sub-result using the above method can improve the completeness and accuracy of the target second reference sub-result, and avoid the problem of poor training effect when using reflection samples to train large models due to incomplete or ambiguous description content.

[0102] According to an embodiment of the present disclosure, the second reference sub-result may include multiple ones.

[0103] Generatively processing the second reference sub-results to obtain a target second reference sub-result may include: determining weights representing the importance of each of the plurality of second reference sub-results based on the respective effect evaluation values of the plurality of first large models; and performing semantic fusion processing on the plurality of second reference sub-results based on the respective weights of the plurality of second reference sub-results to obtain the target second reference sub-result.

[0104] Optionally, a similarity comparison can be performed between the domain information of the first large model and the domain information of the sample question-answer pair to obtain a matching degree. Based on the matching degree, a weight is determined to represent the importance of each of the multiple second reference sub-results. However, this is not limited to this embodiment. The weight of the second reference sub-result can also be determined using the effect evaluation value of the first large model.

[0105] The second reference sub-result and the weight can be used as an input data group, and multiple input data groups can be input into the semantic fusion model, so that the semantic fusion model can be used to perform semantic fusion processing on the multiple input data groups to obtain the target second reference sub-result.

[0106] The semantic fusion large model can be a large language model, as long as it is a large model that can perform semantic fusion operations.

[0107] By combining the respective weights of multiple second reference sub-results and performing semantic fusion processing on the multiple second reference sub-results, it is possible to clarify the importance of each of the multiple second reference sub-results based on multiple weights during the semantic fusion processing, so as to highlight the fusion degree of the content of the second reference sub-results with high importance and reduce the fusion degree of the content of the second reference sub-results with low importance, thereby ensuring the accuracy of the target second reference sub-results while being appropriately detailed and concise.

[0108] Figure 4 The flowchart of determining the reflection result of the target sample according to the embodiment of the present disclosure is schematically shown.

[0109] like Figure 4 As shown, the sample question-answer pair 410 can be input into multiple first-class models, such as first-class model 1, ..., first-class model n, to obtain multiple sample reflection results 420. Based on the first reference sub-results of the classification type of each of the multiple sample reflection results 420, such as the sub-results of "state" and "type", the sample reflection results 421 and the sample reflection results 422 are determined to be intermediate sample reflection results. Other sub-results in the sample reflection result 421, such as the second reference sub-result 4211 of the generative type, and other sub-results in the sample reflection result 422, such as the second reference sub-result 4221 of the generative type, such as the sub-results of "reason" and "hint", are determined. The second reference sub-result 4211 and the second reference sub-result 4221 are semantically fused to obtain a target second reference sub-result 431. The target second reference sub-result 431 is combined with the sub-results in the intermediate sample reflection results that meet the predetermined conditions, such as the sub-results where "state" is 1 and "type" is an image recognition error, to obtain a target sample reflection result 440.

[0110] According to an embodiment of the present disclosure, when executing Figure 2After operation S240, the sample generation method may further include, if it is determined that the test reflection sample contains an error, updating the content of the reflection sample set to which the test reflection sample belongs. The test reflection sample is obtained by screening from the reflection sample set, which is obtained by classifying multiple reflection samples based on the domain information of the sample question-answer pairs.

[0111] Optionally, the original sample question-answer pairs can be preprocessed to obtain sample question-answer pairs. The preprocessing operation is not limited and may include, for example, image screening, image noise reduction, text polishing, text rewriting, etc. Any preprocessing operation that ensures the sample question-answer pairs meet the predetermined quality will suffice.

[0112] Multiple first-class models can be used to annotate sample question-answer pairs using a sample generation method to obtain reflection samples. Multiple reflection samples can be classified according to domain information to obtain multiple reflection sample sets in different domains.

[0113] For example, domain information includes image recognition domain, tool calling domain, instruction understanding domain, etc.

[0114] Multiple reflection samples can be extracted from each of the multiple reflection sample sets to serve as test reflection samples. If it is determined that the test reflection sample contains errors, it can be preliminarily determined that the reflection sample set to which the test reflection sample belongs is likely to contain errors. Content updates can then be performed on the multiple reflection samples in the reflection sample set. Exemplarily, the content update includes updating the reflection results of the target sample.

[0115] Optionally, a test reflection sample with an incorrect reflection result for the target sample can be determined to have an error. A large model trained with the reflection sample set can be used to process the sample question-answer pairs in the test reflection sample, and the output reflection result can be compared with the reflection result in the test reflection sample. If the two are different, it can be preliminarily determined that the test reflection sample has an error. Other instruction evaluation methods, such as rule evaluation or model evaluation, can be used to further evaluate the test reflection sample. If the evaluation result indicates that the test reflection sample has an error, it can be determined that the test reflection sample has an error.

[0116] Figure 5 The figure schematically shows a schematic diagram of updating a reflection sample set according to an embodiment of the present disclosure.

[0117] like Figure 5As shown, the multiple original sample question-answer pairs in the original database 510 are preprocessed respectively to obtain multiple sample question-answer pairs 520. The multiple sample question-answer pairs 520 are labeled respectively to obtain the target sample reflection result of each sample question-answer pair as a label, thereby obtaining multiple reflection samples 530. The multiple reflection samples 530 are classified according to the domain information to obtain multiple reflection sample sets 540. Test reflection samples 550 are extracted from each reflection sample set 540, and the multiple reflection sample sets 540 are respectively quality evaluated to obtain multiple quality evaluation results 560. When the quality evaluation result indicates that the test reflection sample 550 has an error, the target sample reflection result of the reflection sample set to which the test reflection sample 550 belongs is updated.

[0118] The quality of the reflection sample set is evaluated and updated based on the quality assessment results, thereby improving the quality of the reflection samples used for training. In addition, multiple reflection samples are divided into sets based on domain information, thereby improving the efficiency of quality evaluation while reducing the amount of updated data.

[0119] Figure 6 The flowchart of the training method of the large model according to the embodiment of the present disclosure is schematically shown.

[0120] like Figure 6 As shown, the method includes operation S610.

[0121] In operation S610, the second large model is trained using the reflection samples to obtain a reflection large model.

[0122] According to an embodiment of the present disclosure, the reflection sample is generated using the sample generation method provided by an embodiment of the present disclosure.

[0123] The second model can be an open-source pre-trained model, with no restrictions on its architecture or type. As long as the second model can be trained using the reflection samples, and the resulting reflection model has the ability to find errors, it will be sufficient.

[0124] Supervised fine-tuning (SFT) can be used for training. For example, a sample question-answer pair is input into the second-largest model, which then outputs a training reflection result. This training reflection result is then compared with the target sample reflection result corresponding to the sample question-answer pair. Based on this comparison, the model parameters of the second-largest model are adjusted until the training reflection result output by the second-largest model and the target sample reflection result meet the training conditions.

[0125] The comparison can be performed using a loss function, but is not limited thereto. Semantic similarity can also be used for comparison. The training condition can include a training round reaching a threshold, or a semantic similarity between the training reflection result and the target sample reflection result reaching a threshold.

[0126] By training the second large model using the reflection samples provided by the embodiments of the present disclosure, the trained reflection large model can be endowed with multiple reflection capabilities of different types, such as the ability to identify error types and error causes between question-answer pairs, and the ability to output prompt information to guide correct reasoning. This improves the model performance of the reflection large model and the clarity and accuracy of reflection.

[0127] According to the embodiment of the present disclosure, Figure 6 The reflection big model obtained is combined with the interaction big model and applied to the question-answering scenario. It can utilize the effective error correction ability of the reflection big model to further improve the accuracy and robustness of the interaction big model.

[0128] For example, in scenarios involving complex visual interactions such as autonomous driving, medical diagnosis, and intelligent teaching assistance, service quality and user experience will be improved, promoting the in-depth application and development of artificial intelligence technology in a wider range of fields.

[0129] The following will be combined Figure 7 and Figure 8 Further explanation is provided for the question-and-answer scenario.

[0130] Figure 7 The flowchart of the question-answering method according to an embodiment of the present disclosure is schematically shown.

[0131] like Figure 7 As shown, the method includes operations S710 to S750.

[0132] In operation S710 , a question input through an interactive interface is obtained.

[0133] In operation S720 , the question is input into the interactive macro model to obtain an initial answer.

[0134] In operation S730 , the question and the initial answer are input into the reflection model to obtain a reflection result.

[0135] In operation S740 , when the reflection result indicates that there is no correlation between the question and the initial answer, the question, the initial answer, and the reflection result are input into the interactive macro model to obtain a target answer.

[0136] In operation S750 , the target answer is displayed on the interactive interface.

[0137] The reflection big model provided by the embodiment of the present disclosure is trained using the training method of the big model provided by the above embodiment.

[0138] There is no limitation on the type of interactive large model, which can include large language models and / or multimodal large models. Text or a combination of text and images can be input into the interactive large model to obtain an initial answer.

[0139] Input the question and initial answer into the reflection model to obtain the reflection result. If the reflection result shows the correlation between the question and the initial answer, the initial answer can be directly used as the target answer and displayed in the interactive interface.

[0140] When the reflection result represents that there is no correlation between the question and the initial answer, the question, the initial answer and the reflection result can be input into the interactive large model to obtain the target answer.

[0141] The correlation between the question and the initial answer can be understood as a match between the question and the initial answer, and the initial answer is accurate relative to the question.

[0142] Correspondingly, there is no correlation between the question and the initial answer, which can be understood as an error in the initial answer relative to the question.

[0143] Because the reflection model is trained using reflection samples, it has the ability to identify and feedback multiple reflection sub-results of different types. For example, the reflection results include at least two of the following: whether the answer is wrong relative to the corresponding question, the type of error if there is an error, the reason why the answer is wrong relative to the corresponding question, and prompt information used to guide correct reasoning.

[0144] Because the reflective model and the interactive model are independent and separate, the reflection results output by the reflective model can explicitly explain the content related to the error, which helps the interactive model effectively correct errors after receiving prompts. Furthermore, the diverse sub-results within the reflection results enhance the clarity and completeness of the reflection results. This allows the interactive model to output accurate and highly relevant target answers based on the reflection results, initial answers, and questions, avoiding the problem of hallucinations when outputting answers.

[0145] Figure 8 The figure schematically shows a flow chart of a question-answering method according to another embodiment of the present disclosure.

[0146] like Figure 8As shown, question 810 can be input into interactive macro model M810 to obtain initial answer 820. Initial answer 820 is input into reflective macro model M820 to obtain reflective result 830. Based on reflective result 830, it is determined whether question 810 and initial answer 820 are related. If they are related, the initial answer is used as target answer 840. If they are not related, question 810, initial answer 820, and reflective result 830 are input into interactive macro model to obtain target answer 840.

[0147] Figure 9 The block diagram of the sample generation device based on artificial intelligence according to an embodiment of the present disclosure is schematically shown.

[0148] like Figure 9 As shown, the artificial intelligence-based sample generation device 900 includes: a sample input module 910 , a screening module 920 , a sample processing module 930 , and a sample generation module 940 .

[0149] The sample input module 910 is used to input the sample question-answer pairs into different multiple first large models respectively to obtain multiple sample reflection results, wherein the sample reflection results include multiple sub-results for indicating the relevance of the sample question-answer pairs.

[0150] The screening module 920 is configured to determine an intermediate sample reflection result from the multiple sample reflection results when the multiple sample reflection results are different, wherein at least one sub-result of the intermediate sample reflection result meets a predetermined screening condition.

[0151] The sample processing module 930 is used to process other sub-results in the intermediate sample reflection result to obtain the target sample reflection result.

[0152] The sample generation module 940 is used to generate a reflection sample based on the sample question-answer pair and the target sample reflection result.

[0153] According to an embodiment of the present disclosure, the screening module includes: a first screening submodule, a second screening submodule and a third screening submodule.

[0154] The first screening submodule is used to determine multiple first reference sub-results of classification type from multiple sample reflection results, wherein the first reference sub-results of classification type include at least one of the following: whether the answer in the sample question-answer pair is wrong relative to the corresponding question, and the type of error if there is an error.

[0155] The second screening submodule is configured to determine a target first reference sub-result that meets a predetermined screening condition from the plurality of first reference sub-results.

[0156] The third screening submodule is configured to use the sample reflection result including the first target reference sub-result as the intermediate sample reflection result.

[0157] According to an embodiment of the present disclosure, the screening module includes: a fourth screening sub-module and a fifth screening sub-module.

[0158] The fourth screening submodule is used to determine a target first large model from the multiple first large models based on the respective effect evaluation values of the multiple first large models.

[0159] The fifth screening submodule is used to determine whether the sub-results in the sample reflection results output by the target first large model meet the predetermined screening conditions and obtain the intermediate sample reflection results.

[0160] According to an embodiment of the present disclosure, the predetermined screening condition includes at least one of the following: the effect evaluation value of the first large model is greater than a predetermined evaluation threshold; the proportion of the same sub-results is greater than a predetermined proportion threshold.

[0161] According to an embodiment of the present disclosure, the sample processing module includes: a sub-result determination sub-module, a sub-result processing sub-module, and a result determination sub-module.

[0162] The sub-result determination submodule is used to determine a second reference sub-result of the generative type from other sub-results. The second reference sub-result of the generative type includes at least one of the following: the reason why the answer in the sample question-answer pair is wrong relative to the corresponding question, and prompt information for guiding correct reasoning.

[0163] The sub-result processing sub-module is used to perform generative processing on the second reference sub-result to obtain a target second reference sub-result.

[0164] The result determination submodule is used to obtain the target sample reflection result based on the target second reference sub-result.

[0165] According to an embodiment of the present disclosure, the second reference sub-result includes a plurality of sub-results.

[0166] According to an embodiment of the present disclosure, the sub-result processing submodule includes: a weight determination unit and a fusion unit.

[0167] The weight determination unit is used to determine the weight used to characterize the importance of each of the multiple second reference sub-results based on the effect evaluation values of each of the multiple first large models.

[0168] The fusion unit is used to perform semantic fusion processing on the multiple second reference sub-results based on the respective weights of the multiple second reference sub-results to obtain a target second reference sub-result.

[0169] According to an embodiment of the present disclosure, the sample generating apparatus further includes: an updating module.

[0170] A more detailed module is used to update the content of the reflection sample set to which the test reflection sample belongs when it is determined that there is an error in the test reflection sample. The test reflection sample is screened from the reflection sample set, and the reflection sample set is obtained by classifying multiple reflection samples based on the domain information of the sample question and answer pairs.

[0171] According to an embodiment of the present disclosure, the sample generating apparatus further includes: a result determining module.

[0172] The result determination module is used to determine that the multiple sample reflection results are different when multiple sub-results of any type in the multiple sample reflection results are different.

[0173] Figure 10 A block diagram of a large model training device according to an embodiment of the present disclosure is schematically shown.

[0174] like Figure 10 As shown, the large model training device 1000 includes: a training module 1010.

[0175] The training module 1010 is used to train the second large model using the reflection samples to obtain the reflection large model.

[0176] The reflection sample is generated using the sample generation device provided by the embodiment of the present disclosure.

[0177] Figure 11 The block diagram of the question-answering device according to an embodiment of the present disclosure is schematically shown.

[0178] like Figure 11 As shown, the question-answering device 1100 includes: an acquisition module 1110 , an interaction module 1120 , a reflection module 1130 , a re-interaction module 1140 and a presentation module 1150 .

[0179] The acquisition module 1110 is used to acquire questions input through the interactive interface.

[0180] The interactive module 1120 is used to input questions into the interactive model to obtain initial answers.

[0181] The reflection module 1130 is used to input questions and initial answers into the reflection model to obtain reflection results.

[0182] The re-interaction module 1140 is used to input the question, the initial answer and the reflection result into the interaction model to obtain the target answer when the reflection result indicates that the question and the initial answer are irrelevant.

[0183] The display module 1150 is used to display the target answer on the interactive interface.

[0184] Among them, the reflection big model is trained using the big model training device provided by the embodiment of the present disclosure.

[0185] Figure 12 The structural block diagram of an artificial intelligence agent according to an embodiment of the present disclosure is schematically shown.

[0186] In the embodiments of the present disclosure, Figure 12 As shown, the AI agent 1200 may include an input module 1210 , a processing module 1220 and an output module 1230 .

[0187] Input module 1210, for receiving input information;

[0188] Processing module 1220, configured to determine a target task based on input information received by the input module, determine a large model based on the target task, and obtain output information by invoking the large model to execute the artificial intelligence-based sample generation method and large model training method provided in accordance with embodiments of the present disclosure, or by invoking the large model to execute the question-answering method provided in accordance with embodiments of the present disclosure;

[0189] The output module 1230 is used to output the output information obtained by the processing module.

[0190] According to an embodiment of the present disclosure, the input module 1210 is responsible for receiving or perceiving information such as queries, requests, instructions, signals, or data from the outside world (e.g., a user or the external environment) and converting it into a format that can be understood and processed by the AI agent 1200. The input module 1210 is the primary link for the AI agent 1200 to interact with the outside world. It enables the AI agent 1200 to efficiently and accurately obtain the necessary "sensory" information from the outside world and respond to this information.

[0191] In an example, the input module 1210 may input the sample question-answer pairs, reflection samples or questions described above.

[0192] In this example, processing module 1220 is the core support for AI agent 1200's ability to handle complex tasks. Processing module 1220 can execute the artificial intelligence-based sample generation method and large model training method described above, or execute the question-answering method provided by the embodiment of the present disclosure by calling the large model.

[0193] In this example, the performance of processing module 1220 may be closely related to the large model on which AI agent 1200 is based. To fully utilize the capabilities of the large model, the internal structure of processing module 1220 may be designed to be highly configurable and extensible to cope with various types of tasks and requirements in real-world scenarios.

[0194] In this example, after the AI agent 1200 receives the question, the processing module 1220 can use the interactive large model to process the question, obtain an initial answer, and use the reflection large model to reflect on the correlation between the question and the initial answer to obtain a reflection result. If the reflection result indicates that the question and the answer are related, the initial answer is used as the target answer. If the reflection result indicates that the question and the answer are not related, the question, the initial answer, and the reflection result are input into the interactive large model to obtain the target answer, and the target answer is passed to the output module 1230.

[0195] Understandably, while large language models possess excellent language understanding and generation capabilities, like humans, they are limited in the tasks they can perform without tools. However, when AI agent 1200 is empowered with tool-based capabilities, it can perform tasks such as mathematical calculations using a calculator, data analysis using Python, and weather forecasting using search engines.

[0196] In an example, the output module 1230 may output the aforementioned reflection sample, target answer, or trained large model.

[0197] The AI agent 1200 according to the embodiment of the present disclosure can simply and effectively improve the level of intelligence, and enhance flexibility and versatility.

[0198] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0199] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.

[0200] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.

[0201] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.

[0202] Figure 13A schematic block diagram of an example electronic device 1300 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0203] like Figure 13 As shown, device 1300 includes a computing unit 1301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1302 or a computer program loaded from a storage unit 1308 into a random access memory (RAM) 1303. RAM 1303 may also store various programs and data required for the operation of device 1300. Computing unit 1301, ROM 1302, and RAM 1303 are connected to each other via a bus 1304. An input / output (I / O) interface 1305 is also connected to bus 1304.

[0204] Various components in device 1300 are connected to an input / output (I / O) interface 1305, including an input unit 1306, such as a keyboard and mouse; an output unit 1307, such as various types of displays and speakers; a storage unit 1308, such as a magnetic disk and optical disk; and a communication unit 1309, such as a network card, a modem, a wireless communication transceiver, etc. Communication unit 1309 allows device 1300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0205] Computing unit 1301 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of computing unit 1301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit 1301 performs the various methods and processes described above, such as the AI-based sample generation method, the large model training method, and the question-answering method. For example, in some embodiments, the AI-based sample generation method, the large model training method, and the question-answering method may be implemented as a computer software program tangibly embodied in a machine-readable medium, such as storage unit 1308. In some embodiments, part or all of the computer program may be loaded and / or installed onto device 1300 via ROM 1302 and / or communication unit 1309. When the computer program is loaded into RAM 1303 and executed by computing unit 1301, one or more steps of the AI-based sample generation method, large model training method, and question-answering method described above may be performed. Alternatively, in other embodiments, computing unit 1301 may be configured to execute the AI-based sample generation method, large model training method, and question-answering method in any other appropriate manner (e.g., via firmware).

[0206] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0207] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0208] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0209] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0210] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0211] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0212] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.

[0213] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A sample generation method based on artificial intelligence, comprising: Inputting the sample question-answer pairs into different first models respectively to obtain a plurality of sample reflection results, wherein the sample reflection results include a plurality of sub-results indicating the relevance of the sample question-answer pairs; In the case where the plurality of sample reflection results are different, determining an intermediate sample reflection result from the plurality of sample reflection results, wherein at least one sub-result of the intermediate sample reflection result meets a predetermined screening condition; Processing other sub-results in the intermediate sample reflection result to obtain a target sample reflection result; and A reflection sample is generated based on the sample question-answer pair and the target sample reflection result.

2. The method according to claim 1, wherein Determining an intermediate sample reflection result from the plurality of sample reflection results includes: Determining a plurality of first reference sub-results of a classification type from the plurality of sample reflection results, wherein the first reference sub-results of a classification type include at least one of the following: whether an answer in the sample question-answer pair is incorrect relative to the corresponding question, and the error type if an error exists; Determining a target first reference sub-result that meets the predetermined screening condition from the plurality of first reference sub-results; The sample reflection result including the first target reference sub-result is used as the intermediate sample reflection result.

3. The method according to claim 1, wherein Determining an intermediate sample reflection result from the plurality of sample reflection results includes: determining a target first large model from the plurality of first large models based on respective effect evaluation values of the plurality of first large models; and Determine whether the sub-results in the sample reflection results output by the target first large model meet the predetermined screening conditions to obtain the intermediate sample reflection results.

4. The method according to claim 2 or 3, wherein: The predetermined screening condition includes at least one of the following: The effect evaluation value of the first model is greater than a predetermined evaluation threshold; The ratio of identical sub-results is greater than a predetermined ratio threshold.

5. The method according to any one of claims 1 to 4, wherein The processing of other sub-results in the intermediate sample reflection result to obtain the target sample reflection result includes: Determining a second reference sub-result of a generative type from the other sub-results, the second reference sub-result of the generative type including at least one of the following: a reason why an answer in the sample question-answer pair is incorrect relative to the corresponding question, and prompt information for guiding correct reasoning; performing generative processing on the second reference sub-result to obtain a target second reference sub-result; Based on the target second reference sub-result, the target sample reflection result is obtained.

6. The method according to claim 5, wherein: The second reference sub-results include a plurality of; The generative processing of the second reference sub-result to obtain a target second reference sub-result includes: Determining, based on the respective effect evaluation values of the plurality of first large models, a weight for characterizing the importance of each of the plurality of second reference sub-results; and Based on the respective weights of the plurality of second reference sub-results, semantic fusion processing is performed on the plurality of second reference sub-results to obtain the target second reference sub-result.

7. The method according to claim 1, further comprising: When it is determined that there is an error in the test reflection sample, the content of the reflection sample set to which the test reflection sample belongs is updated, wherein the test reflection sample is screened from the reflection sample set, and the reflection sample set is obtained by classifying multiple reflection samples based on the domain information of the sample question and answer pairs.

8. The method according to any one of claims 1 to 7, further comprising: When a plurality of sub-results of any type in a plurality of the sample introspection results are different, it is determined that the plurality of sample introspection results are different.

9. A large model training method comprising: The second largest model is trained using the reflection samples to obtain the reflection large model; The reflection sample is generated using the sample generation method according to any one of claims 1 to 8.

10. A question-answering method, comprising: Get questions entered through the interactive interface; Input the question into the interactive macro model to obtain an initial answer; Inputting the question and the initial answer into the reflection model to obtain a reflection result; In a case where the reflection result indicates that there is no correlation between the question and the initial answer, inputting the question, the initial answer and the reflection result into the interactive macro model to obtain a target answer; as well as Displaying the target answer on the interactive interface; Wherein, the reflection large model is trained using the large model training method according to claim 9.

11. A sample generation device based on artificial intelligence, comprising: a sample input module, configured to input sample question-answer pairs into different first models to obtain a plurality of sample reflection results, wherein the sample reflection results include a plurality of sub-results indicating the relevance of the sample question-answer pairs; a screening module, configured to, when the plurality of sample reflection results are different, determine an intermediate sample reflection result from the plurality of sample reflection results, wherein at least one sub-result of the intermediate sample reflection result meets a predetermined screening condition; a sample processing module, configured to process other sub-results in the intermediate sample reflection result to obtain a target sample reflection result; and The sample generation module is used to generate a reflection sample based on the sample question-answer pair and the target sample reflection result.

12. A large model training device comprising: A training module is used to train the second largest model using the reflection samples to obtain a reflection large model; The reflection sample is generated by using the sample generating device according to any one of claims 1 to 8.

13. A question-answering device, comprising: The acquisition module is used to obtain questions input through the interactive interface; An interactive module, used to input the question into the interactive model to obtain an initial answer; A reflection module, configured to input the question and the initial answer into a reflection model to obtain a reflection result; a re-interaction module, configured to input the question, the initial answer, and the reflection result into the interaction macro model to obtain a target answer when the reflection result indicates that the question and the initial answer are not correlated; as well as A display module, configured to display the target answer on the interactive interface; Wherein, the reflection large model is trained using the large model training device according to claim 9.

14. An artificial intelligence agent, comprising: An input module, used for receiving input information; a processing module, configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by executing the method according to any one of claims 1 to 10 by calling the large model; An output module is used to output the output information obtained by the processing module.

15. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 10.

16. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 10.

17. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10.

Citation Information

Cited By

  • Text classification method and device, computer equipment and storage medium

    CN120744126A

  • Fine adjustment method of preset model, question and answer method and equipment based on preset model

    CN121009964A

  • Method suitable for character behavior analysis based on multi-modal training model

    CN121708648A