Model training method and device, electronic equipment, chip and storage medium

By introducing a search fragment sequence disruption mechanism and dynamic loss mask training in a large language model, the problem of difficulty in responsive attribution traceability is solved, the model generation efficiency and reliability are optimized, and it is suitable for multi-document information integration scenarios.

CN120448816APending Publication Date: 2025-08-08BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510561350.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The large language model has difficulty in responsive attribution tracing, difficult to integrate multi-document information, poor noise robustness, and the existing technical solutions rely on data customization, making it difficult to migrate and apply to large-scale data, and the generation is time-consuming and verbose, affecting the user experience.

Method used

By introducing a search fragment order disruption mechanism and dynamic loss mask training objective function, the model is forced to learn position-independent but content-related semantic correlation capabilities, optimize model training methods, and improve generation efficiency and reliability.

Benefits of technology

It realizes efficient and reliable answers to model generation in multi-document scenarios, reduces the time-consuming generation, and improves user experience and model generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120448816A_ABST
    Figure CN120448816A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device, electronic equipment, a chip and a storage medium. The method comprises the steps of obtaining input data of a current training round and at least one piece of true value data corresponding to the input data; determining prediction data corresponding to the input data by using the initial model according to the input data; determining predicted traceability data corresponding to the truth value data, the predicted traceability data including predicted user query data and predicted knowledge point data, the predicted user query data being predicted user query data, the user query data being user question data, and the predicted knowledge point data being predicted knowledge point data; the knowledge point data is data related to question data of the user; and training the initial model by using the input data, the truth value data, the prediction data and the prediction traceability data. The model generation efficiency and the reliability of the generated content can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing, and in particular to a model training method, device, electronic device, chip and storage medium. Background Art

[0002] Large Language Models (LLMs) have greatly promoted the development of natural language understanding and generation technologies, achieving performance close to that of humans. However, LLMs still face challenges such as hallucinations and difficulty in attributing responses. Summary of the Invention

[0003] The present disclosure provides a model training method, device, electronic device, chip and storage medium to solve problems in related technologies.

[0004] A first aspect embodiment of the present disclosure proposes a model training method, which includes: obtaining input data of a current training round and true value data corresponding to the input data; determining predicted data corresponding to the input data using an initial model based on the input data; determining predicted traceability data corresponding to the true value data, the predicted traceability data including predicted user query data and predicted knowledge point data, the predicted user query data being predicted user query data, the user query data being user's question data, the predicted knowledge point data being predicted knowledge point data, and the knowledge point data being data related to user's question data; and training the initial model using the input data, true value data, predicted data, and predicted traceability data.

[0005] In some embodiments of the present disclosure, the method also includes: obtaining user query data, system prompt data and at least one knowledge point data; splicing the user query data, system prompt data and at least one knowledge point data to obtain first spliced data; performing a reordering operation on the first spliced data to obtain input data.

[0006] In some embodiments of the present disclosure, performing a reordering operation on the first spliced data to obtain input data includes: determining a random number corresponding to the first spliced data; when the random number belongs to a first range, performing random reordering on the label of at least one knowledge point data to obtain input data; when the random number belongs to a second range, performing random reordering on the position of at least one knowledge point data in the first spliced data to obtain input data; when the random number belongs to a third range, performing random reordering on the label of at least one knowledge point data and the position of at least one knowledge point data in the first spliced data to obtain input data.

[0007] In some embodiments of the present disclosure, determining the predicted traceability data corresponding to the true value data includes: determining the first knowledge point data corresponding to the true value data in at least one knowledge point data based on the label data in the true value data; determining the position of the first knowledge point data as the first position, and determining the position of the user query data as the second position; generating predicted knowledge point data based on the first position and the input data; and generating predicted user query data based on the second position and the input data.

[0008] In some embodiments of the present disclosure, the method further includes: performing a demasking process on the first position and the second position.

[0009] In some embodiments of the present disclosure, input data, true value data, predicted data, and predicted traceability data are used to train an initial model, including: determining a first loss function corresponding to the initial model in the current training round based on the difference between the input data and the predicted traceability data, and the difference between the true value data and the predicted data; determining the initial model as the target model when the first loss function meets the convergence condition; and performing the next round of training on the initial model when the first loss function does not meet the convergence condition, until the loss function of the initial model meets the convergence condition.

[0010] In some embodiments of the present disclosure, determining the first loss function corresponding to the current training round based on the difference between the input data and the predicted traceability data, and the difference between the true value data and the predicted data includes: determining the difference between the predicted knowledge point data and the first knowledge point data as the first difference; determining the difference between the predicted user query data and the user query data in the input data as the second difference; determining the difference between the true value data and the predicted data as the third difference; and determining the first loss function based on the first difference, the second difference, and the third difference.

[0011] The second aspect embodiment of the present disclosure proposes a model training device, which includes: a data processing module for obtaining input data of the current training round and true value data corresponding to the input data; a training module for determining, based on the input data, predicted data corresponding to the input data using an initial model; determining predicted traceability data corresponding to the true value data, the predicted traceability data including predicted user query data and predicted knowledge point data, the predicted user query data being predicted user query data, the user query data being the user's question data, the predicted knowledge point data being predicted knowledge point data, and the knowledge point data being data related to the user's question data; and using the input data, true value data, predicted data, and predicted traceability data to train the initial model.

[0012] The third aspect embodiment of the present disclosure proposes an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in the first aspect embodiment of the present disclosure.

[0013] The fourth aspect embodiment of the present disclosure proposes a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in the first aspect embodiment of the present disclosure.

[0014] The fifth aspect embodiment of the present disclosure proposes a chip, characterized in that it includes at least one processor and a communication interface; the communication interface is used to receive signals input into the chip or signals output from the chip, and the processor communicates with the communication interface and implements the method described in the first aspect embodiment of the present disclosure through logic circuits or executing code instructions.

[0015] In summary, the model training method proposed in the present disclosure can determine the predicted data and the predicted traceability data based on the input data and the true value data corresponding to the input data. After that, the initial model can be trained based on the input data, true value data, predicted data, and predicted traceability data. This can enable the model to deeply learn the correspondence between the user's inquiries, quoted fragments, and answers, improve the reliability of the model-generated content, and optimize the efficiency of the model-generated content.

[0016] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0018] Figure 1 A flow chart of a model training method provided in an embodiment of the present disclosure Figure 1 ;

[0019] Figure 2 A flow chart of a model training method provided in an embodiment of the present disclosure Figure 2 ;

[0020] Figure 3 A flow chart of a model training method provided in an embodiment of the present disclosure Figure 3 ;

[0021] Figure 4A A flowchart of a data processing method provided in an embodiment of the present disclosure;

[0022] Figure 4B A flowchart of a model training method based on dynamic loss masking provided in an embodiment of the present disclosure;

[0023] Figure 5 A schematic diagram of the structure of a model training device provided in an embodiment of the present disclosure;

[0024] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure;

[0025] Figure 7 A schematic diagram of the chip structure provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0026] The following describes in detail embodiments of the present disclosure, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0027] Large language models (LLMs) have significantly advanced the development of natural language understanding and generation technologies, achieving near-human-level performance. However, LLMs still face challenges such as hallucinations and difficulty attributing responses. Retrieval-Augmented Generation (RAG) technology integrates the inherent knowledge of large language models with vast, dynamic external database resources by injecting externally retrieved knowledge into the model generation phase. This significantly improves the accuracy and credibility of generated content, while supporting continuous knowledge updates and the integration of domain-specific information, making it more suitable for knowledge-intensive tasks.

[0028] However, current RAG models face challenges integrating multi-document information and poor noise robustness. Furthermore, model responses often lack citations to reliable sources, resulting in the model being unable to provide complete and reliable answers to questions in complex RAG scenarios. Therefore, optimizing model information integration, noise resistance, and multi-source traceability capabilities is a key challenge in RAG research.

[0029] In one approach, the task objectives are usually broken down, a dedicated dataset is constructed for the objective, and reinforcement learning methods such as Supervised Fine Tuning (SFT) or Direct Preference Optimization (DPO) are combined to train and optimize the model on the dedicated dataset.

[0030] In this approach, based on user questions and search snippets, the model is guided to generate responses consisting of a reference to the original text [GROUNDING] and an answer to the question [ANSWER], forming dedicated training data. Training is divided into two phases: 1) First, the model is fine-tuned using this data, allowing it to learn the [GROUNDING] + [ANSWER] response format. 2) A small-parameter model is used to generate low-quality data with poor relevance between the original text and the question answer. This data is then paired with the high-quality data to form preference pairs. The Direct Preference Optimization (DPO) algorithm is then used to train the model on this preference data, allowing it to learn how to provide reliable responses and cite sources.

[0031] Another approach constructs specialized data (training data for ReferGeneration and ClaimGeneration) to train two models: ReferGeneration generates the original reference based on the user question and retrieval fragment; and ClaimGeneration generates the answer based on the preceding input. During the final generation process, these two models are used alternately, and the input to each model is adjusted accordingly to obtain the final answer, thus improving the RAG system's response verifiability and credibility.

[0032] In another solution, in order to enhance the robustness of the model and the order-independence of the retrieval segments, the parallelization of retrieval segments (Chunks) is naturally achieved at the model structure level. However, the change in model structure requires a lot of engineering adaptation in actual deployment, and there is still some distance before it can be put into practice and applied to RAG generation.

[0033] The above solution relies too much on data customization, and when the model generates a response, it needs to first generate the original text and then generate the answer to the question. This approach has the following problems:

[0034] 1) These solutions typically require rebuilding training data, involving the construction of large amounts of proprietary data. This is costly, and the data is not universally applicable, making it difficult to migrate to open source data or new domains. They are also difficult to apply to large-scale data model training, resulting in difficulty in generalization and hindering the rapid updating and iteration of data and models.

[0035] 2) To improve the model's information integration and answer traceability capabilities, the aforementioned methods all choose to first output the original text citations in the search fragment during the generation phase, and then answer the question based on these citations. However, in RAG scenarios, especially when integrating multiple documents for question-answering, the model needs to output a large number of original text citations when generating answers, which is extremely time-consuming and not conducive to online deployment in production environments. Furthermore, the inclusion of original text citations makes responses lengthy, making them difficult for users to understand and read, and thus impacting the user experience.

[0036] 3) It mainly focuses on data construction research and improvement, and lacks improvements in model training methods.

[0037] Furthermore, most existing technical solutions only focus on the impact of retrieval segment ordering on generated results during the inference phase, while ignoring the impact of retrieval segment position during training. This results in a single segment position ordering in the training data (e.g., all correct answers come from the Nth fixed segment). This causes the model to misfit the relationship between document position and source-based answers, reducing the model's source-based capabilities.

[0038] Therefore, to address the above issues, this paper proposes a model training method that introduces a scrambling mechanism for the order of retrieval fragments (Chunks) during the data processing phase. By randomly arranging the input order of Chunks to construct training data, the model is forced to learn semantic association capabilities that are position-independent but content-related. The model needs to extract key information and generate structured answers through global semantic understanding rather than relying on local order. By introducing a training objective function with a dynamic loss mask, the model training method is improved, allowing the model to learn to correctly index the original text, the relationship between the query and the answer through loss constraints. Compared with learning the mapping relationship by generating correct references, this significantly improves the efficiency of model generation.

[0039] The specific contents of this method are as follows.

[0040] Figure 1 A flow chart of a model training method provided in an embodiment of the present disclosure Figure 1 .like Figure 1 As shown, the method may include the following steps.

[0041] Step 101: Obtain input data for the current training round and true value data corresponding to the input data.

[0042] In some embodiments, the current training round may be a round for training the initial model, wherein each training round may use input data corresponding to the current training round, that is, the input data used in the current round, and the input data may be a piece of data in a preset training data set. After obtaining the input data, the true value data corresponding to the input data may be obtained.

[0043] In some embodiments, optionally, the initial model may be, for example, a generative model, such as a question-answering model, a RAG model, etc., wherein the RAG question-answering scenario may be in the form of text question-answering, image document question-answering, code, etc. Wherein, when the initial model is a question-answering model, the input data may include user query data (query), knowledge point data, system prompt data, etc., wherein the user query data may be data of a user's question, data of the user's query, etc., such as "how is the weather today", the name of the knowledge point data may be reference data, retrieval data, reference fragment, retrieval fragment, chunk fragment, etc., and the knowledge point data is data related to the data of the user's question. For example, when the user query data is "how is the weather today", the knowledge point data may be today's meteorological information, such as today's highest and lowest temperatures, today's weather type (such as cloudy, sunny, rainy, etc.), and today's humidity, ultraviolet intensity, etc., that is, one user query data may correspond to at least one knowledge point data, at least one knowledge point data is related to the user query data, and the relevance of at least one knowledge point data to the user query data may be different.

[0044] In some embodiments, the system prompt data is related to the user query data, wherein the system prompt data is used to guide the initial model understanding task. For example, the system prompt data can be used to inform the model that the current task is to generate answer text based on the user query data and knowledge point data, or to provide explanations, etc., or the system prompt can also be used to indicate the format and specifications of the model output data, etc.

[0045] In some embodiments, the true value data may include answer data and label data. For example, the true value data may be represented as (<response,label:1> ), wherein the answer data (response) is the correct response data or answer data corresponding to the user query data determined based on at least one knowledge point data. For example, the true value data can be the data of the correct answer corresponding to the user query data, that is, the true value data can be the correct answer or the standard answer. For example, if the user query data is "How is the weather today?", and the knowledge point data is "Today's weather is sunny", "Today's maximum temperature is 25 degrees Celsius, and today's minimum temperature is 15 degrees Celsius", then the true value data can be "Today's weather is sunny, and today's temperature is 15 to 25 degrees Celsius". Optionally, at least one knowledge point data can be a reference fragment, and multiple reference fragments may include some of the same content. For example, reference fragment 1 is "Today's weather is sunny, and today's maximum temperature is 25 degrees Celsius", and reference fragment 2 is "Today's weather is sunny, and today's minimum temperature is 15 degrees Celsius".

[0046] In some embodiments, label data (label) is used to indicate the source of the answer data, that is, the name of the label data can also be a traceability label, a reference label, a source label, etc. Optionally, the source of the answer data is at least one knowledge point data, so the label data can be used to indicate at least one knowledge point, and the label data can indicate the index of at least one knowledge point, that is, each knowledge point will have an index label or index parameter, such as ref: 1. For example, when the label data is label: 1, it indicates that the source corresponding to the answer data is the knowledge point data corresponding to ref: 1.

[0047] Step 102: Based on the input data, use the initial model to determine the predicted data corresponding to the input data.

[0048] In some embodiments, after receiving the input data, the initial model can be used to make predictions to obtain predicted data corresponding to the input data. The predicted data can be an answer generated based on the user query data and at least one knowledge point data in the input data, that is, the initial model can generate an answer to the user query data based on the user query data and at least one knowledge point data in the input data. The predicted data is predicted. Before the initial model is trained, there may be differences between the predicted data and the true data. For example, for the example of step 101, the predicted data generated by the initial model can be, for example, "Today's weather is sunny, and today's temperature is 16 to 25 degrees Celsius." Therefore, in the model training project, the true data is the preset correct answer, and the predicted data is the predicted answer. There may be differences between the predicted data and the true data.

[0049] Step 103: Determine the predicted traceability data corresponding to the true value data.

[0050] In some embodiments, the predicted tracing data includes predicted user query data and predicted knowledge point data. The predicted user query data is the predicted user query data, the user query data is the user's question data, the predicted knowledge point data is the predicted knowledge point data, and the knowledge point data is data related to the user's question data.

[0051] In other words, the predicted traceability data is obtained by continuing to predict the initial model, and the predicted traceability data includes predicted user query data and predicted knowledge point data, among which the predicted user query data and predicted knowledge point data are generated by the initial model based on the input data. That is, the initial model can generate predicted user query data and predicted knowledge point data based on the input data after receiving the input data. During the generation process, since the model has not been trained, there may be errors, that is, there may be differences between the generated predicted user query data and predicted knowledge point data and the user query data and knowledge point data in the user input data.

[0052] In some embodiments, determining the predicted traceability data corresponding to the true value data includes: determining the first knowledge point data corresponding to the true value data in at least one knowledge point data based on the label data in the true value data; determining the position of the first knowledge point data as the first position, and determining the position of the user query data as the second position; generating predicted knowledge point data based on the first position and the input data; and generating predicted user query data based on the second position and the input data.

[0053] In some embodiments, the location of the user query data in the input data can be determined. For example, the location of the user query data in the input data can be determined. For example, the user query data can have a special identification bit (such as a special token), and the initial model can determine the location of the user query data based on the identification bit, or the location of the user query data can be a preset location.

[0054] In some embodiments, the first knowledge point data corresponding to the true value data in at least one knowledge point data can be determined based on the label data in the true value data. The first knowledge point data is the traceability data of the predicted data, that is, the true value data is generated based on the content of the first knowledge point data. The first knowledge point data is the reference data of the true value data and the first knowledge point data is the source of the true value data.

[0055] In some embodiments, after obtaining the true value data, the initial model can determine the traceability data corresponding to the answer data based on the label data in the true value data, that is, the first knowledge point data corresponding to the answer data. For example, when the label data is label: 1, the source corresponding to the answer data is identified as the knowledge point data corresponding to ref: 1, then the knowledge point data corresponding to ref: 1 is the first knowledge point data. After determining the first knowledge point data, the location of the first knowledge point data can be determined, and then the predicted knowledge point data can be generated based on the location of the first knowledge point data and the input data.

[0056] Step 104: Use the input data, true value data, predicted data, and predicted traceability data to train the initial model.

[0057] In some embodiments, after determining the input data, true value data, predicted data, and predicted traceability data, the loss function of the initial model of the current training round can be determined based on the difference between the input data and the predicted traceability, and the difference between the true value data and the predicted data. Then, it can be determined whether the convergence condition is met based on the loss function. For example, when the loss function meets the convergence condition, it can be determined that the initial model has achieved the training purpose, and the performance of the initial model is better at this time. Therefore, the training of the initial model is completed, and the initial model is deployed as the target model. When it is determined that the loss function of the initial model in the current training round does not meet the convergence condition, the initial model can be trained for the next round until the loss function of the initial model meets the convergence condition.

[0058] In some embodiments, when the loss function does not meet the convergence conditions, the model can be adjusted according to the loss function. For example, the parameters of the initial model can be adjusted, and then the adjusted initial model can be trained for the next round until the loss function of the initial function meets the convergence conditions.

[0059] Alternatively, the number of iterations, that is, the number of training times, can be preset. After determining the loss function corresponding to the current training round, the model can be adjusted according to the loss function, and then the value of the current training round can be judged. When the value of the current training round reaches the preset number of iterations, it can be determined that the model training is completed, and the adjusted initial model is used as the target model. Otherwise, when the preset number of iterations is not reached, the adjusted initial model can continue to be trained for the next round until the value of the training round reaches the preset number of iterations.

[0060] In summary, the above embodiments of the present disclosure can determine the predicted data and the predicted traceability data based on the input data and the true value data corresponding to the input data. Afterwards, the initial model can be trained based on the input data, the true value data, the predicted data, and the predicted traceability data. When calculating the loss function, by using the difference between the predicted user query data and the user query data in the input data, as well as the difference between the predicted knowledge point data and the knowledge point data in the input data, the trained model can more accurately locate the user query and the knowledge point data corresponding to the user query, and can enable the model to deeply learn the correspondence between the user's query, the quoted fragment and the answer, thereby improving the reliability of the model-generated content and optimizing the efficiency of the model-generated content.

[0061] Figure 2 A flow chart of a model training method provided in an embodiment of the present disclosure Figure 2 .like Figure 2 As shown, based on Figure 1 In the illustrated embodiment, the method includes the following steps.

[0062] Step 201: Obtain user query data, system prompt data, and at least one piece of knowledge point data.

[0063] In some embodiments, user query data, system prompt data, and at least one knowledge point data may be obtained, wherein the obtained user query data, system prompt data, and at least one knowledge point data may be training data in a training data set.

[0064] Step 202 : splicing the user query data, the system prompt data, and at least one knowledge point data to obtain first spliced data.

[0065] In some embodiments, user query data, system prompt data and at least one knowledge point data can be spliced to obtain first spliced data. Optionally, the query data, system prompt data and at least one knowledge point data can be organically integrated according to a preset format to obtain first spliced data.

[0066] Step 203: Perform a reordering operation on the first spliced data to obtain input data.

[0067] In some embodiments, since in the training data set, for a user query data, at least one knowledge point data corresponding to the user query data is arranged in a certain order, for example, in the training data set, multiple knowledge point data corresponding to a user query data can be sorted from high to low or from low to high according to the relevance to the user query, or can be sorted from high to low or from low to high according to the richness of the content, therefore, the multiple knowledge point data in the generated first spliced data may also be sorted according to the sorting method in the training data set, for example, the multiple knowledge point data in the first spliced data can be sorted from high to low or from low to high according to the relevance to the user query, or can be sorted from high to low or from low to high according to the richness of the content, and the index value corresponding to each knowledge point may also be arranged in sequence. This arrangement method may result in a single fragment position sorting in the training data, and the model may incorrectly fit the relationship between the document position and the traceability answer after training, resulting in a decrease in the model's traceability ability.

[0068] For example, for a first spliced data, the first spliced data includes 5 knowledge point data, namely knowledge point 1, knowledge point 2, knowledge point 3, knowledge point 4, and knowledge point 5, wherein the relevance of the 5 knowledge point data to the user query data is arranged from high to low as knowledge point 1, knowledge point 2, knowledge point 3, knowledge point 4, and knowledge point 5, wherein the index of knowledge point 1 is index 1, the index of knowledge point 2 is index 2, the index of knowledge point 3 is index 3, the index of knowledge point 4 is index 4, and the index of knowledge point 5 is index 5. At this time, in the first spliced data, a possible knowledge point sorting method from left to right is knowledge point 1, knowledge point 2, knowledge point 3, knowledge point 4, and knowledge point 5. Assuming that there are N first spliced data, and the sorting method of multiple knowledge point data in each of the N first spliced data is the above-mentioned sorting method from high to low according to the relevance of the user query data, then the N first spliced data The knowledge point data with the highest similarity in a spliced data are all arranged first. If the model is directly trained using N first spliced data, then during the training process, the first knowledge point data corresponding to the true value data may be the knowledge point 1 corresponding to index 1, or the first knowledge point data corresponding to the true value data may be the first data in the first spliced data. At this time, the model may mistakenly believe that the determination of the true value data is related to the location of multiple knowledge points. For example, it may mistakenly believe that the true value data must be determined based on the first knowledge point data in the first spliced data, that is, the correct answers all come from the Nth fixed segment, causing the model to incorrectly fit the relationship between the document position and the traceability answer. The model may ignore other knowledge point data. At this time, if the first knowledge point data in the first spliced data is wrong data or irrelevant data, it will cause the final output answer to be wrong or have poor satisfaction.

[0069] In some embodiments, a reordering operation can be performed on the first spliced data to make the data and / or index values of multiple knowledge points in the first spliced data consistent with the specific content of the knowledge points. The model is trained based on the reordered first spliced data, which can avoid the model from relying on the position or index of the knowledge point data to generate answers. The model can better learn the relationship between the user query and the specific content of the knowledge point to obtain more reliable and accurate answers, thereby improving user satisfaction. The reordering operation is performed on the first spliced data to obtain input data, which includes: determining a random number corresponding to the first spliced data; when the random number belongs to the first range, performing random reordering on the label of at least one knowledge point data to obtain input data; when the random number belongs to the second range, performing random reordering on the position of at least one knowledge point data in the first spliced data to obtain input data; when the random number belongs to the third range, performing random reordering on the label of at least one knowledge point data and the position of at least one knowledge point data in the first spliced data to obtain input data.

[0070] In some embodiments, a random number corresponding to the first spliced data can be generated. When the value of the random number belongs to the first range, the label of at least one knowledge point data can be randomly reordered to obtain input data, wherein the random reordering of the label of at least one knowledge point data can be a random reordering of the index value corresponding to the at least one knowledge point data, that is, the index value of at least one knowledge point data can be randomly exchanged. At this time, the positions of multiple knowledge point data in the first spliced data can remain unchanged. For example, multiple knowledge point data are still sorted according to similarity in the first spliced data, but at this time, the labels (that is, index values) of multiple knowledge point data change. Using this scheme to randomly shuffle the data can make the trained model independent of the index values of multiple knowledge point data when determining the knowledge point data related to the user query.

[0071] In some embodiments, when the value of the random number belongs to the second range, the position of at least one knowledge point data in the first spliced data can be randomly reordered to obtain input data. At this time, the sorting method of multiple knowledge point data in the first spliced data can be randomly disrupted. At this time, the index value can be related to the position of the knowledge point data in the first spliced data. For example, the first spliced data includes 5 knowledge point data, namely knowledge point 1, knowledge point 2, knowledge point 3, knowledge point 4, and knowledge point 5. The relevance of the 5 knowledge point data to the user query data is arranged from high to low as knowledge point 1, knowledge point 2, knowledge point 3, knowledge point 4, and knowledge point 5, where the index of knowledge point 1 is index 1, the index of knowledge point 2 is index 2, the index of knowledge point 3 is index 3, and the index of knowledge point 4 is index The index of knowledge point 4 and the index of knowledge point 5 are index 5. When the positions of multiple knowledge point data are shuffled, for example, the order of multiple knowledge point data after shuffling is knowledge point 2, knowledge point 5, knowledge point 3, knowledge point 1, and knowledge point 4, then index 1 is changed to the index value of the first knowledge point data in the first spliced data, that is, the index value of knowledge point 2 is index 1, and index 2 is changed to the index value of the second knowledge point data in the first spliced data, that is, the index value of knowledge point 5 is index 2, and so on. The index value of knowledge point 3 is index 3, the index value of knowledge point 1 is index 4, and the index value of knowledge point 4 is index 5. Using this method for shuffling can make the trained model independent of the index value of the knowledge point data and the position of the knowledge point data in the first spliced data when generating answers.

[0072] In some embodiments, when the random number belongs to the third range, the label of at least one knowledge point data and the position of at least one knowledge point data in the first spliced data are randomly reordered to obtain input data. That is, when the random number belongs to the third range, the label of at least one knowledge point data and the position of at least one knowledge point data in the first spliced data can be randomly reordered at the same time. Using this method for shuffling can make the trained model independent of the index value of the knowledge point data and the position of the knowledge point data in the first spliced data when generating answers.

[0073] In other words, the specific sequencing method used for the first splicing data can be determined based on the range of the random number of the first splicing data. Optionally, for round training, when there are multiple first splicing data, each first splicing data can be made to execute one of the above three sequencing methods with a certain probability. Optionally, the first range, the second range and the third range can be preset and can be determined based on the actual application scenario.

[0074] In summary, the above-mentioned embodiments of the present application can randomly shuffle the labels and / or positions of multiple knowledge point data, so that the trained model does not rely on the index value of the knowledge point data and the position of the knowledge point data in the first spliced data when generating answers. It can enable the model to deeply learn the correspondence between the user's inquiries, quoted fragments and answers, improve the reliability of the model's generated content, and optimize the efficiency of the model's generated content.

[0075] Figure 3 A flow chart of a model training method provided in an embodiment of the present disclosure Figure 3 .like Figure 3 As shown, based on Figure 1 In the illustrated embodiment, the method includes the following steps.

[0076] Step 301: Determine the first loss function corresponding to the initial model in the current training round based on the difference between the input data and the predicted traceability data, and the difference between the true data and the predicted data.

[0077] In some embodiments, the method further includes: performing unmasking processing on the first position and the second position. In other words, after determining the position of the first knowledge point data and the position of the predicted query user data, the position of the first knowledge point data and the position of the predicted query user data can be unmasked, wherein the unmasking processing can be removing the loss mask of the first position and the second position. After the unmasking processing, the first knowledge point data and the predicted query user data can participate in the calculation of the loss function.

[0078] In some embodiments, based on the difference between the input data and the predicted traceability data, and the difference between the true value data and the predicted data, determining the first loss function corresponding to the current training round includes: determining the difference between the predicted knowledge point data and the first knowledge point data as the first difference; determining the difference between the predicted user query data and the user query data in the input data as the second difference; determining the difference between the true value data and the predicted data as the third difference; and determining the first loss function based on the first difference, the second difference, and the third difference.

[0079] In some embodiments, the difference between the predicted knowledge point data and the first knowledge point data may be the difference between the specific content of the predicted knowledge point data and the first knowledge point data, that is, the difference between the predicted knowledge point text and the first knowledge point text, the difference between the semantics of the predicted knowledge point data and the semantics of the first knowledge point data, etc. For example, the difference between the predicted knowledge point data and the first knowledge point data may be determined by using edit distance, longest common subsequence, similarity algorithm, or hybrid algorithm, which is not limited in this disclosure.

[0080] In some embodiments, the difference between the predicted user query data and the user query data in the input data may be the difference between the specific content of the predicted user query data and the user query data in the input data, that is, the difference between the predicted user query text and the user query text in the input data, the difference between the semantics of the predicted user query data and the semantics of the user query data in the input data, and so on. For example, the difference between the predicted user query data and the user query data in the input data may be determined by using an edit distance, a longest common subsequence, a similarity algorithm, or a hybrid algorithm, and the present disclosure is not limited to this.

[0081] In some embodiments, the difference between the true data and the predicted data may be the difference between the specific contents of the true data and the predicted data, that is, the difference between the text of the true data and the text of the predicted data, the difference between the semantics of the true data and the semantics of the predicted data, etc. For example, the difference between the true data and the predicted data may be determined by using edit distance, longest common subsequence, similarity algorithm, or hybrid algorithm, which is not limited in this disclosure.

[0082] In some embodiments, the first difference, the second difference, and the third difference can be used to determine the loss function. In other words, when calculating the loss function, the difference between the user query data in the input data and the predicted user query data, the difference between the first knowledge point data corresponding to the true value data and the predicted knowledge point data, and the difference between the true value data and the predicted data can be used to determine the loss function, wherein the user query data in the input data, the predicted user query data, the first knowledge point data, and the predicted knowledge point data are input type data, and the true value data and the predicted data are output type data. That is, when calculating the loss function, the use of input type data to calculate the loss function is newly added, which can enable the model to better learn the correspondence between input data and output type data, so that the reliability and accuracy of the output answer of the model when it is used are higher. The loss function can be expressed as follows:

[0083]

[0084] in, is the loss function, u represents the input data, u ref is the correct original citation of the input, u query is the input user query, r represents the answer, d represents a single training data, D is the training data set, θ is a parameter used to indicate the current number of training times. Specifically, the loss function can be, for example, mean square error, mean absolute error, cross entropy loss, etc., which is not limited in this disclosure.

[0085] Step 302: When the first loss function satisfies the convergence condition, the initial model is determined to be the target model.

[0086] In some embodiments, when the first loss function meets the convergence condition, it means that the values of the first difference, the second difference and the third difference in the current training round are small, that is, the predicted user query data, predicted knowledge point data and predicted data generated by the initial model meet the expected goals. At this time, there is no need to adjust the initial model, that is, stop training the initial model, and the initial model can be determined as the target model.

[0087] Step 303: If the first loss function does not meet the convergence condition, the initial model is trained in the next round until the loss function of the initial model meets the convergence condition.

[0088] In some embodiments, when the first loss function does not meet the convergence condition, the initial model can be trained for the next round until the loss function of the initial model meets the convergence condition. Before the next round of training of the initial model, the initial model can be adjusted, that is, the initial model can be adjusted according to the loss function. For example, the parameters of the initial function can be adjusted. For example, the initial model can be optimized by methods such as gradient optimization, regularization, loss function design, and training strategies to optimize model performance and reduce model generation loss. The specific adjustment method is not limited in this disclosure.

[0089] In summary, the above embodiments of the present disclosure can add the use of input type data to calculate the loss function when calculating the loss function, which can enable the model to better learn the correspondence between input data and output type data, so that the reliability and accuracy of the output answers of the model will be higher when used.

[0090] The technical solution of the present disclosure is further described in detail below in conjunction with specific application examples.

[0091] The following is a RAG generation traceability model optimization method based on dynamic mask provided by an embodiment of the present disclosure. The method has high versatility and does not require the generation of proprietary training data; the generation efficiency is high, and the generation stage does not require the generation of the pre-referenced original text, and the answer is more readable. It can be applied to optimize the single-modal and multi-modal RAG generation models in the following scenarios: 1. Single-document single-source RAG question-and-answer scenario; 2. Multi-document multi-source RAG question-and-answer scenario (including question-and-answer scenarios with refusal to answer or no traceability labels); 3. In complex information retrieval and generation tasks that require the model to have high robustness and fine-grained traceability capabilities. The RAG question-and-answer scenario can be in the form of text question-and-answer, image document question-and-answer, code, etc.

[0092] This solution proposes a general RAG generation and provenance optimization process, which includes data processing methods and model training methods. First, through a data processing method, we can transform any generated traceability data. The transformed data can force the model to learn semantic association capabilities that are position-independent but content-related. For the model version that has not transformed the data, if the order of the retrieved chunks is not ideal (such as mixed steps, reverse time order), the model may generate logically confusing or wrong answers. After completing the data processing, we introduce a loss function based on a dynamic loss mask in the model training process to constrain the model to learn the semantic mapping relationship between the retrieval fragment (original text) containing the correct answer, the answer, and the original text traceability label, and continue training until the model loss converges.

[0093] 1. Reference location-independent data processing method

[0094] like Figure 4AAs shown in Figure 1, the model's input consists of three parts: system prompts, knowledge points (retrieval fragments or chunks), and user queries. These three are organically integrated according to a specific format to form the model input. The model's output is the correct answer and a traceability label.

[0095] In the data processing phase, the solution introduces three data processing methods:

[0096] 1. Randomly shuffle and rearrange the source tags ([ref_n]) of the reference fragments, such as Figure 4A As shown in the second data of , the traceability labels ([ref_n]) of the reference fragments are randomly shuffled and rearranged without changing the positions of the reference fragments in the original data;

[0097] 2. Randomly shuffle and rearrange the reference fragment positions, such as Figure 4A As shown in the third data of , the traceability label ([ref_n]) of the reference fragment is not changed, and the position of the reference fragment in the original data is randomly shuffled and rearranged;

[0098] 3. Such as Figure 4A As shown in the fourth data, the reference fragment position and traceability tag serial number are shuffled at the same time.

[0099] Generate a random number corresponding to each piece of data, and according to the range of the random number, make each piece of data enter one of the three processing methods with a certain probability.

[0100] 2. Model training method based on dynamic loss mask

[0101] This method mainly optimizes the loss function of traditional instruction fine-tuning training. Specifically, Figure 4B For example:

[0102] 1. First, based on the answer (<response,label:1> ), dynamically locate the user query (query) of the input part, the location of the traceability tag and the location of the retrieval original text where the answer to the question is located.

[0103] 2. Next, we remove (unmask) the loss mask (loss mask) at the corresponding position. During training, this part of the text (token) will be used together with the answer token (<response,label:1> ) to calculate the losses together.

[0104] 3. Use this method to train all data until the model loss converges.

[0105]

[0106] Among them, u represents the input data, u refis the correct original citation of the input, u query is the input user query, r represents the answer, d represents a single training data, D is the training data set, and θ is a parameter used to indicate the current number of training times.

[0107] In summary, the method for constructing training data that is independent of the reference position in the above-mentioned examples disclosed herein is simple and efficient, does not involve the generation of additional data content, and can be applied to any existing RAG traceability data. The training data constructed based on this method is more robust to the position of the traceability label, which can further improve the accuracy of the model traceability. The model training method based on the dynamic loss mask in the above-mentioned examples disclosed herein helps the model learn the mapping relationship between the reference original text and the traceability answer by adding a dynamic loss mask. This solution can avoid the explicit generation of reference fragments while ensuring the improvement of the generated traceability effect, greatly reduce the generation redundancy, improve the efficiency output of the RAG model, and improve the output readability.

[0108] The model obtained by training the above examples disclosed in the present invention avoids the generation of correct citations of the original text, greatly reduces the generation time, and improves the conciseness of the reply; and the model achieves optimal results in multi-information integration, redundant information noise resistance and traceability accuracy evaluation, achieving a dual improvement in RAG generation and traceability capabilities, with stronger robustness, no reliance on specific model structure, and no complex interactive logic, and can be quickly applied.

[0109] Figure 5 Schematic diagram of the structure of a model training device 500 provided in an embodiment of the present disclosure. Figure 5 As shown, the device includes: a data processing module 510, which is used to obtain the input data of the current training round and the true value data corresponding to the input data; a training module 520, which is used to determine the predicted data corresponding to the input data based on the input data using the initial model; determine the predicted traceability data corresponding to the true value data, the predicted traceability data including predicted user query data and predicted knowledge point data, the predicted user query data is the predicted user query data, the user query data is the user's question data, the predicted knowledge point data is the predicted knowledge point data, and the knowledge point data is the data related to the user's question data; the input data, true value data, predicted data, and predicted traceability data are used to train the initial model.

[0110] In some embodiments, the data processing module is also used to obtain user query data, system prompt data and at least one knowledge point data; splice the user query data, system prompt data and at least one knowledge point data to obtain first spliced data; perform a reordering operation on the first spliced data to obtain input data.

[0111] In some embodiments, the data processing module is also used to perform a reordering operation on the first spliced data to obtain input data, including: determining a random number corresponding to the first spliced data; when the random number belongs to a first range, performing random reordering on the label of at least one knowledge point data to obtain input data; when the random number belongs to a second range, performing random reordering on the position of at least one knowledge point data in the first spliced data to obtain input data; when the random number belongs to a third range, performing random reordering on the label of at least one knowledge point data and the position of at least one knowledge point data in the first spliced data to obtain input data.

[0112] In some embodiments, the training module is also used to determine the predicted traceability data corresponding to the true value data, including: determining the first knowledge point data corresponding to the true value data in at least one knowledge point data based on the label data in the true value data; determining the position of the first knowledge point data as the first position, and determining the position of the user query data as the second position; generating predicted knowledge point data based on the first position and the input data; generating predicted user query data based on the second position and the input data.

[0113] In some embodiments, the model training device further includes a processing module for performing demasking processing on the first position and the second position.

[0114] In some embodiments, the training module is also used to determine the first loss function corresponding to the initial model in the current training round based on the difference between the input data and the predicted traceability data, and the difference between the true data and the predicted data; when the first loss function meets the convergence condition, the initial model is determined to be the target model; when the first loss function does not meet the convergence condition, the initial model is trained for the next round until the loss function of the initial model meets the convergence condition.

[0115] In some embodiments, the training module is also used to determine the difference between the predicted knowledge point data and the first knowledge point data as the first difference; determine the difference between the predicted user query data and the user query data in the input data as the second difference; determine the difference between the true value data and the predicted data as the third difference; and determine the first loss function based on the first difference, the second difference and the third difference.

[0116] In summary, the model training device 500 can determine the predicted data and the predicted traceability data based on the input data and the true value data corresponding to the input data. Afterwards, the initial model can be trained based on the input data, the true value data, the predicted data, and the predicted traceability data. This can enable the model to deeply learn the correspondence between the user's inquiries, quoted fragments, and answers, improve the reliability of the model-generated content, and optimize the efficiency of the model-generated content.

[0117] In the embodiments provided above, the methods and devices provided in the embodiments of the present application are introduced. In order to implement the various functions of the methods provided in the embodiments of the present application, the electronic device may include a hardware structure and a software module, and implement the aforementioned functions in the form of a hardware structure, a software module, or a hardware structure plus a software module. One of the aforementioned functions may be executed in the form of a hardware structure, a software module, or a hardware structure plus a software module.

[0118] Figure 6 FIG6 is a block diagram of an electronic device 600 for implementing the above method according to an exemplary embodiment. For example, the electronic device 600 may be a mobile phone, a computer, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0119] Reference Figure 6 , the electronic device 600 may include one or more of the following components: a processing component 602 , a memory 604 , a power component 606 , a multimedia component 608 , an audio component 610 , an input / output (I / O) interface 612 , a sensor component 614 , and a communication component 616 .

[0120] The processing component 602 generally controls the overall operation of the electronic device 600, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 602 may include one or more processors 620 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 602 may include one or more modules to facilitate interaction between the processing component 602 and other components. For example, the processing component 602 may include a multimedia module to facilitate interaction between the multimedia component 608 and the processing component 602.

[0121] The memory 604 is configured to store various types of data to support operations on the electronic device 600. Examples of such data include instructions for any application or method operating on the electronic device 600, contact data, phone book data, messages, pictures, videos, etc. The memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0122] The power supply assembly 606 provides power to the various components of the electronic device 600. The power supply assembly 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 600.

[0123] The multimedia component 608 includes a screen that provides an output interface between the electronic device 600 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 608 includes a front camera and / or a rear camera. When the electronic device 600 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0124] The audio component 610 is configured to output and / or input audio signals. For example, the audio component 610 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device 600 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 604 or transmitted via the communication component 616. In some embodiments, the audio component 610 also includes a speaker for outputting audio signals.

[0125] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.

[0126] The sensor assembly 614 includes one or more sensors for providing various aspects of status assessment for the electronic device 600. For example, the sensor assembly 614 can detect the open / closed state of the electronic device 600, the relative positioning of components, such as the display and keypad of the electronic device 600. The sensor assembly 614 can also detect changes in the position of the electronic device 600 or a component of the electronic device 600, the presence or absence of user contact with the electronic device 600, the orientation or acceleration / deceleration of the electronic device 600, and temperature changes of the electronic device 600. The sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 614 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 614 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0127] The communication component 616 is configured to facilitate wired or wireless communication between the electronic device 600 and other devices. The electronic device 600 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, 4G LTE, 5G NR (NewRadio) or a combination thereof. In an exemplary embodiment, the communication component 616 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 616 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0128] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.

[0129] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by the processor 620 of the electronic device 600 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0130] The embodiments of the present disclosure further provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the method described in the above embodiments of the present disclosure.

[0131] Figure 7 FIG. 7 is a schematic diagram showing a structure of a chip 700 for implementing the above method according to an exemplary embodiment. Figure 7 The chip 700 includes a communication interface 701 and at least one processor 702. The communication interface 701 is used to receive signals input into the chip 700 or signals output from the above chip 700. The processor 702 communicates with the communication interface 701 and implements the method described in the above embodiments of the present disclosure through logic circuits or executing code instructions.

[0132] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0133] Throughout this specification, reference to terms such as "one embodiment," "some embodiments," "illustrative embodiments," "example," "specific example," or "some examples" means that a specific feature, structure, material, or characteristic described in conjunction with the embodiment or example is included in at least one embodiment or example of the present invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in at least one embodiment or example.

[0134] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.

[0135] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processing module, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection having at least one wire (control method), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or otherwise processing it in a suitable manner if necessary, and then storing it in a computer memory.

[0136] It should be understood that various parts of the embodiments of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0137] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0138] Furthermore, the functional units in the various embodiments of the present invention may be integrated into a single processing module, each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in either hardware or software functional modules. If the integrated modules are implemented as software functional modules and sold or used as standalone products, they may also be stored in a computer-readable storage medium. The aforementioned storage medium may be a read-only memory, a magnetic disk, or an optical disk, etc.

[0139] Although the embodiments of the present invention have been shown and described above, it will be understood that the above embodiments are exemplary and are not to be construed as limitations on the present invention. A person skilled in the art may change, modify, replace and modify the above embodiments within the scope of the present invention.

Claims

1. A model training method, characterized in that: The method comprises: Obtain the input data of the current training round and the true value data corresponding to the input data; Determine, based on the input data, predicted data corresponding to the input data using an initial model; Determine predicted traceability data corresponding to the true value data, the predicted traceability data including predicted user query data and predicted knowledge point data, the predicted user query data being predicted user query data, the user query data being user question data, the predicted knowledge point data being predicted knowledge point data, and the knowledge point data being data related to the user question data; The initial model is trained using the input data, the true value data, the predicted data, and the predicted traceability data.

2. The method according to claim 1, characterized in that The method further comprises: Obtain user query data, system prompt data, and at least one knowledge point data; Splicing the user query data, the system prompt data, and the at least one knowledge point data to obtain first spliced data; A reordering operation is performed on the first spliced data to obtain the input data.

3. The method according to claim 2, characterized in that The performing a reordering operation on the first spliced data to obtain the input data includes: Determining a random number corresponding to the first spliced data; When the random number belongs to the first range, performing random reordering on the labels of the at least one knowledge point data to obtain the input data; When the random number belongs to the second range, performing random reordering on the position of the at least one knowledge point data in the first spliced data to obtain the input data; In a case where the random number belongs to the third range, random reordering is performed on the label of the at least one knowledge point data and the position of the at least one knowledge point data in the first spliced data to obtain the input data.

4. The method according to claim 2, characterized in that Determining the predicted traceability data corresponding to the true value data includes: Determining, according to the label data in the true value data, first knowledge point data corresponding to the true value data in the at least one knowledge point data; Determine the location of the first knowledge point data as a first location, and determine the location of the user query data as a second location; generating the predicted knowledge point data according to the first position and the input data; The predicted user query data is generated according to the second position and the input data.

5. The method according to claim 4, characterized in that The method further comprises: The first position and the second position are demasked.

6. The method according to claim 4, characterized in that The training of the initial model using the input data, the true value data, the predicted data, and the predicted traceability data includes: Determining a first loss function corresponding to the initial model in a current training round based on a difference between the input data and the predicted traceability data, and a difference between the true data and the predicted data; When the first loss function satisfies a convergence condition, determining the initial model as a target model; If the first loss function does not meet the convergence condition, the initial model is trained in the next round until the loss function of the initial model meets the convergence condition.

7. The method according to claim 4, characterized in that The determining, based on the difference between the input data and the predicted traceability data, and the difference between the true value data and the predicted data, of a first loss function corresponding to the current training round includes: determining a difference between the predicted knowledge point data and the first knowledge point data as a first difference; determining a difference between the predicted user query data and the user query data in the input data as a second difference; determining a difference between the true value data and the predicted data as a third difference; The first loss function is determined according to the first difference, the second difference, and the third difference.

8. A model training device, comprising: A data processing module is used to obtain the input data of the current training round and the true value data corresponding to the input data; A training module, configured to determine, based on the input data and using an initial model, predicted data corresponding to the input data; Determine predicted traceability data corresponding to the true value data, the predicted traceability data including predicted user query data and predicted knowledge point data, the predicted user query data being predicted user query data, the user query data being user question data, the predicted knowledge point data being predicted knowledge point data, and the knowledge point data being data related to the user question data; The initial model is trained using the input data, the true value data, the predicted data, and the predicted traceability data.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.

11. A chip, characterized in that: The method comprises at least one processor and a communication interface; the communication interface is used to receive a signal input to the chip or a signal output from the chip, and the processor communicates with the communication interface and implements the method as described in any one of claims 1 to 7 through a logic circuit or executing code instructions.