Information processing method and electronic equipment
By introducing multi-skill training and SKE-Learn framework into the language model, explicitly extracting and verifying intrinsic knowledge, the problem of data hallucination amplification in the iterative self-learning process of traditional language models is solved, and the reliability and performance of the model are improved.
Patent Information
- Application Number
- CN202510259580.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-24
AI Technical Summary
Traditional language models have the problem of data hallucination amplification during the iterative self-learning process, which leads to reduced reliability of the model and may cause crashes.
Using a large model based on multiple meta skills training, including the first meta skills used to extract internal knowledge, the second meta skills used to generate reasoning, and the third meta skills used to self-evaluate, the SKE-Learn framework is used to explicitly extract, verify and use inferences using intrinsic knowledge.
It improves the performance of the language model, reduces the hallucination amplification problem in the iterative self-learning process, improves data utilization efficiency, and ensures the quality of output results.
Smart Images

Figure CN120196714A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of computers, and more particularly to information processing methods and electronic devices. Background Art
[0002] Natural Language Processing (NLP), as a key branch in the fields of computer science and artificial intelligence, is widely applied in scenarios such as intelligent dialogue, content creation, information extraction, etc., and based on this, effective human-computer interaction can be achieved. Generative large language models are deep learning models capable of generating natural language text. In the process of NLP, improving the performance of large language models is one of the core goals. A powerful model can understand natural language more accurately, thereby better serving users. Summary of the Invention
[0003] According to an exemplary embodiment of the present disclosure, there is provided a method, apparatus, electronic device, computer-readable storage medium, and computer program product for information processing.
[0004] In a first aspect of the present disclosure, there is provided an information processing method, including: obtaining input information; and inputting the input information into a trained large model to obtain output information corresponding to the input information, where the output information includes internal knowledge output and inference output, and the trained large model is trained based on multiple meta-skills, where the meta-skills include a first meta-skill for extracting internal knowledge, a second meta-skill for generating inferences, and a third meta-skill for self-evaluation.
[0005] In a second aspect of the present disclosure, there is provided an electronic device, including: at least one processing unit; at least one memory, at least one memory being coupled to at least one processing unit and storing instructions for execution by at least one processing unit, the instructions when executed by at least one processing unit causing the electronic device to execute the method described in the first aspect of the present disclosure.
[0006] In a third aspect of the present disclosure, there is provided an information processing apparatus, including: an obtaining unit configured to obtain input information; and an output unit configured to input the input information into a trained large model to obtain output information corresponding to the input information, where the output information includes internal knowledge output and inference output, and the trained large model is trained based on multiple meta-skills, where the meta-skills include a first meta-skill for extracting internal knowledge, a second meta-skill for generating inferences, and a third meta-skill for self-evaluation.
[0007] In a fourth aspect of the present disclosure, there is provided a computer-readable storage medium having machine-executable instructions stored thereon, which when executed by a device cause the device to perform the method described in the first aspect of the present disclosure.
[0008] In a fifth aspect of the present disclosure, there is provided a computer program product including computer-executable instructions, wherein the computer-executable instructions, when executed by a processor, implement the method described in the first aspect of the present disclosure.
[0009] In a sixth aspect of the present disclosure, there is provided an electronic device including: a processing circuit configured to perform the method described in the first aspect of the present disclosure.
[0010] The Summary of the Invention is provided to introduce a series of concepts in a simplified form, which will be further described in the Detailed Description below. The Summary of the Invention is not intended to identify the key features or essential features of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In conjunction with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more apparent. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0012] Figure 1 shows a schematic structural diagram of an example environment to which some embodiments of the present disclosure can be applied;
[0013] Figure 2 shows a schematic flow chart of a model training process according to some embodiments of the present disclosure;
[0014] Figure 3 shows a schematic diagram of model training according to some embodiments of the present disclosure;
[0015] Figure 4 shows a schematic diagram of an example usage process according to some embodiments of the present disclosure;
[0016] Figure 5 shows a schematic flow chart of a process of model output according to some embodiments of the present disclosure;
[0017] Figure 6 shows a block diagram of an example device according to some embodiments of the present disclosure; and
[0018] Figure 7 shows a block diagram of an example device that can be used to implement the embodiments of the present disclosure. DETAILED DESCRIPTION
[0019] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0020] In the field of NLP, language models are capable of deeply understanding and processing natural language. For example, according to the user's needs, a language model can accurately extract corresponding information from numerous materials. The information may include, but is not limited to, daily knowledge, travel guides, inspiration and materials for creation, etc. It can be understood that, as an assistant, the language model can help users save time and effort and improve work and learning efficiency.
[0021] However, in the iterative self-learning process of traditional language models, there is also a problem of data hallucination amplification, resulting in a decrease in the reliability of the model and possibly causing crashes.
[0022] As the number of iterations increases, the output of the language model gradually deviates from the true knowledge, reducing the reliability of the language model.
[0023] To solve the above problems and potential other problems, embodiments of the present disclosure provide a method for information processing. In the embodiments of the present disclosure, a trained large model can be used to obtain output information corresponding to the input information, where the output information includes internal knowledge output and inference output, and where the trained large model is trained based on multiple meta-skills, and the meta-skills include a first meta-skill for extracting internal knowledge, a second meta-skill for generating inferences, and a third meta-skill for self-evaluation.
[0024] In this way, a self-knowledge explicitation learning (SKE-Learn) framework is proposed, enabling the large model to explicitly extract, verify, and utilize internal knowledge for reasoning. Knowledge explanations are provided for the inferences of the model, thereby ensuring the quality of the output results. In addition, the large model can also automatically select reliable data, alleviating the problem of hallucination amplification in the iterative self-learning process, thereby improving the performance of the language model and showing higher data utilization efficiency.
[0025] The large model in the embodiments of the present disclosure is also referred to as a large language model, a large model based on self-knowledge explicitation, a large model based on internal knowledge explicitation, or other names, and the present disclosure is not limited thereto.
[0026] Figure 1 FIG. 1 shows a schematic structural diagram of an exemplary environment to which some embodiments of the present disclosure can be applied. The exemplary environment 100 includes input information 102, a trained large model 104, and output information 106.
[0027] As Figure 1 shown, the input information 102 is input into the trained large model 104, and the large model 104 analyzes and processes the input information 102 to determine the output information 106 corresponding to the input information 102. The output information 106 includes two parts: internal knowledge output and inference output. The input information 102 may refer to the user's requirements and instructions, including but not limited to text types (question consultation, instruction description, etc.), voice types, and image types, etc. For example, in terms of text-based written consultation, a user may input "What is the order of the eight major planets in the solar system?" expecting the large model to give accurate popular science content. Among them, the internal knowledge may refer to the knowledge that the large model 104 autonomously extracts from a large-scale unsupervised knowledge corpus or its own knowledge system and is associated with the input information 102. In the embodiments of the present disclosure, this part of the internal knowledge does not require manual annotation and is completely automatically extracted by the large model.
[0028] In some embodiments of the present disclosure, the trained large model 104 is trained based on multiple meta-skills. The multiple meta-skills include: a first meta-skill for extracting internal knowledge, a second meta-skill for generating inferences, and a third meta-skill for self-evaluation.
[0029] Exemplarily, the first meta-skill can extract relevant explicit internal knowledge (denoted as k) based on a question (denoted as q), expressed as Equation (1) below:
[0030] k = M(q, p extract ) (1)
[0031] Exemplarily, the second meta-skill can generate an inference (denoted as r) based on the extracted knowledge, expressed as Equation (2) below:
[0032] r = M(q, k, p reason ) (2)
[0033] Exemplarily, the third meta-skill can self-evaluate the quality of the knowledge extraction and inference processes, expressed as Equations (3)-(4) below:
[0034]
[0035] In Equations (1)-(4), M represents the large model, is the reference knowledge corresponding to the question q, s k and s rThey are the score values of the self-assessment of knowledge and reasoning, p extract , p reason , and are the relevant prompt information of meta-skills.
[0036] The meta-skill for extracting internal knowledge helps the large model 104 to autonomously identify and extract knowledge associated with a given problem from a large-scale unsupervised knowledge corpus (such as a knowledge collection accessible via the Internet) or its own learned knowledge system, without relying on manual annotation of a large amount of knowledge data. With the help of this skill, the large model 104 can deeply understand the semantics and context of the problem, and accurately locate and extract the knowledge information helpful for solving the problem from the vast knowledge reserve. It can be understood that with the relevant knowledge extracted as the reasoning basis of the output information, the large model 104 can perform effective reasoning and analysis. The large model 104 combines the input information 102 and the existing knowledge to obtain a more reliable and reasonable reasoning output, rather than simply matching and retrieving knowledge.
[0037] In some embodiments of the present disclosure, the meta-skill for generating reasoning means that the large model 104 can perform in-depth logical deduction and correlation analysis based on the extracted relevant knowledge. When facing complex problems, the large model 104 can integrate, process and apply knowledge, and perform effective reasoning and judgment in combination with the context and conditions of the specific problem. On the one hand, it improves the quality and depth of the answer, and the reasoning results given have rigorous logic and rich connotations; on the other hand, it enhances the model's ability to handle complex tasks. Whether it is a scientific research problem or a decision-making dilemma in real life, it can provide more valuable reasoning results and bring more efficient and accurate services to users. For example, in the medical field, when a doctor inputs information such as a patient's symptoms (such as cough, fever, fatigue), medical history (having had pneumonia) and examination results (blood routine shows an increase in white blood cells) into the large model 104, the large model 104 can combine knowledge about the characteristics and pathogenesis of various diseases in the medical knowledge base to infer the possible diseases the patient may have, such as it may be bacterial pneumonia, and further infer a suitable treatment plan, such as which antibiotic to recommend, the dosage and the course of treatment, etc., providing strong reference for the doctor's diagnosis and treatment. It can be understood that the relevant knowledge and reasoning results generated by the large model 104 according to the input information 102 can be selected by the user to view.
[0038] In some embodiments of the present disclosure, the meta - skills for self - assessment refer to the ability of the large - model 104 to score the quality of the internal knowledge it extracts and the reasoning results it generates, and to measure the accuracy, relevance degree of the internal knowledge, as well as the logic and rationality of the reasoning results according to preset criteria and rules. Further, in terms of knowledge extraction, when the model obtains relevant knowledge from a large amount of information, the self - assessment mechanism can determine whether this knowledge is accurate and comprehensive. The self - assessment mechanism is the key to assisting the model in stable iterative learning. Specifically, the parts with high scoring values of the self - assessment mechanism can be used for learning, while the parts with low scoring values can be deleted, thereby improving the learning efficiency of the model.
[0039] In previous learning processes, the large - model may generate some content that seems reasonable but actually deviates from the real world, that is, the so - called "hallucinations". The three meta - skills of the large - model 104 in the embodiments of the present disclosure can ensure the quality of the dataset during model training, thereby alleviating the "hallucination" phenomenon of the model. In addition, it can also reduce the dependence on irrelevant information, improve the processing ability of key information, enable the model to run more stably and efficiently in various complex tasks, and provide better services for users.
[0040] Figure 2 FIG. shows a schematic flow chart of a model training process according to some embodiments of the present disclosure. At block 202, based on an initial model and a knowledge base, a first model is generated, where the initial model does not have the self - assessment ability and the first model has the self - assessment ability. At block 204, based on the first model and the knowledge base, a trained large - model is generated through iterative training.
[0041] In some embodiments of the present disclosure, the initial model can be trained on meta - skills to obtain the first model, and then iterative training is performed based on the first model to finally obtain the trained large - model (such as Figure 1 104 in). Exemplarily, the process of training the initial model to obtain the first model can be referred to as the first training stage, and the iterative training process of obtaining the trained large - model from the first model can be referred to as the second training stage. Optionally, the iterative process in the second training stage is a process of self - learning and self - evolution of the large - model. And the training set of the training process can be obtained based on the knowledge base. Optionally, the initial model can be represented as M 0 , and the first model is represented as (or M1). In the second training stage, the large - model generated after the i - th iteration is represented as M i+1 (i = 1, 2,...), and the trained large - model is represented as M n . Among them, M 0Without the meta-skill of self-evaluation, after the first training stage, the first model has the first meta-skill of extracting internal knowledge, the second meta-skill of generating inferences, and the third meta-skill of self-evaluation, so that it can further achieve self-iterative optimization in the second training stage. It can be understood that with respect to the third meta-skill of self-evaluation, the initial model does not have this third meta-skill, while the first model after the first training stage has this third meta-skill.
[0042] In the embodiments of the present disclosure, the knowledge base contains reference knowledge in many fields, including but not limited to medicine, industry, economy, history, and human culture, etc. Each field contains multiple branches. Taking a knowledge field about tourism as an example, the knowledge base may have reference knowledge such as introductions of various tourist attractions, local cuisine, and transportation information. In the embodiments of the present disclosure, by training the model, the model masters three meta-skills of extracting internal knowledge, reasoning, and self-evaluation. By using these meta-skills, the large model can establish an iterative self-learning method to ensure that the model can reliably select self-synthesized data each time, alleviate the hallucination problem, and improve the model performance.
[0043] In some embodiments, the training process of the first training stage at block 202 may include: obtaining a knowledge base, where the knowledge base includes multiple reference knowledge; constructing a training data set for the first training stage based on the knowledge base, where the training data set for the first training stage includes a first data set corresponding to the first meta-skill, a second data set corresponding to the second meta-skill, and a third data set corresponding to the third meta-skill; and obtaining the first model based on the initial model (M 0 ) and the training data set of the first training stage
[0044] It can be understood that M 0 has the ability to understand and analyze the content of the knowledge base. In the first training stage, for the reference knowledge (or knowledge fragments) in the knowledge base, the initial model can reversely generate question samples corresponding to the knowledge fragment, then extract internal knowledge samples involved in the question itself based on the question sample, and then generate corresponding inference samples based on the extracted internal knowledge samples.
[0045] M 0 The internal knowledge extracted by M will have differences in terms of relevance and accuracy with the question, and the inference process and the obtained inference results based on these knowledge will also have different quality levels in terms of logic and rationality. For example, when asked "Why do plants release carbon dioxide at night?", M 0The extracted internal knowledge may contain some irrelevant plant growth cycle information, and the inference result may be that "plants release carbon dioxide because they carry out photosynthesis at night", which is obviously contrary to scientific facts, lacking both relevance and reasonable inference. To improve the performance of the model and enable M 0 to learn the meta-skill of self-assessment, hint information can be introduced during the construction of the training dataset to score knowledge samples and inference samples. In some embodiments, the hint information may include, but is not limited to, accuracy hints, relevance hints, logical hints, and integrity hints, etc. In some embodiments, a scoring model (such as an external supervision model) can be used to score samples in combination with the hint information.
[0046] Specifically, for the reference knowledge (denoted as c i ) in the knowledge base, assuming the question sample is denoted as q i , the relevant knowledge sample extracted by the initial model is denoted as k i , and the inference sample generated based on the relevant knowledge is denoted as r i . Then, further based on the hint information, a scoring model can be used to determine the first scoring sample of the knowledge sample (such as the score value of the knowledge sample, denoted as ) and the second scoring sample of the inference sample (such as the score value of the inference sample, denoted as ). Subsequently, a training dataset for the first training stage can be constructed based on the reference knowledge, question sample, knowledge sample, inference sample, first scoring sample, and second scoring sample. Exemplarily, the data item in the first dataset corresponding to the first meta-skill can be denoted as (q i , k i ), the data item in the second dataset corresponding to the second meta-skill can be denoted as (q i , k i , r i ), and the data item in the third dataset corresponding to the third meta-skill can be denoted as
[0047] In some examples, a first model can be trained based on the first dataset corresponding to the first meta-skill, the second dataset corresponding to the second meta-skill, and the third dataset corresponding to the third meta-skill.
[0048] In some examples, a first subset can also be selected from the first dataset, a second subset from the second dataset, and a third subset from the third dataset. And a first model can be obtained through training based on the first subset, the second subset, and the third subset. Optionally, this process can also be referred to as a dataset filtering process. By filtering the training dataset in the first training stage, training can be performed based on a part (such as the high-quality part) rather than all data items in the training dataset of the first training stage.
[0049] Specifically, to filter the training dataset, a first threshold for knowledge sample scoring and a second threshold for inference sample scoring can be preset in advance. Then, based on the predefined thresholds, filtering of the training dataset can be achieved, and a first subset (e.g., denoted as ) and a second subset (e.g., denoted as ) can be selected from the first dataset and the second dataset. As an example, the score value range can be from 0 to 10, and the first and second thresholds can be 8.
[0050] For example, for a certain data item if is greater than (or greater than or equal to) the first threshold, and is greater than (or greater than or equal to) the second threshold, then the first subset includes the data item (q1, k1), and the second subset includes the data item (q1, k1, r1). On the contrary, if is less than or equal to (or less than) the first threshold, and is less than or equal to (or less than) the second threshold, then the first subset does not include the data item (q1, k1), and the second subset does not include the data item (q1, k1, r1). In this way, knowledge samples k i with scores higher than the preset threshold and inference samples r i can be selected respectively to construct a training dataset for training to extract internal knowledge and a training dataset for training inference
[0051] Optionally, a third subset can be determined from the third dataset through uniform sampling. Further, by uniformly sampling the scores s k and s r , a training dataset corresponding to the third meta-skill can be determined
[0052] In some embodiments, in the first training stage, a first model can be obtained based on a first loss function. The first loss function in the first training stage can be expressed as Equation (5) below
[0053]
[0054] In Equation (5), the first loss function includes the training objectives for three meta-skills.
[0055] The first model M generated after the first training stage 1 has three meta-skills: intrinsic knowledge extraction, reasoning, and self-assessment. In the subsequent second training stage, it can iteratively self-train using self-synthesized data and select high-quality data through its self-assessment ability to improve the performance of the model.
[0056] In some embodiments, the training process in the second training stage at block 204 may include: in the first iteration of the iterative training, using the first model as the base model to obtain a result model based on the knowledge base; and in any iteration after the first iteration of the iterative training, using the result model obtained in the previous iteration as the base model to obtain the result model for the current iteration based on the knowledge base.
[0057] Exemplarily, during the second training stage, through the 1st iteration, the first model is updated to model M 2 ; through the 2nd iteration, model M 2 is updated to M 3 ;... through the nth iteration, model M n is updated to M n+1 .
[0058] Taking the jth iteration as an example, model M j is trained through the jth iteration to generate model M j+1 . Specifically, a training data set for the jth iteration process can be constructed based on the knowledge base, where the training data set for the jth iteration process includes a first training data set corresponding to the first meta-skill and a second training data set corresponding to the second meta-skill.
[0059] Specifically, for the reference knowledge (e.g., c j ) in the knowledge base, model M j can be used to generate a question sample, a knowledge sample, a reasoning sample, a first scoring sample for the knowledge sample, and a second scoring sample for the reasoning sample corresponding to the reference knowledge. It can be understood that in this process, model M j deeply mines and synthesizes the information in the knowledge base based on its own knowledge reserve and processing ability. Model M j can generate a question sample q j , a knowledge sample k j , and a reasoning sample r j corresponding to the reference knowledge (e.g., c j ) according to the knowledge and reasoning patterns learned previously. Model M jAfter generating a new dataset, the self - evaluation meta - skill is utilized to score the generated knowledge sample k j and the inference sample r j For example, a first scoring sample can be determined for the knowledge sample and a second scoring sample can be determined for the inference sample It can be understood that in the second training stage, the large - model M j has the third meta - skill of self - evaluation, so it can autonomously score based on the prompt information. It can be seen that no other scoring models are needed in the second training stage.
[0060] Furthermore, if exceeds the first threshold and exceeds the second threshold, it indicates that the quality of this sample is high, with high accuracy and reliability. Based on these problem samples, knowledge samples, and inference samples that meet the threshold requirements, a training dataset for the j - th iteration process is constructed. Among them, the data items in the first training dataset corresponding to the first meta - skill can be represented as (q j , k j ), which can be used to train the model's intrinsic knowledge extraction ability; the data items in the second training dataset corresponding to the second meta - skill can be represented as (q j , k j , r j ), which can be used to strengthen the model's inference ability.
[0061] Based on the basic model M j and the constructed training dataset for the j - th iteration process, training operations are carried out. During the training process, the model continuously learns the knowledge and inference patterns contained in the dataset, adjusts its own parameters and structure to optimize the intrinsic knowledge extraction and inference abilities. After this round of training, the basic model M j evolves into the result model M j+1 .
[0062] The second loss function in any iteration process of the second training stage is shown in the following formula (6):
[0063]
[0064] In formula (6), the second loss function includes the cross - entropy based on the first meta - skill and the second meta - skill.
[0065] It can be understood that the model M j is trained and generated by M j-1 in the (j - 1) - th iteration process. The training dataset for the (j - 1) - th iteration process is synthesized by M j-1 . M j-1Knowledge is verified through its self - assessment ability and reference knowledge from the knowledge base. This verification mitigates hallucinated data in the model's self - learning process by establishing a more reliable data selection process. Additionally, since all data comes from the model itself, the iterative training process can ensure reliability and automation. In each iterative training, the model can gradually refine its knowledge boundaries and meta - skills from the experience of the model generated in the previous iterative training, thus achieving iterative improvement in performance.
[0066] Figure 3 FIG. 4 shows a schematic structural diagram of an example environment 300 for model training according to some embodiments of the present disclosure. The example environment 300 includes a first training stage 302 and a second training stage 304.
[0067] As Figure 3 shown, in the first training stage 302, the initial model M 0 can determine a problem sample q i based on the knowledge base C, and further extract the associated knowledge k i . According to the knowledge k i , reason about the problem q i to generate r i . To enable the model to learn three meta - skills: intrinsic knowledge extraction, reasoning, and self - assessment, a meta - skill training dataset needs to be constructed. As shown above, the model M 0 determines problem - knowledge - reasoning samples from the knowledge base and can score the knowledge samples and reasoning samples based on the hint information using equations (3) and (4). Optionally, a scoring model can be used, in combination with the hint information, to score the knowledge samples and reasoning samples. Optionally, the scoring model can be a higher - performance external supervision model. Optionally, the hint information can indicate the way of scoring, such as the score range, etc. Specifically, the scoring model can evaluate the accuracy, integrity, and relevance of the knowledge samples. For example, in mathematical knowledge, if the formula given in the knowledge sample does not match the standard formula in the reference knowledge, it will be judged as inaccurate. The scoring model can also evaluate the logic and rationality of the reasoning samples. According to the above evaluation results, corresponding scores are assigned to the knowledge samples and reasoning samples. For example, a percentage - based scoring system can be adopted. Knowledge samples with high accuracy, good integrity, and strong relevance can get higher scores; reasoning samples with strict logic and high rationality will also receive high scores. Through quantitative scoring, the model can more intuitively compare the quality of different samples, providing a basis for subsequent training and optimization.
[0068] It can be understood that, in order to further improve the performance of the model and enable the model to master the three meta - skills, high - quality k i sets and r iSets, respectively construct training data sets and where is used for training of internal knowledge extraction, is used for training of reasoning. For the meta-skill of self-assessment, the scores s k and s r are uniformly sampled to construct the training data set Based on the model is trained. To further increase the diversity of questions, a small number of existing questions can also be added in this process to construct the knowledge and reasoning processes in a similar way.
[0069] Combined with the above formula (5), by training the model M 0 the generated M 1 has three meta-skills of internal knowledge extraction, reasoning and self-assessment, can realize self-iterative training, gradually optimize the performance of the model, and at the same time alleviate the problem of data hallucination amplification in the self-learning process.
[0070] In the second training stage 304, the model M j that already has the self-assessment ability does not need to rely on an external scoring model to score the knowledge samples and reasoning samples. Therefore, M j can select the sample data by itself, and the high-quality sample data constitutes the training data set and Train M j to generate M j+1 .
[0071] Exemplarily, each iterative training will generate a model with better performance. Referring to Figure 3 in the second training stage 304, taking the j-th iterative training as an example, taking the model M j generated in the (j - 1)-th iterative process as the base model, based on the database, M j generates the model M j+1 through the current iterative training. Among them, the training data set in the j-th iterative training process is synthesized by M j based on the knowledge base.
[0072] Specifically, the model M j can determine the problem sample q j based on the knowledge base. As described above, M j has three meta-skills. Among them, the internal knowledge extraction meta-skill can extract the knowledge sample k j associated with the problem sample q j . Further, according to the reasoning meta-skill, the model M j is based on the problem sample q jand the knowledge sample k j Infer the result r j Utilize the model M j 's self-evaluation ability to provide j the scores of k j and r respectively denoted as Select the k j and r j with higher scores to construct a new set of training data and for training M j to generate a more mature model M j+1 The training objective of the iterative training process is as shown in the above formula (6).
[0073] Figure 4 FIG. shows a schematic diagram of a process 400 of model output according to some embodiments of the present disclosure. In Figure 4 a question 402, internal knowledge 404, an inference result 406, and a trained large model 408 are shown.
[0074] As Figure 4 shown, the entire information processing flow demonstrates the efficient and accurate processing ability of the trained large model 408 when facing user questions. The user inputs the question 402 into the trained large model 408. Here, the question 402 is a specific query that the user expects to get an answer to, which can come from various fields and cover various topics. For example, assume the question 402 is "Given x = 1, what does x << 3 mean in Python 3", and the user, with a question about Python programming syntax, inputs it into the model to seek an answer.
[0075] After receiving the question 402, the large model 408 extracts relevant internal knowledge 404 according to the question. The internal knowledge 404 is explicit knowledge that the model learns and stores from a large knowledge base during the training process. These knowledges are like the "wisdom treasure house" of the model, providing a solid foundation for solving problems. In the above Python question, the internal knowledge 404 may include the definitions and operation rules of bitwise operators in Python. It can be understood that the explicitization of internal knowledge is of great significance. Through this process, users can view the relevant internal knowledge, increasing the transparency of the model processing process and allowing users to clearly know what the model is based on to answer questions.
[0076] After that, the model 408 performs reasoning based on the extracted internal knowledge 404 to obtain the reasoning result 406. The reasoning process is a process in which the model uses logical thinking to combine and analyze the internal knowledge with the problem. For the question "Given x = 1, what does x << 3 mean in Python 3", the model will reason based on the operation rule of the bitwise operator "<<" in the internal knowledge (the left shift operator shifts the binary bits of a number to the left by the specified number of bits, and each left shift is equivalent to multiplying by 2). Since the binary representation of x = 1 is 0001, and after shifting 3 bits to the left it becomes 01000, and the corresponding decimal number is 8, so the reasoning result 406 is "x << 3 in Python 3 means shifting the binary bits of x to the left by 3 bits, and the result is 8".
[0077] Different from the information processing of traditional large models, Figure 4 the shown reasoning process is verified by clear and viewable knowledge. Traditional large models may give results when processing problems, but users are not clear about their basis and reasoning process, which has a certain "black box" nature. While the trained large model 408 can cite clear internal knowledge as evidence, which greatly reduces the "hallucination" phenomenon. By clearly citing internal knowledge, the reliability and accuracy of the model output are improved, thereby enhancing the overall quality of the generated output and providing more credible and valuable answers for users.
[0078] Figure 5 FIG. shows a schematic flow chart of a process 500 for information processing according to some embodiments of the present disclosure. At block 502, input information is obtained. At block 504, the input information is input into the trained large model to obtain output information corresponding to the input information, and the output information includes an internal knowledge output and a reasoning output, where the trained large model is trained based on multiple meta-skills, and the meta-skills include a first meta-skill for extracting internal knowledge, a second meta-skill for generating reasoning, and a third meta-skill for self-evaluation.
[0079] In some embodiments, the input information obtained at block 502 may come from a user input. For example, the input information may be a specific question raised by the user. For example, in the medical field, the user may input "What are the early symptoms of diabetes"; it may also be a text content that needs to be analyzed and interpreted, such as a news report, a scientific research literature, etc.
[0080] At block 504, after the input information is input into the trained large model, the meta-skills possessed by the large model will analyze and process the input information. Among them, the first meta-skill is used to extract internal knowledge, and quickly screen out the knowledge content related to the input information from the model's huge knowledge base according to the input information. For example, when the input is "How many major planets are there in the solar system?", the model will accurately extract the internal knowledge "There are eight major planets in the solar system, namely Mercury, Venus, Earth, Mars, Jupiter, Saturn, Uranus, and Neptune" and output it based on the first meta-skill. The second meta-skill is used to generate reasoning. Based on the extracted internal knowledge, the model will use logical analysis and reasoning abilities to deeply deduce the input information. If the input information is "The temperature has dropped suddenly recently. What impact may it have on crops?", after the model extracts relevant meteorological and agricultural knowledge, it will use the second meta-skill to reason out inferences such as "The sudden drop in temperature may cause crops to suffer from cold damage, affect their growth and development, reduce yields, and may also cause some cold-intolerant crops to die".
[0081] In some examples, the trained large model can be applied to multiple fields. For example, in the intelligent customer service scenario, the large model can accurately understand the user's consultation information and output professional knowledge and reasonable solutions; in the academic research field, it can help researchers quickly obtain relevant knowledge and conduct reasoning and analysis. Compared with traditional models, the trained large model has significant advantages in the accuracy, comprehensiveness, and intelligence of information processing, and can better meet the information processing needs in different scenarios.
[0082] In some embodiments of the present disclosure, in order to verify the SKE-Learn model optimization scheme trained based on meta-skills, experiments were carried out for multiple natural language processing benchmark tasks and compared with existing traditional large models. The experimental results are shown in Table 1, where "SKE-Learn" represents the scheme of the present disclosure. In Table 1, Massive Multitask Language Understanding (MMLU), BIG Bench Hard (BBH), AI2 Reasoning Challenge-Easy (ARC-E), AI2 Reasoning Challenge-Challenge (ARC-C), and Natural Questions (NQ) are the benchmark test sets for verification. And the values in Table 1 represent the prediction accuracy (%) of the benchmark test sets.
[0083] Table 1
[0084]
[0085] In Table 1, Zero-shot, Chain-of-Thought, Plan-and-Solve, and Re-reading all belong to prompt-based methods and all implicitly utilize the internal knowledge of large language models. From the data in the table, it can be seen that SKE-Learn(M 0 ) performs relatively prominently in terms of the average score and multiple specific tasks, higher than most comparison methods. This indicates that the initial SKE-Learn(M 0 ) model already has certain advantages.
[0086] Compared with the traditional self-learning baseline , the meta-skill training model has an average performance improvement of 3.83%, while the training data used is reduced by 90%. This means that large models with meta-skills can achieve better results without relying on a large amount of data, greatly saving resources. Further, the untrained model M 0 has poor performance, but after iterative training, the performance of the model continues to improve, achieving performance enhancement in all benchmark tests and training rounds. The average score in the meta-skill training stage M 1 increases by 1.85%. Compared with M 4 , the average performance in the fourth iterative training (M 0 ) is improved by 4.41%. After completing the entire self-learning process, the results of certain tasks (such as BBH) are significantly improved by 7.69%. Therefore, the trained large model has a significant improvement effect, can achieve a substantial performance improvement with more efficient data utilization, and has obvious advantages in the field of self-learning.
[0087] Moreover, the performance improvement of the model in the embodiments of the present disclosure does not stem from knowledge injection, but from the meta-skills of gradually strengthening the extraction, verification, and utilization of explicit internal knowledge in the high-quality self-learning process.
[0088] It should be understood that in the embodiments of the present disclosure, "first", "second", "third", etc. are only used to indicate that multiple objects may be different, but at the same time do not exclude the possibility that two objects are the same, and should not be construed as any limitation to the embodiments of the present disclosure.
[0089] It should also be understood that the classification of the methods, situations, categories, and embodiments in the embodiments of the present disclosure is only for the convenience of description and should not constitute a special limitation. The features in various methods, categories, situations, and embodiments can be combined with each other under logical circumstances.
[0090] It should also be understood that the above content is only to help those skilled in the art better understand the embodiments of the present disclosure, rather than to limit the scope of the embodiments of the present disclosure. Those skilled in the art can make various modifications, changes, combinations, etc. according to the above content. The solutions after such modifications, changes, or combinations are also within the scope of the embodiments of the present disclosure.
[0091] It should also be understood that the description of the above content focuses on emphasizing the differences between the various embodiments. The same or similar parts can be referred to or borrowed from each other. For the sake of brevity, they will not be elaborated here.
[0092] Figure 6 Fig. shows a schematic block diagram of an exemplary device 600 according to some embodiments of the present disclosure. The device 600 can be implemented by software, hardware, or a combination of both. As Figure 6 shown, the device 600 includes an acquisition unit 602 and an output unit 604.
[0093] The acquisition unit 602 is configured to acquire input information. The output unit 604 is configured to input the input information into a trained large model to obtain output information corresponding to the input information. The output information includes internal knowledge output and inference output, where the trained large model is trained based on multiple meta-skills, and the meta-skills include a first meta-skill for extracting internal knowledge, a second meta-skill for generating inferences, and a third meta-skill for self-evaluation.
[0094] Optionally or additionally, the device 600 may further include a training unit, configured to: in a first training stage, generate a first model based on an initial model and a knowledge base, where the initial model does not have the self-evaluation ability and the first model has the self-evaluation ability; and in a second training stage, generate a trained large model through iterative training based on the first model and the knowledge base.
[0095] In some embodiments, the training unit may be configured to: acquire a knowledge base, where the knowledge base includes multiple reference knowledges; construct a training data set for the first training stage based on the knowledge base, where the training data set for the first training stage includes a first data set corresponding to the first meta-skill, a second data set corresponding to the second meta-skill, and a third data set corresponding to the third meta-skill; and obtain a first model based on the initial model and the training data set for the first training stage.
[0096] Exemplarily, the training unit may be configured to: for the reference knowledge in the knowledge base, use the initial model to generate question samples, knowledge samples, and inference samples corresponding to the reference knowledge; based on the prompt information, use the reference knowledge as a reliable reference to determine the first scoring sample of the knowledge sample and the second scoring sample of the inference sample; and generate the first sample data in the first dataset, the second sample data in the second dataset, and the third sample data in the third dataset, where the first sample data includes question samples and knowledge samples, the second sample data includes question samples, knowledge samples, and inference samples, and the third sample data includes reference knowledge, question samples, knowledge samples, inference samples, the first scoring sample of the knowledge sample, and the second scoring sample of the inference sample.
[0097] Exemplarily, the training unit may be configured to determine a first subset of the first dataset and a second subset of the second dataset based on predefined first and second thresholds; determine a third subset from the third dataset by uniform sampling; and obtain a first model through training based on the first subset, the second subset, and the third subset.
[0098] Exemplarily, the training unit may be configured to: in the first iteration of the iterative process, use the first model as the base model and obtain a result model based on the knowledge base; and in any iteration of the iterative training, use the result model obtained in the previous iteration as the base model and obtain the result model of the current iteration based on the knowledge base.
[0099] Exemplarily, the training unit may be configured to: construct a training dataset for the current iteration process based on the knowledge base, where the training dataset for the current iteration process includes a first training dataset corresponding to the first meta-skill and a second training dataset corresponding to the second meta-skill; and obtain the result model of the current iteration based on the base model and the training dataset for the current iteration process.
[0100] Exemplarily, the training unit may be configured to: for the reference knowledge in the knowledge base, use the base model to generate question samples, knowledge samples, inference samples, the first scoring sample of the knowledge sample, and the second scoring sample of the inference sample; and if the first scoring sample exceeds the first threshold and the second scoring sample exceeds the second threshold, construct the first sample data in the first training dataset and the second sample data in the second training dataset based on the question samples, knowledge samples, and inference samples.
[0101] Figure 6 The device 600 can be used to implement the above-mentioned process in combination with Figures 1 to 5 the process described above. For the sake of brevity, it will not be elaborated here.
[0102] In the embodiments of the present disclosure, the division of modules or units is illustrative, merely a logical function division. In actual implementation, there may be other division methods. Additionally, in the embodiments of the present disclosure, each functional unit may be integrated into one unit, may exist independently physically, or two or more units may be integrated into one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0103] Figure 7 A block diagram of an example device 700 that may be used to implement the embodiments of the present disclosure is shown. It should be understood that Figure 7 the device 700 shown is merely exemplary and should not constitute any limitation on the functions and scope of the implementations described herein. For example, the device 700 may be used to perform the Figures 1 to 5 processes described above.
[0104] As Figure 7 shown, the device 700 is in the form of a general-purpose computing device. The components of the computing device 700 may include, but are not limited to, one or more processors or processing units 710, a memory 720, a storage device 730, one or more communication units 740, one or more input devices 750, and one or more output devices 760. The processing unit 710 may be an actual or virtual processor and is capable of performing various processes according to the programs stored in the memory 720. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the computing device 700.
[0105] The computing device 700 generally includes multiple computer storage media. Such media may be any accessible media available to the computing device 700, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 720 may be volatile memory (such as registers, caches, random access memory (RAM)), non-volatile memory (such as read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 730 may be removable or non-removable media and may include machine-readable media, such as a flash drive, a magnetic disk, or any other media that can be used to store information and / or data (such as training data for training) and can be accessed within the computing device 700.
[0106] The computing device 700 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 7 , a disk drive for reading from or writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading from or writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to the bus (not shown) by one or more data media interfaces. The memory 720 may include a computer program product 725 having one or more program modules configured to execute the various methods or actions of the various implementations of the present disclosure.
[0107] The communication unit 740 enables communication with other computing devices via a communication medium. Additionally, the functions of the components of the computing device 700 may be implemented in a single computing cluster or multiple computer machines capable of communicating via a communication connection. Thus, the computing device 700 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0108] The input device 750 may be one or more input devices, such as a mouse, keyboard, trackball, etc. The output device 760 may be one or more output devices, such as a display, speaker, printer, etc. The computing device 700 may also communicate with one or more external devices (not shown) as needed via the communication unit 740, such as storage devices, display devices, etc., communicate with one or more devices that enable a user to interact with the computing device 700, or communicate with any device that enables the computing device 700 to communicate with one or more other computing devices (e.g., a network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0109] According to an exemplary implementation of the present disclosure, a computer-readable storage medium is provided, on which computer-executable instructions are stored, where the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is also provided, the computer program product being tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions being executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, a computer program product is provided, on which a computer program is stored, and when the program is executed by a processor, the method described above is implemented.
[0110] Aspects of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0111] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, create a means for implementing the functions / acts specified in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable medium storing the instructions comprises a manufacture including instructions for implementing various aspects of the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0112] The computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more blocks of the flowchart and / or block diagram.
[0113] The flowcharts and block diagrams in the figures illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the figures. For example, two consecutive blocks may in fact be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0114] The various implementations of the present disclosure have been described above. The above description is exemplary and not exhaustive, and is also not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, practical applications, or improvements to technologies in the market, or to enable other ordinary skilled artisans in the art to understand the various implementation manners disclosed herein.
Claims
1. An information processing method, comprising: Get input information; as well as Input the input information into the trained large model to obtain output information corresponding to the input information, wherein the output information includes internal knowledge output and reasoning output, The trained large model is obtained based on multiple meta-skills, wherein the meta-skills include a first meta-skill for extracting internal knowledge, a second meta-skill for generating reasoning, and a third meta-skill for self-assessment.
2. The method according to claim 1, further comprising: In a first training phase, a first model is generated based on an initial model and a knowledge base, wherein the initial model does not have a third meta-skill of self-assessment, and the first model has the first meta-skill, the second meta-skill, and the third meta-skill; as well as In the second training stage, the trained large model is generated through iterative training based on the first model and the knowledge base.
3. The method according to claim 2, wherein generating the first model based on the initial model and the knowledge base comprises: Acquire the knowledge base, wherein the knowledge base includes a plurality of reference knowledge; constructing a training data set for the first training phase based on the knowledge base, wherein the training data set for the first training phase includes a first data set corresponding to the first meta-skill, a second data set corresponding to the second meta-skill, and a third data set corresponding to the third meta-skill; as well as The first model is obtained based on the initial model and the training data set of the first training stage.
4. The method according to claim 3, wherein constructing the training data set of the first training stage based on the knowledge base comprises: For reference knowledge in the knowledge base, using the initial model to generate question samples, knowledge samples, and reasoning samples corresponding to the reference knowledge; Based on the prompt information, the reference knowledge is used as a reliable reference to determine a first scoring sample of the knowledge sample and a second scoring sample of the reasoning sample; as well as Generate first sample data in the first data set, second sample data in the second data set, and third sample data in the third data set, wherein the first sample data includes the problem sample and the knowledge sample, the second sample data includes the problem sample, the knowledge sample and the reasoning sample, and the third sample data includes the reference knowledge, the problem sample, the knowledge sample, the reasoning sample, the first scoring sample of the knowledge sample and the second scoring sample of the reasoning sample.
5. The method according to claim 4, wherein obtaining the first model comprises: Determining a first subset of the first data set and a second subset of the second data set based on a predefined first threshold and a second threshold; Determine a third subset from the third data set by uniform sampling; as well as The first model is obtained through training based on the first subset, the second subset, and the third subset.
6. The method according to claim 2, wherein generating the trained large model through iterative training based on the first model and the knowledge base comprises: In a first iteration of the iterative training, the first model is used as a basic model, and a result model is obtained based on the knowledge base; as well as In any iteration process after the first iteration process of the iterative training, the result model obtained in the previous iteration process is used as the basic model, and the result model of the current iteration process is obtained based on the knowledge base.
7. The method according to claim 6, wherein obtaining the result model of the current iteration process comprises: Constructing a training data set for the current iteration process based on the knowledge base, wherein the training data set for the current iteration process includes a first training data set corresponding to the first meta-skill and a second training data set corresponding to the second meta-skill; as well as Based on the basic model of the current iterative process and the training data set of the current iterative process, a result model of the current iterative process is obtained.
8. The method according to claim 7, wherein constructing the training data set of the current iteration process based on the knowledge base comprises: For the reference knowledge in the knowledge base, using the basic model of the current iteration process to generate a question sample, a knowledge sample, a reasoning sample, a first scoring sample of the knowledge sample, and a second scoring sample of the reasoning sample corresponding to the reference knowledge; as well as If the first scoring sample exceeds a first threshold and the second scoring sample exceeds a second threshold, first sample data in the first training data set and second sample data in the second training data set are constructed based on the question sample, the knowledge sample and the reasoning sample.
9. An electronic device, comprising: at least one processing unit; At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 8.
10. An information processing device, comprising: An acquisition unit, configured to acquire input information; as well as An output unit is configured to input the input information into a trained big model to obtain output information corresponding to the input information, wherein the output information includes internal knowledge output and reasoning output, wherein the trained big model is obtained based on multiple meta-skills training, wherein the meta-skills include a first meta-skill for extracting internal knowledge, a second meta-skill for generating reasoning, and a third meta-skill for self-assessment.
11. A computer-readable storage medium having a computer program stored thereon, wherein the computer program implements the method according to any one of claims 1 to 8 when executed by a processor.