Method and apparatus for generating companion reading manuscripts based on knowledge points extraction for different grades
By receiving the book names and companion reading requirements entered by the user, searching outlines of knowledge points that meet the grade requirements, and creating a prompt information set to guide the large language model to generate companion reading manuscripts, it solves the problem that it is difficult for the existing technology to generate suitable companion reading manuscripts, and realizes high-quality companion reading manuscripts automatically generated based on the grade and cognitive level, effectively improving children's reading ability and interest.
Patent Information
- Application Number
- CN202411250297.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-09-06
AI Technical Summary
It is difficult for the existing technology to automatically generate suitable companion reading manuscripts based on different grades and cognitive levels, resulting in the generated manuscripts that do not meet the learning needs of children of specific ages and cannot effectively improve their reading ability and interests.
By receiving the book names and companion reading requirements entered by the user, search the outline of knowledge points that meet the grade requirements, and create a prompt information set to guide the large language model to generate companion reading documents according to the manuscript generation requirements to ensure that the generated manuscript meets the learning needs of children of a specific age group.
The large language model automatically generates corresponding difficulty and depth knowledge points based on children's different grades and cognitive levels. The generated companion reading manuscripts can effectively improve children's reading ability and interest.
Smart Images

Figure CN118760763B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and particularly to a method and device for generating accompanying reading manuscripts based on knowledge points extraction for different grades. Background Art
[0002] In the modern education field, with the growth of personalized learning needs, educational technology is developing towards a more customized and hierarchical direction. Especially for children's reading assistance tools, such as a system for generating accompanying reading manuscripts based on knowledge points extraction for different grades, its purpose is to provide appropriate reading materials and knowledge point explanations for children through an automated method according to their age and cognitive level. Such a system not only needs to sort out the content structure of books, but also be able to stimulate children's thinking, and at the same time summarize the ability knowledge points suitable for children at different stages to achieve long-term and systematic training and improve children's reading ability and interest.
[0003] In related technologies, natural language processing technology has made remarkable progress in the field of text generation, especially in using large language models for text generation. Large language models have been widely used in fields such as dialogue systems, machine translation, and text generation. However, when it comes to generating accompanying reading manuscripts for children of different grades, large language models are difficult to automatically generate corresponding difficulty and depth of knowledge points according to the different grades and cognitive levels of children, resulting in the generated accompanying reading manuscripts not meeting the learning needs of children of a specific age group and unable to effectively improve their reading ability and interest. Summary of the Invention
[0004] Embodiments of this application provide a method and device for generating accompanying reading manuscripts based on knowledge points extraction for different grades. To provide a basic understanding of some aspects of the disclosed embodiments, a simple summary is given below. This summary part is not a general review, nor is it intended to identify key / important constituent elements or delineate the protection scope of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description below.
[0005] In a first aspect, embodiments of this application provide a method for generating accompanying reading manuscripts based on knowledge points extraction for different grades, the method including:
[0006] Receiving a book name and an accompanying reading requirement input by a user, the accompanying reading requirement including a grade requirement and a manuscript generation requirement;
[0007] According to the book name, retrieving a knowledge point outline that meets the grade requirement from a pre-constructed knowledge base, the pre-constructed knowledge base containing knowledge points at different grade stages;
[0008] According to the manuscript generation requirement, creating a set of prompt information for guiding a large language model to generate an accompanying reading manuscript;
[0009] Generate a companion reading manuscript corresponding to the book title based on the knowledge point outline and the hint information set.
[0010] Optionally, according to the requirements for generating the manuscript, create a hint information set for guiding the large language model to generate a companion reading manuscript, including:
[0011] Decompose the requirements for generating the manuscript to obtain multiple sub-questions;
[0012] Match each sub-question with the knowledge points in the pre-constructed knowledge base to obtain the knowledge point matching results for each sub-question;
[0013] Generate a hint information set for guiding the large language model to generate a companion reading manuscript according to the knowledge point matching results of each sub-question.
[0014] Optionally, match each sub-question with the knowledge points in the pre-constructed knowledge base to obtain the knowledge point matching results for each sub-question, including:
[0015] Analyze the specified part of the specific content corresponding to each sub-question and the learning requirements of the specified grade. The specific content is the book content corresponding to the book title;
[0016] Use the specified part of the book content corresponding to each sub-question and the learning requirements of the specified grade as the retrieval conditions for each sub-question;
[0017] Retrieve the knowledge points that meet the retrieval conditions of each sub-question from the pre-constructed knowledge base;
[0018] Use the retrieved knowledge points that meet the retrieval conditions of each sub-question as the knowledge point matching results for each sub-question.
[0019] Optionally, the grade requirements include a knowledge point type parameter and a cognitive level parameter corresponding to the student's grade;
[0020] Generate a hint information set for guiding the large language model to generate a companion reading manuscript according to the knowledge point matching results of each sub-question, including:
[0021] Determine the goals and expected outputs of each sub-question according to the knowledge point type parameter and the cognitive level parameter corresponding to the student's grade;
[0022] Input the goals and expected outputs of each sub-question and the knowledge point matching results of each sub-question into the large language model for learning, and output the hint information for each sub-question;
[0023] Summarize the hint information for each sub-question to obtain a hint information set for guiding the large language model to generate a companion reading manuscript.
[0024] Optionally, the prompt information set includes prompt information for each sub-question, and each sub-question is obtained by decomposing the manuscript generation requirements;
[0025] Based on the knowledge point outline and the prompt information set, generate a companion reading manuscript corresponding to the book name, including:
[0026] According to the prompt information of each sub-question and the knowledge point outline, use a pre-trained Chinese graded reading large model to reply to obtain the reply result corresponding to each sub-question;
[0027] Obtain the content logical order and importance degree of the book content corresponding to the book name;
[0028] According to the content logical order and importance degree, sort and integrate the reply results corresponding to each sub-question to obtain the final reply result;
[0029] Use the final reply result as the companion reading manuscript corresponding to the book name.
[0030] Optionally, according to the prompt information of each sub-question and the knowledge point outline, use a pre-trained Chinese graded reading large model to reply to obtain the reply result corresponding to each sub-question, including:
[0031] Through a preset search engine, search for relevant information on the book content corresponding to the book name, and the relevant information includes author information, book introduction, and reading experience;
[0032] Input the author information, book introduction, and reading experience into the pre-trained Chinese graded reading large model to output the extended text of the book content;
[0033] Input the prompt information of each sub-question, the knowledge point outline, and the book content into the large language model to output the initial reply corresponding to each sub-question;
[0034] Based on the extended text of the book content, expand the initial reply corresponding to each sub-question to obtain the reply result corresponding to each sub-question.
[0035] Optionally, generate a pre-trained Chinese graded reading large model according to the following steps, including:
[0036] Obtain the manuscript generation requirements of the companion reading manuscript and the local corpus, and the local corpus is constructed based on the companion reading text data in the children's language field;
[0037] According to the manuscript generation requirements, determine the parameters to be fine-tuned in the large language model, and the parameters to be fine-tuned are part of the parameters in the large language model applicable to the manuscript generation requirements;
[0038] Fine-tune and optimize the large language model according to the training corpus in the local corpus and the parameters to be fine-tuned, to obtain a pre-trained Chinese graded reading large model.
[0039] Optionally, fine-tuning and optimizing the large language model according to the training corpus in the local corpus and the parameters to be fine-tuned includes:
[0040] Preprocess the training corpus in the local corpus to obtain a training data set and a test data set;
[0041] Perform matrix low-rank decomposition on the parameters to be fine-tuned to obtain a decomposed first parameter matrix and a second parameter matrix. The first parameter matrix contains transformation parameters related to the columns of the original weight matrix in the large language model, and the second parameter matrix contains transformation parameters related to the rows of the original weight matrix in the large language model;
[0042] Fine-tune the large language model according to the training data set, the first parameter matrix, and the second parameter matrix;
[0043] Optimize the fine-tuned large language model according to the test data set;
[0044] Where , is the original weight matrix, is a matrix of size , represents the number of rows of matrix , represents the number of columns of matrix , is the rank in the low-rank decomposition, used to control the accuracy of the approximation and the number of parameters, is the first parameter matrix, is the second parameter matrix.
[0045] Optionally, fine-tuning the large language model according to the training data set, the first parameter matrix, and the second parameter matrix includes:
[0046] Calculate the forward propagation result of the model according to the first training data, the first parameter matrix, and the second parameter matrix. The first training data is each training data in the training data set;
[0047] Calculate the model loss value of the large language model according to the forward propagation result and a preset loss function. The preset loss function is obtained by freezing the model parameters in the original loss function of the large language model and replacing them with the first parameter matrix and the second parameter matrix;
[0048] When the model loss value reaches the minimum, obtain the fine-tuned large language model; where
[0049] The calculation formula for the forward propagation result is:
[0050]
[0051] Among them, is the forward propagation result, is the first training data, is the result of matrix low-rank decomposition;
[0052] The preset loss function is:
[0053]
[0054] Among them, is the forward propagation result at time step , is the original weight matrix, is the bias term, is the activation function, is the time step of the first training data, is the parameter to be fine-tuned, The optimization objective is to maximize the loss value of the parameter set , is the first training data, is the label of the training data, is the training data set, represents the length of the output sequence , is the log probability, are the original parameters of the large language model, is the low-rank update parameter, is the time step label.
[0055] In a second aspect, an embodiment of the present application provides a companion reading manuscript generation device based on knowledge point extraction for different grades. The device includes:
[0056] A receiving module, configured to receive the book name and the companion reading requirement input by the user, and the companion reading requirement includes grade requirements and manuscript generation requirements;
[0057] A retrieval module, configured to retrieve a knowledge point outline that meets the grade requirements from a pre-constructed knowledge base according to the book name, and the pre-constructed knowledge base contains knowledge points at different grade stages;
[0058] A creation module, configured to create a set of prompt information for guiding the large language model to generate a companion reading manuscript according to the manuscript generation requirements;
[0059] A generation module, configured to generate a companion reading manuscript corresponding to the book name based on the knowledge point outline and the set of prompt information.
[0060] The technical solutions provided by the embodiments of this application may include the following beneficial effects:
[0061] In the embodiments of this application, on the one hand, according to the book name, this application can retrieve the knowledge point outline that meets the grade requirements from the pre-constructed knowledge base, which contains knowledge points at different grade levels, ensuring the suitability and accuracy of the knowledge points, adapting to the cognitive levels and learning progress of different students, so as to realize that the large language model can automatically generate corresponding difficulty and depth of knowledge points according to the different grades and cognitive levels of children; on the other hand, according to the manuscript generation requirements, this application can create a set of prompt information for guiding the large language model to generate a companion reading manuscript. The manuscript generation requirements clarify the goals and expected results of the generated manuscript, so that the set of prompt information can guide the large language model to generate a companion reading manuscript that meets the learning needs of children of a specific age group, and can effectively improve their reading ability and interest.
[0062] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings
[0063] The drawings here are incorporated into the description and form a part of this description, showing the embodiments consistent with this application, and are used together with the description to explain the principles of this application.
[0064] Figure 1 is a schematic flowchart of a method for generating a companion reading manuscript based on knowledge point extraction for different grades provided by the embodiments of this application;
[0065] Figure 2 is a schematic diagram of the front-end page of an intelligent companion reading system provided by the embodiments of this application;
[0066] Figure 3 is a schematic diagram of the data representation of a knowledge point outline provided by this application;
[0067] Figure 4 is a schematic diagram of the data representation of a set of prompt information for guiding the large language model to generate a companion reading manuscript provided by this application;
[0068] Figure 5 is a schematic diagram of the front-end page of another intelligent companion reading system provided by this application;
[0069] Figure 6 is a schematic flowchart of a model training method provided by this application;
[0070] Figure 7 is a schematic diagram of the structure of a device for generating a companion reading manuscript based on knowledge point extraction for different grades provided by this application;
[0071] Figure 8 This is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0072] The following description and the accompanying drawings fully disclose specific implementation manners of the present application, enabling those skilled in the art to practice them.
[0073] It should be clear that the described embodiments are only a part of the embodiments of the present application, rather than all of them. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.
[0074] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are only examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0075] In the description of the present application, it should be understood that terms such as "first", "second", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0076] Currently, natural language processing technology has made remarkable progress in the field of text generation, especially in text generation using large language models. Large language models have been widely used in fields such as dialogue systems, machine translation, and text generation.
[0077] The inventors have realized that when it comes to generating accompanying reading manuscripts for children of different grades, large language models are difficult to automatically generate corresponding knowledge points with appropriate difficulty and depth according to the different grades and cognitive levels of children, resulting in the generated accompanying reading manuscripts not meeting the learning needs of children of a specific age group and being unable to effectively improve their reading ability and interest.
[0078] To solve the problem of being unable to effectively improve their reading ability and interest, the present application provides a method and device for generating accompanying reading manuscripts based on knowledge points extraction for different grades to solve the problems existing in the above-mentioned related technical problems. In the embodiments of the present application, on the one hand, according to the book name, the present application can retrieve the knowledge point outline that meets the grade requirements from a pre-constructed knowledge base. This knowledge base contains knowledge points at different grade levels, ensuring the suitability and accuracy of the knowledge points and adapting to the cognitive levels and learning progress of different students, so as to realize that the large language model can automatically generate corresponding difficulty and depth of knowledge points according to the different grades and cognitive levels of children. On the other hand, according to the manuscript generation requirements, the present application can create a set of prompt information for guiding the large language model to generate accompanying reading manuscripts. The manuscript generation requirements clarify the goals and expected results of the generated manuscript, so that the set of prompt information can guide the large language model to generate accompanying reading manuscripts that meet the learning needs of children of a specific age group, and can effectively improve their reading ability and interest. The following will be described in detail with exemplary embodiments.
[0079] The following will be combined with the attached Figure 1 - attached Figure 6 , and a method for generating accompanying reading manuscripts based on knowledge points extraction for different grades provided by the embodiments of the present application will be introduced in detail. This method can be implemented depending on a computer program and can run on a device for generating accompanying reading manuscripts based on knowledge points extraction for different grades based on the von Neumann architecture. This computer program can be integrated into an application or run as an independent tool application.
[0080] Please refer to Figure 1 , which is a schematic flowchart of a method for generating accompanying reading manuscripts based on knowledge points extraction for different grades provided by the embodiments of the present application. As Figure 1 shown, the method of the embodiments of the present application may include the following steps:
[0081] S101, receive the book name and accompanying reading requirements input by the user, and the accompanying reading requirements include grade requirements and manuscript generation requirements;
[0082] Among them, the book name is the specific book title input by the user, and the system will retrieve relevant knowledge points and information from the knowledge base according to this name. The accompanying reading requirements are the specific requirements of the user for the accompanying reading manuscript, and these requirements guide the system on how to generate the manuscript to meet the specific purposes of the user. The grade requirements are the grade levels specified by the user, which usually affect the difficulty and depth of the knowledge points, as well as the language and content complexity of the manuscript. The manuscript generation requirements are the specific expectations of the user for the generated manuscript, which may include the style, structure, content focus, etc. of the manuscript.
[0083] In some embodiments of the present application, the user opens an application capable of generating a companion reading manuscript, enters the book name in the search bar of the application, selects the grade requirement and manuscript generation requirement in the companion reading requirement options, and after the user finally triggers the submission component, the computer processor can receive the book name and companion reading requirements input by the user, and the companion reading requirements include grade requirements and manuscript generation requirements.
[0084] For example, an application capable of generating a companion reading manuscript is, for example, the "Intelligent Companion Reading System". The user opens the "Intelligent Companion Reading System" application, enters the book name "The Little Prince" in the search bar, selects the "Companion Reading Requirements" option, and selects "Grade 3" as the grade requirement in the subsequent form that appears, and enters "Needs to include an overview of the plot and an analysis of the main characters" as the manuscript generation requirement. After receiving the submission instruction, the system recognizes that the book is "The Little Prince", the grade specified by the user is "Grade 3", and determines that "Needs to include an overview of the plot and an analysis of the main characters" is the manuscript generation requirement.
[0085] S102, According to the book name, retrieve a knowledge point outline that meets the grade requirement from a pre-constructed knowledge base, and the pre-constructed knowledge base contains knowledge points at different grade levels;
[0086] Among them, the pre-constructed knowledge base is a database or information collection that has been established, storing knowledge points at different grade levels. Different grade levels refer to different learning stages that change as students' grades increase during the education process, and each stage may have different learning objectives and requirements. Meeting the grade requirement means that the retrieved knowledge point outline should match the grade level specified by the user to ensure that the difficulty and depth of the knowledge points are suitable for students at a specific grade.
[0087] In some embodiments of the present application, the specific process of retrieving a knowledge point outline that meets the grade requirement from a pre-constructed knowledge base according to the book name includes: Searching for the book content by the book name, using the original sentences in the book content and the grade requirement as search parameters, and searching for knowledge points from the pre-constructed knowledge base to obtain a knowledge point outline that meets the grade requirement.
[0088] For example, the "Intelligent Companion Reading System" retrieves a knowledge point outline suitable for third-grade students from the pre-constructed knowledge base according to the book name "The Little Prince" input by the user and the companion reading requirements.
[0089] Among them, the data table of the knowledge point outline is, for example Figure 3 as shown, including knowledge point ID, book name, grade level, knowledge point description, importance, and related concepts.
[0090] S103, Create a set of prompt information for guiding the large language model to generate a companion reading manuscript according to the manuscript generation requirement;
[0091] Among them, the large language model (LLM for short) refers to a deep learning model trained with a large amount of text data, which can generate natural language text or understand the meaning of language text. The large language model can handle various natural language tasks, such as text classification, question answering, dialogue, etc.
[0092] In some embodiments of the present application, according to the manuscript generation requirements, the specific process of creating a set of prompt information for guiding the large language model to generate a companion reading manuscript includes: decomposing the manuscript generation requirements to obtain multiple sub-questions; matching each sub-question with the knowledge points in a pre-constructed knowledge base to obtain the knowledge point matching result of each sub-question; and generating a set of prompt information for guiding the large language model to generate a companion reading manuscript according to the knowledge point matching result of each sub-question.
[0093] Among them, a sub-question is a more specific small question disassembled from the original manuscript generation requirements, and each sub-question targets a certain part or a specific requirement of the manuscript. The knowledge point matching result is the result of the matching process, that is, the corresponding knowledge point found for each sub-question. The set of prompt information is a set of prompts or instructions used to guide or stimulate the large language model to generate a specific output.
[0094] In the embodiments of the present application, the present application decomposes the complex manuscript generation requirements into multiple sub-questions, so that the system can more accurately locate the specific knowledge points required for each part. Using the pre-constructed knowledge base, the system finds the most suitable knowledge points for each sub-question to ensure the accuracy and strong pertinence of the generated manuscript content. Then, according to these knowledge point matching results, the system creates a series of precise sets of prompt information, which guide the large language model to generate high-quality companion reading manuscripts. This embodiment not only improves the efficiency and accuracy of manuscript generation, but also ensures that the manuscript can meet the personalized needs of different readers, thereby enhancing the reading experience and learning effect.
[0095] Among them, the data table of the set of prompt information for guiding the large language model to generate a companion reading manuscript is as follows Figure 4 shown. The sub-question description is the description of the sub-question disassembled from the manuscript generation requirements. The knowledge point keyword is the keyword related to the sub-question, which is used to match the knowledge points in the knowledge base. The prompt content is the prompt or question that specifically guides the large language model to generate a manuscript.
[0096] In some embodiments of the present application, the process of matching each sub-question with the knowledge points in the pre-constructed knowledge base to obtain the knowledge point matching results for each sub-question includes: analyzing the specified part of the specific content corresponding to each sub-question and the learning requirements of the specified grade, where the specific content is the book content corresponding to the book name; using the specified part of the book content corresponding to each sub-question and the learning requirements of the specified grade as the retrieval conditions for each sub-question; retrieving the knowledge points that meet the retrieval conditions of each sub-question from the pre-constructed knowledge base; and using the retrieved knowledge points that meet the retrieval conditions of each sub-question as the knowledge point matching results for each sub-question.
[0097] Among them, the grade requirements include the knowledge point type parameter and the cognitive level parameter corresponding to the student grade.
[0098] In some embodiments of the present application, the process of generating a set of prompt information for guiding the large language model to generate a companion reading manuscript according to the knowledge point matching results of each sub-question specifically includes: determining the goal and expected output of each sub-question according to the knowledge point type parameter and the cognitive level parameter corresponding to the student grade; inputting the goal and expected output of each sub-question and the knowledge point matching results of each sub-question into the large language model for learning, and outputting the prompt information for each sub-question; and summarizing the prompt information for each sub-question to obtain a set of prompt information for guiding the large language model to generate a companion reading manuscript.
[0099] Among them, the knowledge point type parameter refers to the parameter that defines the characteristics of the knowledge point, such as the theme, field, or complexity of the knowledge point, etc. These parameters help the system understand the nature of the knowledge point. The cognitive level parameter corresponding to the student grade refers to the estimated cognitive development level according to the grade where the student is located, and these parameters affect the presentation method and depth of the knowledge point. The goal of the sub-question refers to the specific educational or information transmission purpose that each sub-question needs to achieve. The expected output refers to the type or format of the result that the system or model should generate when processing the sub-question.
[0100] For example, assume that when generating a companion reading manuscript for the book *The Little Prince* for third-grade students, the knowledge point types are "story summary" and "character analysis", and the cognitive level corresponds to the reading and comprehension abilities of third-grade students. Sub-question 1 is "Who is the protagonist of *The Little Prince*?", the goal is to introduce the character of the Little Prince, and the expected output is a short character description. Sub-question 2 is "What is the main plot of *The Little Prince*?", the goal is to summarize the main plot of the story, and the expected output is a story summary. Search for knowledge points related to the sub-questions in the knowledge base and find the description of the Little Prince and the main plot of the story. Input the goals, expected outputs, and knowledge point matching results of the sub-questions into the large language model for learning. The large language model outputs prompt information for each sub-question, such as "Please describe the appearance of the Little Prince and the reason he left his planet" and "Please summarize the Little Prince's travel experiences on different planets". Finally, summarize all the prompt information of the sub-questions to form a complete set of prompt information, which will be used to guide the large language model to generate the final companion reading manuscript.
[0101] S104, generate a companion reading manuscript corresponding to the book name based on the knowledge point outline and the set of prompt information.
[0102] Among them, the companion reading manuscript is a text written to assist reading, usually including explanations, summaries, or analyses of the book content, aiming to help readers better understand and enjoy the book.
[0103] Among them, the set of prompt information includes the prompt information for each sub-question, and each sub-question is obtained by decomposing the manuscript generation requirements.
[0104] In some embodiments of the present application, the specific process of generating a companion reading manuscript corresponding to the book name based on the knowledge point outline and the set of prompt information includes: according to the prompt information and knowledge point outline of each sub-question, use a pre-trained Chinese graded reading large model to generate a response, and obtain the response result corresponding to each sub-question; obtain the content logical order and importance of the book content corresponding to the book name; according to the content logical order and importance, sort and integrate the response results corresponding to each sub-question to obtain the final response result; use the final response result as the companion reading manuscript corresponding to the book name.
[0105] Among them, the pre-trained Chinese graded reading large model is an artificial intelligence system that has been trained with a large amount of Chinese text data and can generate text according to the input prompt information and knowledge point outline. The response result is the text output generated by the large language model according to the sub-question and the prompt information. The content logical order is the structured arrangement of the book content, reflecting the organization and flow of information. The importance is the evaluation and sorting of each part of the content according to its importance or relevance.
[0106] In the embodiments of the present application, by utilizing the hint information and knowledge point outlines of each sub-question, combined with a pre-trained Chinese hierarchical reading large model, the system can generate accurate and targeted response results. These results are then intelligently sorted and integrated according to the logical order and importance of the book content, ensuring that the final accompanying reading manuscript is not only rich in content and reasonable in structure, but also prominent in key points and clear in logic. This method improves the quality and practicality of the manuscript, making it a powerful tool to help readers deeply understand the book content, especially suitable for readers with different cognitive levels, thus greatly enhancing the reading experience and learning effect.
[0107] In some embodiments of the present application, according to the hint information and knowledge point outlines of each sub-question, the specific process of obtaining the response results corresponding to each sub-question through a pre-trained Chinese hierarchical reading large model includes: searching for relevant information on the book content corresponding to the book name through a preset search engine, where the relevant information includes author information, book introduction, and reading experience; inputting the author information, book introduction, and reading experience into the pre-trained Chinese hierarchical reading large model to output extended text of the book content; inputting the hint information and knowledge point outlines of each sub-question and the book content into a large language model to output the initial response corresponding to each sub-question; and expanding the initial response corresponding to each sub-question based on the extended text of the book content to obtain the response result corresponding to each sub-question.
[0108] Among them, the preset search engine is a configured search tool or service for online searching for specific information, such as the Google search engine. The extended text is additional text generated based on the original book content and may include book introduction, author information, and reading experience.
[0109] In the embodiments of the present application, by using the preset search engine, the system can search for and collect relevant information on a specific book, such as author information, book introduction, and readers' reading experience. These information are then input into a pre-trained Chinese hierarchical reading large model to generate extended text about the book content. At the same time, the hint information and knowledge point outlines of each sub-question are also input into the large language model to generate initial responses to these sub-questions. Then, the system uses the generated extended text to expand these initial responses to obtain more comprehensive and detailed response results. This process not only enriches the content of the accompanying reading manuscript, but also improves its quality and depth, providing readers with more comprehensive reading assistance materials and enhancing their reading experience and understanding.
[0110] In the embodiments of the present application, on the one hand, according to the book name, the present application can retrieve the knowledge point outline that meets the grade requirements from a pre-constructed knowledge base. This knowledge base contains knowledge points at different grade levels, ensuring the suitability and accuracy of the knowledge points and adapting to the cognitive levels and learning progress of different students. Thus, it realizes that the large language model can automatically generate knowledge points with corresponding difficulty and depth according to the different grades and cognitive levels of children. On the other hand, according to the manuscript generation requirements, the present application can create a set of prompt information for guiding the large language model to generate accompanying reading manuscripts. The manuscript generation requirements clarify the goals and expected results of the generated manuscript, enabling the set of prompt information to guide the large language model to generate accompanying reading manuscripts that meet the learning needs of children of a specific age group and effectively improving their reading ability and interest.
[0111] For example Figure 5 As shown, after the user clicks the submission component, relevant accompanying reading manuscripts can be generated.
[0112] In some embodiments of the present application, the training process of the pre-trained Chinese graded reading large model includes: obtaining the manuscript generation requirements of the accompanying reading manuscript and the local corpus. The local corpus is constructed based on the accompanying reading text data in the children's language field; determining the parameters to be fine-tuned from the large language model according to the manuscript generation requirements. The parameters to be fine-tuned are some parameters in the large language model that are applicable to the manuscript generation requirements; fine-tuning and optimizing the large language model according to the training corpus in the local corpus and the parameters to be fine-tuned to obtain the pre-trained Chinese graded reading large model.
[0113] In the embodiments of the present application, on the one hand, the local corpus is constructed based on the accompanying reading text data in the children's language field, enabling the large language model to be optimized specifically for the children's language field based on the local corpus, ensuring that the model can more accurately capture and adapt to the context requirements of children when generating accompanying reading manuscripts, and thus generating manuscript content that is more suitable for children to read and understand. On the other hand, fine-tuning and optimizing the large language model according to the training corpus in the local corpus and the parameters to be fine-tuned. The parameters to be fine-tuned enable the model to specifically adapt to the characteristics of specific book content, including complex grammar and rich literary embellishments, so that the model will not have misunderstandings when dealing with complex grammar structures and literary embellishments. At the same time, fine-tuning and optimizing the large language model can extract more available information from the training corpus while making full use of the generation ability of the large model, thus ensuring the generation of manuscript content with a consistent style.
[0114] Please refer to Figure 6 , which is a schematic flow chart of a model training method for an elevator operation state recognition model provided by the embodiments of the present application. As Figure 6 shown, the method of the embodiments of the present application may include the following steps:
[0115] S201, obtain the manuscript generation requirements of the accompanying reading manuscript and the local corpus, where the local corpus is constructed based on the accompanying reading text data in the children's language field;
[0116] In some embodiments, when generating the local corpus, collect and screen books and text materials suitable for children of the target age group to ensure that the corpus content is rich and diverse, covering different literary genres and knowledge points. Clean the collected and screened books and text materials to remove irrelevant content such as advertisements and copyright information to obtain standard content. Text analysis can be performed on the standard content to extract keywords, phrases, and sentence structures to obtain the corpus. Label the corpus, such as difficulty level, theme classification, language style, etc., and finally obtain the local corpus.
[0117] S202, determine the parameters to be fine-tuned from the large language model according to the manuscript generation requirements, where the parameters to be fine-tuned are part of the parameters in the large language model applicable to the manuscript generation requirements;
[0118] In some embodiments of the present application, the specific process of determining the parameters to be fine-tuned from the large language model according to the manuscript generation requirements includes: obtaining the book content of the target book; performing feature analysis on the book content according to the manuscript generation requirements to obtain the content features of the target book; obtaining the model parameters corresponding to the content features of the target book from the large language model; and using the obtained model parameters as the parameters to be fine-tuned.
[0119] Among them, the target book refers to a specific book or text material for which an accompanying reading manuscript needs to be generated. The book content refers to the actual text content in the target book, including the plot, description, character dialogue, etc. The manuscript generation requirements refer to the specific standards or conditions that need to be met when generating the accompanying reading manuscript, such as language style, educational purpose, intended readers, etc. Feature analysis refers to in-depth analysis of the book content to identify and extract key features, such as vocabulary usage, sentence structure, theme elements, etc. Content features refer to the features extracted from the book content, and these features define the unique attributes and styles of the book. Model parameters refer to the variables and weights in the large language model, and these parameters determine the behavior of the model and the way of generating text.
[0120] In the embodiments of the present application, by accurately obtaining the content of the target book and performing detailed feature analysis on it, the core features and styles of the book can be captured. Combined with clear manuscript generation requirements, it ensures that the selection of model parameters is highly relevant to the book content, so that the parameters to be fine-tuned can optimize the model specifically to generate an accompanying reading manuscript that meets specific requirements.
[0121] Among them, the manuscript generation requirements include the children's age range, reading purpose, and knowledge points.
[0122] In some embodiments of the present application, according to the requirements for manuscript generation, the specific process of analyzing the features of the book content to obtain the content features of the target book includes: cleaning the text of the book content and segmenting sentences and paragraphs to obtain a standard text; using natural language processing techniques to extract multi-dimensional features of the standard text, where the multi-dimensional features include lexical features, syntactic features, semantic features, and structural features; respectively matching and associating the lexical features, syntactic features, semantic features, and structural features with the children's age range, reading purpose, and knowledge points to determine the key features related to the multi-dimensional features; and using the key features related to the multi-dimensional features as the content features of the target book.
[0123] In some embodiments of the present application, the specific process of obtaining the model parameters corresponding to the content features of the target book from the large language model includes: according to the content features of the target book, determining the specific aspects to which the model parameters of the large language model are adapted, where the specific aspects include lexical usage, sentence structure, and topic consistency; based on lexical usage, sentence structure, and topic consistency, reviewing the large language model to determine the parameters that affect the feature expression of the book content; for the lexical features, syntactic features, semantic features, and structural features, respectively extracting the word embedding matrix, self-attention and feed-forward network parameters, and layer normalization coefficients from the parameters that affect the feature expression of the book content; and using the word embedding matrix, self-attention and feed-forward network parameters, and layer normalization coefficients as the model parameters corresponding to the content features of the target book.
[0124] S203, according to the training corpus in the local corpus and the parameters to be fine-tuned, fine-tune and optimize the large language model to obtain a pre-trained Chinese graded reading large model.
[0125] In some embodiments of the present application, the specific process of fine-tuning and optimizing the large language model according to the training corpus in the local corpus and the parameters to be fine-tuned includes: preprocessing the training corpus in the local corpus to obtain a training data set and a test data set; performing matrix low-rank decomposition on the parameters to be fine-tuned to obtain a decomposed first parameter matrix and a second parameter matrix, where the first parameter matrix contains transformation parameters related to the columns of the original weight matrix in the large language model, and the second parameter matrix contains transformation parameters related to the rows of the original weight matrix in the large language model; fine-tuning the large language model according to the training data set, the first parameter matrix, and the second parameter matrix; and optimizing the fine-tuned large language model according to the test data set.
[0126] Among them, , is the original weight matrix, is a matrix of size , represents the matrix the number of rows of, represents the matrix The number of columns, is the rank in the low-rank factorization, used to control the accuracy of the approximation and the number of parameters, is the first parameter matrix, is the second parameter matrix.
[0127] In some embodiments of the present application, the specific process of fine-tuning the large language model according to the training dataset, the first parameter matrix, and the second parameter matrix includes: calculating the forward propagation result of the model according to the first training data, the first parameter matrix, and the second parameter matrix, where the first training data is each training data in the training dataset; calculating the model loss value of the large language model according to the forward propagation result and a preset loss function, where the preset loss function is obtained by freezing the model parameters in the original loss function of the large language model and replacing them with the first parameter matrix and the second parameter matrix; obtaining the fine-tuned large language model when the model loss value reaches the minimum.
[0128] Among them, the calculation formula for the forward propagation result is:
[0129] Among them, is the forward propagation result, is the first training data, is the result of matrix low-rank factorization; the preset loss function is:
[0130]
[0131] Among them, is the forward propagation result at time step , is the original weight matrix, is the bias term, is the activation function, is the time step of the first training data, is the parameter to be fine-tuned, is the optimization objective to maximize the loss value of the parameter set , is the first training data, is the label of the training data, is the training dataset, represents the length of the output sequence , is the log probability, are the original parameters of the large language model, are the low-rank update parameters, is the time step of the label.
[0132] In an embodiment of the present application, on the one hand, the local corpus is constructed based on the accompanying reading text data in the children's language field, enabling the large language model to be optimized specifically for the children's language field based on the local corpus, ensuring that the model can more accurately capture and adapt to the context requirements of children when generating accompanying reading manuscripts, and thus generating manuscript content more suitable for children to read and understand; on the other hand, according to the training corpus in the local corpus and the parameters to be fine-tuned, the large language model is fine-tuned and optimized. The parameters to be fine-tuned enable the model to specifically adapt to the characteristics of specific book content, including complex grammar and rich literary modifications, so that the model will not have misunderstandings when dealing with complex grammar structures and literary modifications. At the same time, fine-tuning and optimizing the large language model can extract more available information from the training corpus while making full use of the generation ability of the large model, thus ensuring the generation of manuscript content with a consistent style.
[0133] The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.
[0134] Please refer to Figure 7 , which shows a schematic structural diagram of an accompanying reading manuscript generation device based on knowledge point extraction for different grades provided by an exemplary embodiment of the present application. The accompanying reading manuscript generation device based on knowledge point extraction for different grades can be implemented as all or part of an electronic device through software, hardware, or a combination of both. The device 1 includes a receiving module 10, a retrieval module 20, a creation module 30, and a generation module 40.
[0135] The receiving module 10 is configured to receive the book name and accompanying reading requirements input by the user, and the accompanying reading requirements include grade requirements and manuscript generation requirements;
[0136] The retrieval module 20 is configured to retrieve a knowledge point outline that meets the grade requirements from a pre-constructed knowledge base according to the book name, and the pre-constructed knowledge base contains knowledge points at different grade stages;
[0137] The creation module 30 is configured to create a set of prompt information for guiding the large language model to generate an accompanying reading manuscript according to the manuscript generation requirements;
[0138] The generation module 40 is configured to generate an accompanying reading manuscript corresponding to the book name based on the knowledge point outline and the set of prompt information.
[0139] It should be noted that when the above-described companion reading manuscript generation device based on knowledge point extraction for different grades executes the companion reading manuscript generation method based on knowledge point extraction for different grades, only the above-mentioned division of each functional module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the above-described companion reading manuscript generation device based on knowledge point extraction for different grades and the companion reading manuscript generation method embodiment belong to the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.
[0140] The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0141] In the embodiments of the present application, on the one hand, according to the book name, the present application can retrieve the knowledge point outline that meets the grade requirements from the pre-constructed knowledge base. This knowledge base contains knowledge points at different grade levels, ensuring the suitability and accuracy of the knowledge points and adapting to the cognitive levels and learning progress of different students. Thus, it is realized that the large language model can automatically generate knowledge points with corresponding difficulty and depth according to the different grades and cognitive levels of children. On the other hand, according to the manuscript generation requirements, the present application can create a set of prompt information for guiding the large language model to generate companion reading manuscripts. The manuscript generation requirements clarify the goals and expected results of the generated manuscripts, enabling the set of prompt information to guide the large language model to generate companion reading manuscripts that meet the learning needs of children of a specific age group and effectively improving their reading ability and interest.
[0142] The present application also provides a computer-readable medium on which program instructions are stored. When the program instructions are executed by a processor, the companion reading manuscript generation method based on knowledge point extraction for different grades provided by the above-mentioned various method embodiments is implemented.
[0143] The present application also provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute the companion reading manuscript generation method based on knowledge point extraction for different grades of the above-mentioned various method embodiments.
[0144] Please refer to Figure 8 , which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 8 shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.
[0145] Among them, the communication bus 1002 is used to realize the connection and communication between these components.
[0146] Among them, the user interface 1003 may include a display screen and a camera. Optionally, the user interface 1003 may further include standard wired interfaces and wireless interfaces.
[0147] Among them, the network interface 1004 may optionally include standard wired interfaces and wireless interfaces (such as WI-FI interfaces).
[0148] Among them, the processor 1001 may include one or more processing cores. The processor 1001 connects various parts within the entire electronic device 1000 through various interfaces and circuits. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling data stored in the memory 1005, the processor 1001 performs various functions of the electronic device 1000 and processes data. Optionally, the processor 1001 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 1001 may integrate a combination of one or several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, and application programs, etc.; the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen; the modem is used for processing wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor 1001 and may be implemented separately through a single chip.
[0149] Among them, the memory 1005 may include a Random Access Memory (RAM), or may also include a Read-Only Memory. Optionally, the memory 1005 includes a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments. Optionally, the memory 1005 may also be at least one storage system located far from the aforementioned processor 1001. As Figure 8 shown, the memory 1005, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and a companion reading manuscript generation application program extracted based on knowledge points of different grades.
[0150] In Figure 8 the electronic device 1000 shown, the user interface 1003 is mainly used to provide an input interface for the user to obtain the data input by the user; while the processor 1001 can be used to call the companion reading manuscript generation application program stored in the memory 1005 and specifically perform the following operations:
[0151] Receive the book name and companion reading requirements input by the user, and the companion reading requirements include grade requirements and manuscript generation requirements;
[0152] According to the book name, retrieve the knowledge point outline that meets the grade requirements from the pre-constructed knowledge base, and the pre-constructed knowledge base contains knowledge points at different grade stages;
[0153] According to the manuscript generation requirements, create a set of prompt information for guiding the large language model to generate a companion reading manuscript;
[0154] Based on the knowledge point outline and the set of prompt information, generate a companion reading manuscript corresponding to the book name.
[0155] In one embodiment, when the processor 1001 creates a set of prompt information for guiding the large language model to generate a companion reading manuscript according to the manuscript generation requirements, it specifically performs the following operations:
[0156] Decompose the manuscript generation requirements to obtain multiple sub-questions;
[0157] Match each sub-question with the knowledge points in the pre-constructed knowledge base to obtain the knowledge point matching results for each sub-question;
[0158] Generate a set of prompt information for guiding the large language model to generate a companion reading manuscript based on the knowledge point matching results of each sub-question.
[0159] In one embodiment, when the processor 1001 executes matching each sub-question with the knowledge points in the pre-constructed knowledge base to obtain the knowledge point matching results for each sub-question, it specifically performs the following operations:
[0160] Analyze the specified part of the specific content corresponding to each sub-question and the learning requirements of the specified grade, where the specific content is the book content corresponding to the book name;
[0161] Use the specified part of the book content corresponding to each sub-question and the learning requirements of the specified grade as the retrieval conditions for each sub-question;
[0162] Retrieve the knowledge points that meet the retrieval conditions of each sub-question from the pre-constructed knowledge base;
[0163] Use the retrieved knowledge points that meet the retrieval conditions of each sub-question as the knowledge point matching results for each sub-question.
[0164] In one embodiment, when the processor 1001 executes generating a set of prompt information for guiding the large language model to generate a companion reading manuscript based on the knowledge point matching results of each sub-question, it specifically performs the following operations:
[0165] Determine the goals and expected outputs of each sub-question according to the knowledge point type parameter and the cognitive level parameter corresponding to the student grade;
[0166] Input the goals and expected outputs of each sub-question and the knowledge point matching results of each sub-question into the large language model for learning, and output the prompt information for each sub-question;
[0167] Summarize the prompt information of each sub-question to obtain a set of prompt information for guiding the large language model to generate a companion reading manuscript.
[0168] In one embodiment, when the processor 1001 executes generating a companion reading manuscript corresponding to the book name based on the knowledge point outline and the set of prompt information, it specifically performs the following operations:
[0169] According to the prompt information of each sub-question and the knowledge point outline, reply through the pre-trained Chinese graded reading large model to obtain the reply results corresponding to each sub-question;
[0170] Obtain the content logical order and importance degree of the book content corresponding to the book name;
[0171] Sort and integrate the response results corresponding to each sub-question according to the content logical order and importance to obtain the final response result;
[0172] Use the final response result as the accompanying reading manuscript corresponding to the book name.
[0173] In one embodiment, when the processor 1001 executes to obtain the response results corresponding to each sub-question by performing a response through a pre-trained Chinese hierarchical reading large model according to the hint information and knowledge point outline of each sub-question, the following operations are specifically performed:
[0174] Search for relevant information on the book content corresponding to the book name through a preset search engine, where the relevant information includes author information, book introduction, and reading experience;
[0175] Input the author information, book introduction, and reading experience into the pre-trained Chinese hierarchical reading large model to output an extended text of the book content;
[0176] Input the hint information and knowledge point outline of each sub-question and the book content into the large language model to output an initial response corresponding to each sub-question;
[0177] Expand the initial response corresponding to each sub-question based on the extended text of the book content to obtain the response result corresponding to each sub-question.
[0178] In one embodiment, when the processor 1001 generates a pre-trained Chinese hierarchical reading large model, the following operations are specifically performed:
[0179] Obtain the manuscript generation requirements of the accompanying reading manuscript and the local corpus, where the local corpus is constructed based on the accompanying reading text data in the children's language field;
[0180] Determine the parameters to be fine-tuned from the large language model according to the manuscript generation requirements, where the parameters to be fine-tuned are part of the parameters in the large language model applicable to the manuscript generation requirements;
[0181] Fine-tune and optimize the large language model according to the training corpus in the local corpus and the parameters to be fine-tuned to obtain the pre-trained Chinese hierarchical reading large model.
[0182] In one embodiment, when the processor 1001 fine-tunes and optimizes the large language model according to the training corpus in the local corpus and the parameters to be fine-tuned, the following operations are specifically performed:
[0183] Preprocess the training corpus in the local corpus to obtain a training data set and a test data set;
[0184] Perform matrix low-rank decomposition on the fine-tuning parameters to obtain the decomposed first parameter matrix and second parameter matrix. The first parameter matrix contains transformation parameters related to the columns of the original weight matrix in the large language model, and the second parameter matrix contains transformation parameters related to the rows of the original weight matrix in the large language model;
[0185] Fine-tune the large language model according to the training dataset, the first parameter matrix, and the second parameter matrix;
[0186] Optimize the fine-tuned large language model according to the test dataset;
[0187] Wherein, , is the original weight matrix, is a matrix of size , represents the number of rows of matrix , represents the number of columns of matrix , is the rank in the low-rank decomposition, which is used to control the accuracy of the approximation and the number of parameters, is the first parameter matrix, is the second parameter matrix.
[0188] In one embodiment, when the processor 1001 executes the fine-tuning of the large language model according to the training dataset, the first parameter matrix, and the second parameter matrix, it specifically performs the following operations:
[0189] Calculate the forward propagation result of the model according to the first training data, the first parameter matrix, and the second parameter matrix. The first training data is each training data in the training dataset;
[0190] Calculate the model loss value of the large language model according to the forward propagation result and the preset loss function. The preset loss function is obtained by freezing the model parameters in the original loss function of the large language model and replacing them with the first parameter matrix and the second parameter matrix;
[0191] When the model loss value reaches the minimum, obtain the fine-tuned large language model; wherein,
[0192] The calculation formula for the forward propagation result is:
[0193]
[0194] Wherein, is the forward propagation result, is the first training data, is the result of matrix low-rank decomposition;
[0195] The preset loss function is:
[0196]
[0197] Among them, is the forward propagation result at time step ; is the original weight matrix; is the bias term; is the activation function; is the first training data at time step ; is the parameter to be fine-tuned; The optimization objective is to maximize the loss value of the parameter set ; is the first training data; is the label of the training data; is the training data set; represents the length of the output sequence ; is the log probability; are the original parameters of the large language model; are the low-rank update parameters; is the label at time step .
[0198] In the embodiments of the present application, on the one hand, according to the book name, the present application can retrieve the knowledge point outline that meets the grade requirements from the pre-constructed knowledge base. The knowledge base contains knowledge points at different grade levels, ensuring the suitability and accuracy of the knowledge points, adapting to the cognitive levels and learning progress of different students, so as to realize that the large language model can automatically generate knowledge points with corresponding difficulty and depth according to the different grades and cognitive levels of children; on the other hand, according to the manuscript generation requirements, the present application can create a set of prompt information for guiding the large language model to generate a companion reading manuscript. The manuscript generation requirements clarify the goals and expected results of the generated manuscript, so that the prompt information set can guide the large language model to generate a companion reading manuscript that meets the learning needs of children of a specific age group, and can effectively improve their reading ability and interest.
[0199] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program for generating a companion reading manuscript based on the extraction of knowledge points at different grades can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, the storage medium of the program for generating a companion reading manuscript based on the extraction of knowledge points at different grades can be a magnetic disk, an optical disk, a read-only memory, or a random access memory, etc.
[0200] The above disclosure is only for the preferred embodiments of the present application. Of course, it cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.
Claims
1. A method for generating accompanying reading manuscripts based on the extraction of knowledge points of different grades, characterized in that: The method comprises: Receive the book title and accompanying reading requirements input by the user, wherein the accompanying reading requirements include grade requirements and manuscript generation requirements; According to the book title, a knowledge point outline that meets the grade requirement is retrieved from a pre-built knowledge base, wherein the pre-built knowledge base contains knowledge points at different grade levels; According to the manuscript generation requirements, a prompt information set is created for guiding the large language model to generate a companion reading manuscript; Based on the knowledge point outline and the prompt information set, a reading companion manuscript corresponding to the book title is generated; wherein, The prompt information set includes prompt information of each sub-question, and each sub-question is obtained by decomposing the document generation requirement; The step of generating a reading companion manuscript corresponding to the book title based on the knowledge point outline and the prompt information set includes: According to the prompt information of each sub-question and the outline of the knowledge points, a pre-trained Chinese graded reading model is used to answer the questions, thereby obtaining an answer result corresponding to each sub-question; Obtaining the logical order and importance of the book content corresponding to the book name; According to the logical order and importance of the content, the response results corresponding to each sub-question are sorted and integrated to obtain the final response result; The final reply result is used as the accompanying reading manuscript corresponding to the book title; wherein, Follow the steps below to generate a pre-trained Chinese graded reading model, including: Obtaining the manuscript generation requirements and the local corpus of the companion reading manuscript, wherein the local corpus is constructed based on the companion reading text data in the field of children's language; According to the document generation requirement, determining parameters to be fine-tuned from the large language model, wherein the parameters to be fine-tuned are some parameters in the large language model that are applicable to the document generation requirement; According to the training corpus in the local corpus and the parameters to be fine-tuned, the large language model is fine-tuned and optimized to obtain a pre-trained Chinese graded reading large model; wherein, The step of fine-tuning and optimizing the large language model according to the training corpus in the local corpus and the parameters to be fine-tuned includes: Preprocessing the training corpus in the local corpus to obtain a training data set and a test data set; Performing matrix low-rank decomposition on the parameters to be fine-tuned to obtain a first parameter matrix and a second parameter matrix after decomposition, wherein the first parameter matrix includes transformation parameters related to columns of an original weight matrix in the large language model, and the second parameter matrix includes transformation parameters related to rows of the original weight matrix in the large language model; Fine-tuning the large language model according to the training data set, the first parameter matrix and the second parameter matrix; According to the test data set, the fine-tuned large language model is optimized; wherein, The fine-tuning of the large language model according to the training data set, the first parameter matrix and the second parameter matrix includes: Calculate the forward propagation result of the model according to the first training data, the first parameter matrix and the second parameter matrix, wherein the first training data is each training data in the training data set; Calculating a model loss value of the large language model according to the forward propagation result and a preset loss function, wherein the preset loss function is obtained by freezing model parameters in an original loss function of the large language model and replacing them with the first parameter matrix and the second parameter matrix; When the model loss value reaches the minimum, a fine-tuned large language model is obtained.
2. The method according to claim 1, characterized in that The step of creating a prompt information set for guiding the large language model to generate a companion reading manuscript according to the manuscript generation requirement includes: Decomposing the document generation requirement to obtain multiple sub-problems; Match each sub-question with the knowledge points in the pre-built knowledge base to obtain the knowledge point matching results for each sub-question; According to the knowledge point matching result of each sub-question, a prompt information set is generated for guiding the large language model to generate a companion text.
3. The method according to claim 2, characterized in that The matching of each sub-question with a knowledge point in a pre-built knowledge base to obtain a knowledge point matching result for each sub-question includes: Analyze the designated part of the specific content corresponding to each sub-question and the learning requirements of the designated grade, wherein the specific content is the book content corresponding to the book name; Using the designated part of the book content corresponding to each sub-question and the learning requirements of the designated grade as the search conditions for each sub-question; Retrieving knowledge points that meet the retrieval conditions of each sub-question from a pre-built knowledge base; The retrieved knowledge points that meet the retrieval conditions of each sub-question are used as the knowledge point matching results of each sub-question.
4. The method according to claim 2, characterized in that: The grade requirement includes a knowledge point type parameter and a cognitive level parameter corresponding to the student grade; Generating a prompt information set for guiding the large language model to generate a companion text according to the knowledge point matching result of each sub-question includes: Determine the goal and expected output of each sub-question according to the knowledge point type parameter and the cognitive level parameter corresponding to the student grade; Inputting the target and expected output of each sub-question and the knowledge point matching result of each sub-question into the large language model for learning, and outputting prompt information of each sub-question; The prompt information of each sub-question is summarized to obtain a prompt information set for guiding the large language model to generate a companion text.
5. The method according to claim 1, characterized in that: According to the prompt information of each sub-question and the outline of the knowledge points, a reply is given through a pre-trained Chinese graded reading model to obtain a reply result corresponding to each sub-question, including: Searching for relevant information about the book content corresponding to the book title through a preset search engine, the relevant information including author information, book introduction and reading experience; Input the author information, book introduction and reading experience into a pre-trained Chinese graded reading model, and output an extended text of the book content; Input the prompt information of each sub-question, the knowledge point outline, and the book content into the large language model, and output the initial response corresponding to each sub-question; Based on the extended text of the book content, the initial response corresponding to each sub-question is expanded to obtain a response result corresponding to each sub-question.
6. The method according to claim 1, characterized in that The calculation formula of the forward propagation result is: in, is the forward propagation result, is the first training data, is the result of low-rank decomposition of the matrix; , is the original weight matrix, For a size of The matrix of Representation Matrix the number of rows, Representation Matrix The number of columns, is the rank in the low-rank decomposition, which is used to control the accuracy of the approximation and the number of parameters, is the first parameter matrix, is the second parameter matrix; The preset loss function is: in, is at the time step The forward propagation result is is the original weight matrix, is the bias term, is the activation function, is the time step The first training data, is the parameter to be fine-tuned, The optimization goal is to maximize the set of parameters The loss value, is the first training data, is the label of the training data, is the training data set, Represents the output sequence Length, is the logarithmic probability, is the original parameter of the large language model, is the low-rank update parameter, is the time step .
7. A device for generating accompanying reading manuscripts based on extraction of knowledge points for different grades using the method described in any one of claims 1 to 6, characterized in that: The device comprises: A receiving module, used to receive the book title and accompanying reading requirements input by the user, wherein the accompanying reading requirements include grade requirements and manuscript generation requirements; A retrieval module, for retrieving, according to the book title, a knowledge point outline that meets the grade requirement from a pre-built knowledge base, wherein the pre-built knowledge base contains knowledge points at different grade levels; A creation module, used to create a prompt information set for guiding the large language model to generate a companion reading manuscript according to the manuscript generation requirements; A generation module is used to generate a reading companion manuscript corresponding to the book title based on the knowledge point outline and the prompt information set.
Citation Information
Patent Citations
Reading accompanying method and device, all-in-one machine and storage medium
CN117435707A
Lever-behind child education auxiliary system based on large language model
CN117786083A