Courseware generation method and courseware generation device based on language model

Through the courseware generation method based on the language model, the courseware page title and content are automatically obtained using the chapter name, which solves the problem of teacher's low courseware production efficiency and achieves efficient and high-quality courseware generation.

CN119962486APending Publication Date: 2025-05-09GUANGZHOU SHIYUAN ELECTRONICS CO LTD +2
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311473183.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-07
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

The efficiency of existing teachers in making courseware is low, resulting in long and inefficient production of courseware. Especially for new teachers, it will take more time.

Method used

The courseware generation method based on the language model is adopted to obtain the courseware page title and content through the chapter name, combine the knowledge storage and text generation capabilities of the language model, and automatically combine and optimize the courseware content to generate a complete courseware file.

Benefits of technology

It greatly improves the efficiency of teachers' courseware production and reduces the time and energy invested by teachers in courseware production. The generated courseware is of high quality and unified, and can be directly used for teaching or modified.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119962486A_ABST
    Figure CN119962486A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to the field of data models, and discloses a courseware generation method and device based on a language model, and the method comprises the steps: obtaining a plurality of courseware page titles according to chapter names; obtaining courseware page content corresponding to each courseware page title through a target language model; combining each courseware page title and the corresponding courseware page content into one courseware page to obtain a plurality of courseware pages; and converting each courseware page into a demonstration page to obtain a plurality of demonstration pages, and combining the plurality of demonstration pages to obtain a courseware file. According to the method, a courseware which can be used for chapter teaching is generated for front-line teachers in an end-to-end manner by utilizing knowledge storage and text generation capabilities of a language model. When the courseware is used by a teacher, the complete courseware can be obtained only by inputting the chapter name of the courseware to be made, and after the courseware is obtained, the teacher can be directly used for teaching or used for teaching after being modified, so that the courseware making efficiency of the teacher is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of data models, and in particular to a courseware generation method and a courseware generation device based on a language model. Background Art

[0002] With the advent of the information age, multimedia demonstration courseware came into being. Multimedia demonstration courseware can be understood as the use of multiple multimedia materials such as pictures, texts, sounds and images to form a teaching courseware that can replace traditional blackboard writing.

[0003] In the process of classroom teaching, multimedia demonstration courseware is the carrier of teachers' teaching knowledge. The content design of multimedia demonstration courseware largely determines the scope of knowledge taught by teachers in the classroom, as well as the methods of teaching knowledge, which in turn affects the quality of a class. Courseware production is a relatively important part of teachers' pre-class preparation. When preparing multimedia demonstration courseware for a certain chapter, teachers not only need to design the content structure of the courseware according to their own knowledge background, but also refer to the courseware designed and produced by experienced teachers, and also listen to the advice of teaching and research experts, and finally form their own courseware. In this process, teachers need to invest a long time and more energy, and the efficiency of making courseware is low. For new teachers, it will take more time. Summary of the invention

[0004] The embodiment of the present invention provides a courseware generation method and a courseware generation device based on a language model to solve the problem of low efficiency of existing teachers in making courseware.

[0005] The purpose of the embodiment of the present invention is achieved through the following technical solutions:

[0006] In order to solve the above technical problems, in a first aspect, an embodiment of the present invention provides a courseware generation method based on a language model, comprising:

[0007] Get multiple courseware page titles based on chapter names;

[0008] Acquire the courseware page content corresponding to each courseware page title through the target language model;

[0009] Combining each of the courseware page titles with the corresponding courseware page content into one courseware page to obtain multiple courseware pages;

[0010] Each of the courseware pages is converted into a demonstration page to obtain multiple demonstration pages, and the multiple demonstration pages are combined to obtain a courseware file.

[0011] In one embodiment, obtaining multiple courseware page titles according to the chapter name includes:

[0012] Acquire the teaching link corresponding to the chapter name from a classroom teaching link library according to the chapter name;

[0013] Acquiring the teaching objective corresponding to the chapter name from a chapter teaching objective library according to the chapter name;

[0014] The chapter title, the teaching link and the teaching goal constitute a first input text;

[0015] The first input text is learned and analyzed by the target language model to obtain a corresponding courseware page title.

[0016] In one embodiment, the learning analysis of the first input text by the target language model to obtain the corresponding courseware page title includes:

[0017] Converting the first input text into a vector sequence prefix by using the target language model;

[0018] Predicting, according to the vector sequence prefix, a word vector that appears after the vector sequence prefix;

[0019] Transform all predicted word vectors to obtain the courseware page title.

[0020] In one embodiment, converting the first input text into a vector sequence prefix by the target language model includes:

[0021] Segmenting the first input text using the target language model to obtain a plurality of subwords and position information of each of the subwords;

[0022] Obtain the word vector corresponding to each subword from the word vector matrix, and obtain the position vector corresponding to each position information from the position vector matrix;

[0023] The word vector and the position vector corresponding to each of the sub-words are combined together to form a vector representation of each of the sub-words to obtain multiple vector representations, and the multiple vector representations are formed into a vector sequence prefix.

[0024] In one embodiment, predicting a word vector that appears after the vector sequence prefix according to the vector sequence prefix includes:

[0025] Predicting the occurrence probability of each word vector in the word vector matrix according to the vector sequence prefix, taking the word vector with the highest occurrence probability as the next word vector, and determining the position vector of the next word vector according to the position vector represented by the last vector in the vector sequence prefix;

[0026] Combining the next word vector and the position vector of the next word vector to form a next vector representation, and adding the next vector representation to the vector sequence prefix to update the vector sequence prefix;

[0027] The updated vector sequence prefix is ​​re-input into the target language model, and the occurrence probability of each word vector in the word vector matrix is ​​predicted again according to the updated vector sequence prefix to obtain the next vector representation, and then the vector sequence prefix is ​​updated until the end identifier is received or the predicted word vector reaches a predetermined length to obtain all word vectors appearing after the vector sequence prefix.

[0028] In one embodiment, the courseware page title obtained by converting all predicted word vectors includes:

[0029] All predicted word vectors are converted into corresponding sub-words, the order of all sub-words is determined according to the position vectors of all predicted word vectors, and all sub-words are arranged according to the order to obtain the courseware page title.

[0030] In one embodiment, obtaining the courseware page content corresponding to each courseware page title through the target language model includes:

[0031] Obtain the chapter textbook content, chapter supplementary teaching content, chapter exercise content and teaching suggestions corresponding to the chapter name;

[0032] The chapter name, the courseware page title, the chapter teaching material content, the chapter supplementary teaching content, the chapter exercise content and the teaching suggestion constitute a second input text;

[0033] The second input text is learned and analyzed by the target language model to obtain the courseware page content corresponding to the courseware page title.

[0034] In one embodiment, the learning analysis of the second input text by the target language model to obtain the courseware page content corresponding to the courseware page title includes:

[0035] Converting the second input text into a vector sequence prefix by using the target language model;

[0036] Predicting, according to the vector sequence prefix, a word vector that appears after the vector sequence prefix;

[0037] Convert all predicted word vectors into courseware page content.

[0038] In one embodiment, converting each of the courseware pages into a demonstration page to obtain multiple pages of demonstration pages, and combining the multiple pages of demonstration pages to obtain a courseware file includes:

[0039] Filling the courseware page title and courseware page content of each courseware page into a predefined demonstration template to obtain a process demonstration page;

[0040] By adjusting the attribute tags of the process demonstration page, the font, font size and / or color are adjusted to obtain a final demonstration page, so as to obtain multiple pages of demonstration pages;

[0041] The multiple demonstration pages are combined according to the demonstration order to obtain a courseware file.

[0042] In one embodiment, the method for constructing the target language model includes:

[0043] Establish pre-training data based on textbooks, teacher's books, supplementary teaching books, exercise books and existing courseware files;

[0044] Train the initial language model using pre-training data;

[0045] Establishing instruction fine-tuning data according to a preset input text and an expected output text corresponding to the preset input text;

[0046] The trained initial language model is trained using the preset input text in the instruction fine-tuning data to obtain a predicted output text;

[0047] The cross entropy loss is calculated according to the expected output text and the predicted output text in the instruction fine-tuning data, and the parameters of the initial language model are fine-tuned according to the cross entropy loss until the cross entropy loss is less than a set loss threshold, so as to obtain the target language model.

[0048] In order to solve the above technical problems, in a second aspect, an embodiment of the present invention provides a courseware generation device based on a language model, comprising:

[0049] at least one processor; and,

[0050] a memory communicatively connected to the at least one processor; wherein,

[0051] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the language model-based courseware generation method described in the first aspect.

[0052] To solve the above technical problems, in the third aspect, an embodiment of the present invention provides a non-volatile computer storage medium, wherein the computer storage medium stores computer executable instructions, which are executed by one or more processors to complete the language model-based courseware generation method described in the first aspect.

[0053] Compared with the prior art, the present invention has the following beneficial effects: the present invention utilizes the knowledge storage and text generation capabilities of the language model to generate an end-to-end courseware for frontline teachers that can be used for chapter teaching. When using it, teachers only need to enter the name of the chapter for which they want to make the courseware, and they can get a complete courseware. After getting the courseware, teachers can use it directly for teaching or modify it for teaching, which can greatly improve the efficiency of teachers' courseware production.

[0054] Furthermore, the language model will combine excellent existing courseware and massive knowledge related to the chapter to generate courseware with high and uniform quality, which can improve the quality of teaching.

[0055] Furthermore, by setting the attribute tags in the presentation template, the font, font size, color and other information can be processed in a unified style. The courseware file obtained after processing can be opened and played using presentation playback software. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements / modules and steps with the same reference numerals in the drawings are represented as similar elements / modules and steps. Unless otherwise specified, the figures in the drawings do not constitute proportional limitations.

[0057] Figure 1 It is a flow chart of a method for generating courseware based on a language model provided by an embodiment of the present invention;

[0058] Figure 2 It is a framework diagram of a courseware generation method based on a language model provided in an embodiment of the present invention;

[0059] Figure 3 The embodiment of the present invention provides Figure 1 Specific flow diagram of step 10 in FIG.

[0060] Figure 4 The embodiment of the present invention provides Figure 3 Specific flow diagram of step 101;

[0061] Figure 5 The embodiment of the present invention provides Figure 3 Specific flow diagram of step 102;

[0062] Figure 6 is a structural diagram of a language model provided by an embodiment of the present invention;

[0063] Figure 7 The embodiment of the present invention provides Figure 1 Specific flow diagram of step 20;

[0064] Figure 8 The embodiment of the present invention provides Figure 1 Specific flow diagram of step 40;

[0065] Fig. 9 It is a structural schematic diagram of a courseware generation device based on a language model provided by an embodiment of the present invention;

[0066] Fig.10 It is a structural diagram of another language model-based courseware generation device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0067] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several variations and improvements may be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0068] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0069] It should be noted that, if there is no conflict, the various features in the embodiments of the present invention can be combined with each other, and all are within the protection scope of the present invention. In addition, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart.

[0070] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0071] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.

[0072] In the existing teaching process, teachers are required to produce multimedia demonstration courseware for a certain chapter by themselves, which takes a long time of teachers. The efficiency of teachers in producing courseware is low, and the quality of courseware produced by different teachers is also uneven. Courseware of poor quality will affect the quality of classroom teaching and lead to poor classroom effects.

[0073] Teachers are inefficient in producing courseware. On the one hand, teachers’ knowledge reserves are limited after all. They need to consult various materials to acquire relevant knowledge. At the same time, they also need to screen and combine the acquired knowledge to produce high-quality courseware. On the other hand, different teachers’ teaching hours, teaching qualifications and proficiency in courseware production tools will also affect teachers’ production efficiency.

[0074] The language model has the ability to store knowledge. It can store a large amount of knowledge, retrieve and combine knowledge based on the input text, and screen out effective knowledge points. The language model has the ability to generate text, and can generate corresponding output results based on user input. Therefore, the language model can be used to generate a courseware that can be used for chapter teaching for front-line teachers end-to-end. At the same time, the courseware content can also be input into a predefined presentation template, and the corresponding courseware can be obtained by adjusting the properties of the presentation template. Teachers only need to adjust the content in the template to render the courseware content into the corresponding presentation file, which is not limited by the proficiency of the courseware production tool.

[0075] In order to improve the efficiency of courseware production and to improve the quality of courseware, in an embodiment of the present invention, a courseware generation method based on a language model is provided, including: obtaining multiple courseware page titles according to the chapter name; obtaining the courseware page content corresponding to each of the courseware page titles through a target language model; combining each of the courseware page titles with the corresponding courseware page content into a courseware page to obtain multiple courseware pages; converting each of the courseware pages into a demonstration page to obtain multiple demonstration pages, and combining the multiple demonstration pages to obtain a courseware file.

[0076] The embodiment of the present invention utilizes the knowledge storage and text generation capabilities of the language model to generate a courseware that can be used for chapter teaching for front-line teachers end-to-end. When using it, teachers only need to enter the name of the chapter for which they want to make the courseware, and they can get a complete courseware. After obtaining the courseware, teachers can use it directly for teaching or modify it for teaching, which can greatly improve the efficiency of teachers' courseware production. Moreover, the language model will combine excellent stock courseware and massive knowledge related to the chapter to generate courseware, and the courseware quality is high and uniform, which can improve the quality of teaching.

[0077] Embodiment 1:

[0078] See also Figure 1 and Figure 2 This embodiment provides a courseware generation method based on a language model, and the courseware generation method based on a language model includes:

[0079] Step 10: Get multiple courseware page titles based on the chapter name.

[0080] Among them, the chapter name is the name of the content that needs to be learned in a certain section of the textbook, such as "The Crow Drinks Water", "The Tadpole Looking for Its Mother", etc.

[0081] A courseware file includes multiple courseware pages (it can also be the demonstration page described below). The courseware page title here refers to the title corresponding to the courseware page. The courseware page title is intended to describe the content to be presented in this courseware page in short text. For example, the courseware page title is "Guess the Riddle", "Background Knowledge", "Character Introduction" and "Background Culture".

[0082] In one embodiment, the teaching link corresponding to the chapter name can be obtained from a classroom teaching link library based on the chapter name; the teaching goal corresponding to the chapter name can be obtained from a chapter teaching goal library based on the chapter name; the chapter name, the teaching link and the teaching goal form a first input text; the first input text is studied and analyzed through the target language model to obtain a corresponding courseware page title.

[0083] Among them, the classroom teaching link library is a pre-established knowledge base used to store the teaching links used in the teaching process; the chapter teaching goal library is a pre-established knowledge base used to store the teaching goals that need to be achieved in the teaching process. The chapter teaching goal library can be obtained by extracting the teaching goals of each chapter from the teacher's book; the classroom teaching link library can be determined by front-line teachers or professional teaching and research personnel.

[0084] In actual use, the teacher only needs to input the chapter name, the teaching links, and the teaching objectives at the preset positions, and the target language model will form the first input text with the chapter name, the teaching links, and the teaching objectives; through the learning and analysis of the first input text by the target language model, the corresponding courseware page title can be obtained.

[0085] Among them, there is one or more teaching links. When there are multiple teaching links, the teaching links are filled in the preset positions in sequence, and the target language model will form the corresponding first input text with the chapter name, the teaching links, and the teaching objectives to obtain the courseware page title. That is, each teaching link will correspond to at least one courseware page title.

[0086] Input the teaching objective of the chapter and a certain teaching link into the target language model, and the target language model outputs the courseware page title used for this teaching link.

[0087] The input text input into the target language model can be in the following format:

[0088] {

[0089] "instruction": "Please generate a list of courseware titles required for the classroom teaching link {teaching link} of {chapter name}, requiring teaching with the corresponding courseware and being able to achieve the teaching objective: {teaching objective}",

[0090] }

[0091] Among them, the content to be filled and set in {} is for different chapters and teaching links.

[0092] All the text after the "instruction" field constitutes the first input text, and the target language model extracts the first input text to obtain the courseware page title matching it through the first input text.

[0093] Taking the chapter "The Crow Drinks Water" as an example, the following teaching objectives can be extracted from the first-grade teacher's book of the unified textbook for primary school Chinese:

[0094] (1) Recognize 11 new characters such as "wu, ya" and 1 radical of the inverted radical; be able to write 5 new characters such as "zhi, shi".

[0095] (2) Read the text correctly and fluently, understand the process of the crow drinking water, and recognize natural paragraphs.

[0096] (3) Understand the principle that when encountering difficulties, one should think carefully and actively find ways to solve them.

[0097] At the same time, for the classroom of the first-grade primary school Chinese subject, the following classroom teaching links can be defined by front-line teachers or teaching and research personnel:

[0098] (1) Course introduction: Through the course introduction link, students are guided into the learning situation, their learning interest is stimulated, and the connection between learning tasks is established, which is in line with modern teaching suggestions.

[0099] (2) New course teaching: The new course teaching session focuses on the internal connections and differences between students’ various subjects, creates a learning environment, guides students to pay attention to life experience, and reflects the requirements of teaching suggestions.

[0100] (3) Course summary: The course summary section organizes the learning content, refines the teaching points, and provides conditions for students' personalized growth, which is in line with the teaching suggestions on focusing on deep thinking and feedback.

[0101] (4) Homework: Homework focuses on students’ characteristics, provides learning support, and guides students to solve problems in real life, reflecting the emphasis of teaching suggestions on practical learning.

[0102] For example, for the "Course Introduction" teaching link of "The Crow Drinks Water", the input text input into the target language model can be in the following format:

[0103] {

[0104] "instruction": "Please generate a list of courseware titles for the classroom teaching link {Course introduction: through the course introduction link...} for {The crow drinks water}, and require the use of the corresponding courseware for teaching to achieve the teaching objectives: {Teaching objectives: (1) Recognize 11 new characters such as "crow, crow...}"

[0105] }

[0106] The target language model extracts text from the "instruction" field and generates a courseware page title based on the above generation process.

[0107] In actual application scenarios, for the teaching link of "Course Introduction", the courseware page title can be "Guess the Riddle", "Background Knowledge" and other courseware page titles; for the teaching link of "New Lesson Teaching", courseware page titles such as "First Reading of the Text" and "Word Learning" can be generated.

[0108] Step 20: Obtain the courseware page content corresponding to each courseware page title through the target language model.

[0109] The target language model described in this embodiment may be a language model based on Llama 2, a language model based on Baichuan, a language model based on LLaMA, etc., or a language model including a multi-layer transformer structure.

[0110] In one embodiment, one courseware page title corresponds to one courseware page content, and the courseware page content corresponding to each courseware page title is obtained through the target language model, thereby obtaining multiple pages of courseware page content.

[0111] Step 30: Combine each of the courseware page titles and the corresponding courseware page content into one courseware page to obtain multiple courseware pages; convert each of the courseware pages into a demonstration page to obtain multiple demonstration pages.

[0112] The demonstration page may be a presentation in the form of PPT (Microsoft Office PowerPoint).

[0113] In one embodiment, courseware pages in document form are converted into presentation pages.

[0114] Step 40: Combine the multiple demonstration pages to obtain a courseware file.

[0115] In one embodiment, multiple pages of demonstration pages are typeset and combined according to the display order of the demonstration pages during the teaching process to obtain a courseware file.

[0116] In an embodiment of the present invention, the knowledge storage and text generation capabilities of the language model are utilized to generate a courseware that can be used for chapter teaching for front-line teachers end-to-end. When using it, teachers only need to enter the name of the chapter for which they want to make the courseware, and they can get a complete courseware. After obtaining the courseware, teachers can use it directly for teaching or modify it for teaching, which can greatly improve the efficiency of teachers' courseware production. Moreover, the language model will combine excellent stock courseware and massive knowledge related to the chapter to generate courseware, and the courseware quality is high and uniform, which can improve the quality of teaching.

[0117] In one embodiment, see Figure 3 In step 10, the first input text is learned and analyzed by the target language model to obtain the corresponding courseware page title, which specifically includes:

[0118] Step 101: Convert the first input text into a vector sequence prefix through the target language model.

[0119] The first input text consists of multiple Chinese characters, and needs to be converted into a word vector that can be calculated by the model, and the word vector and its corresponding position vector form a vector representation, and the vector representation is integrated into a vector sequence prefix. That is, all the characters after the "instruction" field in the above are converted into a vector sequence prefix.

[0120] Step 102: predicting the word vector that appears after the vector sequence prefix according to the vector sequence prefix.

[0121] Step 103: Convert all predicted word vectors to obtain the courseware page title.

[0122] In one embodiment, the word vector needs to be converted into text to obtain the courseware page title.

[0123] Continue reading Figure 4 In step 101, it specifically includes:

[0124] Step 1011: segment the first input text using the target language model to obtain a plurality of sub-words and position information of each of the sub-words.

[0125] In one embodiment, a tokenizer (BPE, WordPiece, etc.) is used to segment the first input text into token sequences.

[0126] Combination Figure 6 , assuming that the token sequence composed of multiple subwords is t1, t2, t3, ... t k , where t1, t2, t3, ... t k is a subword, predict the next subword (token) t k+1 The probability of occurrence, that is, p(t k+1 |t1,t2,t3,...t k ), select the word vector corresponding to the subword with the highest probability as the next word vector, add the next word vector to the current vector sequence prefix, and continuously fill the vector sequence prefix. This cycle is repeated in this way until the generated vector sequence reaches the predetermined length or meets other loop end conditions, and a complete string t1, t2, t3, ...t is obtained. k ,t k+1 ,...t n Among them, t k+1 ,...t n is the predicted character string. The predetermined length may be determined according to the actual situation and is not specifically limited here. The loop end condition may be the loop end symbol of the language model.

[0127] Step 1012: Obtain the word vector corresponding to each sub-word from the word vector matrix, and obtain the position vector corresponding to each position information from the position vector matrix.

[0128] Each row of the word vector matrix (also called token embedding matrix) is a vector corresponding to a word in the vocabulary. Assuming there are 20,000 words in the vocabulary, and the vector corresponding to each word is 100-dimensional, the size of the word vector matrix is ​​20,000X100.

[0129] The position vector matrix (also called position embedding matrix) is the number of words in a sequence, such as 1, 2, 3, etc. Each position number corresponds to a position vector.

[0130] Step 1013: Combine the word vector and the position vector corresponding to each of the sub-words to form a vector representation of each of the sub-words to obtain multiple vector representations, and form a vector sequence prefix with the multiple vector representations.

[0131] The combination of word vectors and position vectors can be in a concatenated form or an array form, which is not specifically limited in this embodiment. For example, the word vector corresponding to subword t1 is a, and the position vector is b. A and b are concatenated together to form a vector representation ab of each of the subwords; or a and b are combined together in the form of vectors to form a vector representation [a, b] of each of the subwords.

[0132] Continue reading Figure 5 In step 102, it specifically includes:

[0133] Step 1021: predict the occurrence probability of each word vector in the word vector matrix according to the prefix of the vector sequence, take the word vector with the highest occurrence probability as the next word vector, and determine the position vector of the next word vector according to the position vector represented by the last vector in the prefix of the vector sequence.

[0134] In this embodiment, the occurrence probability of each word vector is predicted, and the word vector with the highest occurrence probability is used as the next word vector.

[0135] Assuming that the order of appearance of the characters is taken as the position of the characters in the input text, the position vector of the next word vector is obtained as follows: the position vector represented by the last vector in the prefix of the vector sequence is the vector corresponding to 200, and the position vector of the next word vector is the vector corresponding to 201.

[0136] Step 1022: Combine the next word vector and the position vector of the next word vector to form a next vector representation, and add the next vector representation to the vector sequence prefix to update the vector sequence prefix.

[0137] Step 1023: Re-input the updated vector sequence prefix into the target language model, and predict the occurrence probability of each word vector in the word vector matrix again based on the updated vector sequence prefix to obtain the next vector representation, and then update the vector sequence prefix until the end identifier is received or the predicted word vector reaches a predetermined length to obtain all word vectors that appear after the vector sequence prefix.

[0138] The predetermined length may be determined according to actual conditions and is not specifically limited here. The end identifier may be a loop end identifier of the language model.

[0139] In this embodiment, the vector sequence prefix is ​​continuously updated in a text chain manner to obtain the courseware page title.

[0140] In one embodiment, step 103 specifically includes: converting all predicted word vectors into corresponding sub-words, determining the sequence of all sub-words according to the position vectors of all predicted word vectors, arranging all sub-words according to the sequence, and obtaining the courseware page title.

[0141] In one embodiment, see Figure 7 In step 20, it specifically includes:

[0142] Step 201: Obtain the chapter teaching material content, chapter teaching supplementary content, chapter exercise content and teaching suggestions corresponding to the chapter name.

[0143] Step 202: The chapter name, the courseware page title, the chapter teaching material content, the chapter supplementary teaching content, the chapter exercise content and the teaching suggestion constitute a second input text.

[0144] Step 203: Perform learning analysis on the second input text through the target language model to obtain the courseware page content corresponding to the courseware page title.

[0145] In one embodiment, the second input text is converted into a vector sequence prefix through the target language model; the word vector that appears after the vector sequence prefix is ​​predicted based on the vector sequence prefix; and all predicted word vectors are converted into courseware page content. The specific implementation method can refer to step 101-step 103. The difference between the two is that the input text is different, the predicted output text is different, and the specific execution method is the same, which will not be repeated here.

[0146] In this embodiment, based on the above-generated courseware page title, the language model is used to generate courseware page content that meets the teaching suggestions by comprehensively utilizing the chapter textbook content, supplementary teaching materials, exercise books, curriculum standards issued by the Ministry of Education, and teaching suggestions for the subject section to which the chapter belongs. For each courseware page title, the language model can be used to generate the corresponding courseware page content through the following prompt:

[0147] {

[0148] "instruction": "Please generate specific content for the courseware page {courseware page title} used in the classroom teaching of chapter {chapter name}. When using the generated courseware page content for teaching, it can meet the {teaching suggestions in the curriculum standard}. When generating courseware content, you can refer to the following information. Chapter content: {chapter content}; Chapter corresponding teaching materials content: {chapter corresponding teaching materials content}; Chapter related exercises: {chapter related exercises}"

[0149] }

[0150] The {} contains the content that needs to be filled in different chapters.

[0151] In one embodiment, see Figure 8 In step 40, it specifically includes:

[0152] Step 401: Fill the courseware page title and courseware page content of each courseware page into a predefined presentation template to obtain a process presentation page.

[0153] Step 402: By adjusting the attribute tags of the process demonstration page, the font, font size and / or color are adjusted to obtain a final demonstration page, so as to obtain multiple pages of demonstration pages.

[0154] The presentation template can be in XML format. By setting the attribute tags in XML, the font, font size, color and other information can be processed in a unified style. The courseware file obtained after the above processing can be opened and played using the presentation playback software.

[0155] Step 403: Combine the multiple demonstration pages according to the demonstration order to obtain a courseware file.

[0156] In one embodiment, the courseware page title and the courseware page content can be combined into a complete courseware page. Combine multiple courseware pages to obtain a completed courseware. The courseware page title and the courseware page content are text information and cannot be opened using presentation playback software (e.g., wps, powerpoint, etc.). The courseware page title and the courseware page content are filled into a predefined template adapted for speech software, in which there are fields corresponding to the title and the text content, such as "title" and "content". The courseware page title is filled into the "title" field, and the courseware page content is filled into the "content" field. Finally, by setting the attribute tags in the presentation template, the information such as font, font size, color, etc. is processed in a unified style. The courseware file obtained after the above processing can be opened and played using presentation playback software.

[0157] In one embodiment, it is also necessary to obtain the target language model used above through training and instruction fine-tuning. The method for constructing the target language model includes: establishing pre-training data based on textbooks, teacher's books, supplementary teaching books, exercise books and existing courseware files; training the initial language model through the pre-training data to obtain a trained initial language model; establishing instruction fine-tuning data based on a preset input text and an expected output text corresponding to the preset input text; training the trained initial language model through the preset input text in the instruction fine-tuning data to obtain a predicted output text; calculating the cross-entropy loss based on the expected output text and the predicted output text in the instruction fine-tuning data, and fine-tuning the parameters of the initial language model based on the cross-entropy loss until the cross-entropy loss is less than the set loss threshold to obtain the target language model. Among them, the loss threshold can be set according to the actual situation and is not specifically limited here.

[0158] In actual application scenarios, the language model for obtaining the courseware page title and the language model for the courseware page content can be the same language model or two different language models. If they are two different language models, you only need to change the samples and perform training and instruction fine-tuning as described above. If they are the same language model, you need to go through at least two stages of training and instruction fine-tuning so that the language model can generate the courseware page title and courseware page content.

[0159] The details are as follows:

[0160] To obtain a language model that can generate courseware page titles, two stages of model training are required: pre-training and instruction fine-tuning. When pre-training and instruction fine-tuning the language model, the training data is input into Figure 5 In the structure shown, the language model calculates the cross entropy loss based on the probability distribution of the softmax output and the actual expected output to optimize the training model and adjust the model parameters.

[0161] Pre-training: Use unsupervised pre-training data to train the language model so that the language model has the relevant knowledge to generate the courseware page title. The pre-training data usually required includes textbooks used by students, teacher's books, supplementary books, exercise books and courseware in the stock courseware library.

[0162] Instruction fine-tuning: Use supervised training data (i.e. the preset input text mentioned above) to train the language model so that the language model has the ability to generate courseware page titles using existing knowledge. The data format required for instruction fine-tuning is usually:

[0163] {

[0164] “Instruction”: “Please generate a list of courseware titles for the classroom teaching session {teaching session name} for {chapter name}, and require the use of corresponding courseware for teaching to achieve the teaching objectives: {teaching objective content}”,

[0165] "output": {list of courseware page titles corresponding to the teaching session}

[0166] }

[0167] The courseware page title list corresponding to the output field can be obtained from the stock courseware library. The instruction fine-tuning data composed of instruction and output can be obtained through manual annotation.

[0168] In actual application scenarios, to obtain a model that can generate courseware page content, two stages of model training are also required, namely: pre-training and instruction fine-tuning.

[0169] Pre-training: Use unsupervised pre-training data to train the language model so that the language model has the relevant knowledge to generate the courseware page content corresponding to the courseware page title. The pre-training data usually required includes textbooks used by students, teacher books, supplementary books, exercise books, and courseware in the stock courseware library.

[0170] Instruction fine-tuning: Use supervised training data (i.e. the preset input text mentioned above) to train the language model so that the language model has the ability to generate the corresponding courseware page content based on the courseware page title. The data format required for instruction fine-tuning is usually:

[0171] {

[0172] "instruction": "Please generate specific content for the courseware page {courseware page title} used in the classroom teaching of chapter {chapter name}. When using the generated courseware page content for teaching, it can meet the {teaching suggestions in the curriculum standard}. When generating courseware content, you can refer to the following information. Chapter content: {chapter content}; Chapter corresponding teaching materials content: {chapter corresponding teaching materials content}; Chapter related exercises: {chapter related exercises}"

[0173] "output": {the content of the courseware page corresponding to the courseware page title}

[0174] }

[0175] Among them, in the output field, the courseware page content corresponding to the courseware page title can be obtained from the existing courseware library.

[0176] The present invention utilizes the knowledge storage and text generation capabilities of the language model to generate a courseware that can be used for chapter teaching for frontline teachers end-to-end. When using it, teachers only need to enter the name of the chapter to be produced, and they can get a complete courseware. After obtaining the courseware, teachers can use it directly for teaching or modify it for teaching, which can greatly improve the efficiency of teachers' courseware production.

[0177] Embodiment 2:

[0178] Teachers need to constantly polish courseware according to actual teaching situations. In the prior art, after a certain version of courseware is produced, teachers use this version of courseware for classroom teaching. After class, they optimize the courseware in a targeted manner according to students' responses, classroom atmosphere, teacher-student interaction, etc. Generally, teachers optimize courseware manually, which is not only inefficient, but also depends on the teacher's knowledge reserve. If the teacher's knowledge reserve is small, the optimized content may not meet the actual teaching needs.

[0179] In order to solve the aforementioned problems, an embodiment of the present invention provides a method for optimizing courseware, including: collecting recording information during a teacher's lecture and page-turning events for the courseware, and obtaining the voice text content corresponding to each page of the courseware according to the recording information and the page-turning events; obtaining the chapter teaching objectives of the courseware, and obtaining the teaching objective items corresponding to each page of the courseware according to the chapter teaching objectives; for each page of the courseware, using a target language model to perform learning and analysis on the courseware page content, the voice text content, the teaching objective items, and chapter-related information of the courseware page to obtain optimized courseware page content; and optimizing the courseware through the optimized courseware page content to obtain optimized courseware.

[0180] In one embodiment, the collecting of recording information during the teacher's teaching process and page turning events for the courseware, and obtaining the voice text content corresponding to each page of the courseware according to the recording information and the page turning event include: obtaining the dwell time interval [t start_i ,t end_i ], where t start_i Indicates the time when the i-th page of the courseware is turned, t end_i Indicates the moment of turning over the i-th courseware page; converts the recording information into a continuous audio file; for each courseware page, extracts the corresponding audio slice from the audio file according to the dwell time interval corresponding to the courseware page; identifies the voice text content from the audio slice to obtain the voice text content corresponding to each courseware page.

[0181] In one embodiment, the method of obtaining the chapter teaching objectives of a courseware and obtaining the teaching objective items corresponding to each page of the courseware according to the chapter teaching objectives includes: dividing the chapter teaching objectives into at least one teaching objective item, and converting each of the teaching objective items into a teaching objective vector; converting the courseware page content of each page of the courseware page into a courseware page content vector; calculating the similarity between each of the courseware page content vectors and each of the teaching objective vectors in turn, and obtaining the teaching objective vector with the highest similarity; and using the teaching objective item corresponding to the teaching objective vector with the highest similarity as the teaching objective item corresponding to the courseware page.

[0182] In one embodiment, the target language model is used to perform learning and analysis on the courseware page content, the voice text content, the teaching objective items and the chapter association information of the courseware page to obtain the optimized courseware page content, including: obtaining the chapter teaching material content, chapter teaching aid content, chapter exercise content and teaching suggestions corresponding to the courseware to obtain chapter association information; the courseware page content, the voice text content, the teaching objective items and the chapter association information constitute an input text; the input text is learned and analyzed by the target language model to obtain the optimized courseware page content.

[0183] In one embodiment, the learning and analysis of the input text by the target language model to obtain optimized courseware page content includes: converting the input text into a vector sequence prefix by the target language model; predicting the word vector that appears after the vector sequence prefix based on the vector sequence prefix; and converting all predicted word vectors into optimized courseware page content.

[0184] In one embodiment, converting the input text into a vector sequence prefix through the target language model includes: segmenting the input text through the target language model to obtain multiple sub-words and position information of each sub-word; obtaining the word vector corresponding to each sub-word from the word vector matrix, and obtaining the position vector corresponding to each position information from the position vector matrix; combining the word vector and the position vector corresponding to each sub-word to form a vector representation of each sub-word to obtain multiple vector representations, and forming the multiple vector representations into a vector sequence prefix.

[0185] In one embodiment, predicting the word vector that appears after the vector sequence prefix according to the vector sequence prefix includes: predicting the occurrence probability of each word vector in the word vector matrix according to the vector sequence prefix, taking the word vector with the highest occurrence probability as the next word vector, and determining the position vector of the next word vector according to the position vector of the last vector representation in the vector sequence prefix; combining the next word vector and the position vector of the next word vector to form a next vector representation, and adding the next vector representation to the vector sequence prefix to update the vector sequence prefix; re-inputting the updated vector sequence prefix into the target language model, predicting the occurrence probability of each word vector in the word vector matrix again according to the updated vector sequence prefix to obtain the next vector representation, and then updating the vector sequence prefix until an end identifier is received or the predicted word vector reaches a predetermined length to obtain all word vectors that appear after the vector sequence prefix.

[0186] In one embodiment, the method for constructing the target language model includes: establishing pre-training data based on textbooks, teacher's books, teaching supplementary books, exercise books, classroom transcripts and existing courseware files; training the initial language model through the pre-training data; establishing instruction fine-tuning data based on preset input text and expected output text corresponding to the preset input text; training the trained initial language model through the preset input text in the instruction fine-tuning data to obtain predicted output text; calculating the cross-entropy loss based on the expected output text and the predicted output text in the instruction fine-tuning data, and fine-tuning the parameters of the initial language model based on the cross-entropy loss until the cross-entropy loss is less than a set loss threshold to obtain the target language model.

[0187] In one embodiment, the recording information includes one or more of teacher teaching voice information, student feedback voice information, and teacher-student interaction voice information.

[0188] In the embodiment of the invention, the recording information of the teacher's teaching process is obtained, and the voice text content of each page of the courseware is obtained according to the recording information. With the help of the text understanding and text generation capabilities of the language model, the language model can integrate the voice text content of the classroom, the teaching target items and the relevant information of the chapter, modify and optimize the existing courseware pages, continuously iterate to improve the quality of the courseware, and then improve the quality of classroom teaching. This way of optimizing courseware is highly efficient, and the knowledge reserve of the language model is greater than the knowledge reserve of the teacher, which can ensure the quality of the optimized content. After multiple iterations of optimization, it can ensure that the actual teaching needs are met.

[0189] Embodiment 3:

[0190] Based on the above-mentioned embodiments 1 to 2, Fig. 9 As shown, this embodiment also provides a courseware generation device based on a language model, including: at least one processor 21; and a memory 22 communicatively connected to the at least one processor 21; wherein the memory 22 stores instructions executable by the at least one processor 21, and the instructions are executed by the at least one processor 21 so that the at least one processor 21 can execute the courseware generation method based on the language model described in the aforementioned embodiment.

[0191] The processor 21 and the memory 22 may be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.

[0192] The memory 22 is a non-volatile computer-readable storage medium, which can be used to store non-volatile software programs, non-volatile computer executable programs and modules, such as program instructions / modules corresponding to the language model-based courseware generation method in the embodiment of the present invention. The processor 21 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions and modules stored in the memory 22, that is, the language model-based courseware generation method of the above method embodiment is implemented, which specifically includes: obtaining multiple courseware page titles according to the chapter name; obtaining the courseware page content corresponding to each of the courseware page titles through the target language model; combining each of the courseware page titles with the corresponding courseware page content into a courseware page to obtain multiple courseware pages; converting each of the courseware pages into a demonstration page to obtain multiple demonstration pages, and combining the multiple demonstration pages to obtain a courseware file.

[0193] The memory 22 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and an application required by at least one function; the data storage area may store data created according to the use of a courseware generation device based on a language model, etc. In addition, the memory 22 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other non-volatile solid-state storage device. In some embodiments, the memory 22 may optionally include a memory remotely arranged relative to the processor 21, and these remote memories may be connected to the courseware generation device based on the language model via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0194] The one or more modules are stored in the memory 22, and when executed by the one or more processors 21, execute the language model-based courseware generation method in any of the above method embodiments, for example, execute the method steps of the language model-based courseware generation method described above.

[0195] The above product can execute the method provided by the embodiment of the present invention, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not described in detail in this embodiment, please refer to the method provided by the embodiment of the present invention.

[0196] Furthermore, the courseware generation device based on the language model may also include: Wi-Fi device 23, display screen 24, Bluetooth device 25, audio circuit 26, power system 27, peripheral interface 28, sensor module 29, data conversion module 30 and other components. These components can communicate through one or more communication buses or signal lines. Those skilled in the art can understand that Fig.10 The hardware structure shown in the figure does not constitute a limitation on the courseware generation device based on the language model. The courseware generation device based on the language model may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0197] The processor 21 is the control center of the courseware generation device based on the language model, and uses various interfaces and lines to connect various parts of the courseware generation device based on the language model. By running or executing the application program stored in the memory 22, and calling the data and instructions stored in the memory 22, the processor 21 performs various functions and processes data of the courseware generation device based on the language model. In some embodiments, the processor 21 may include one or more processing units; the processor 21 may also integrate an application processor 21 and a modem processor 21; wherein the application processor 21 mainly processes the operating system, user interface and application program, and the modem processor 21 mainly processes wireless communication. It can be understood that the modem processor 21 may not be integrated into the processor 21.

[0198] In some other embodiments of the present invention, the processor 21 may also include an artificial intelligence (AI) chip. The learning and processing capabilities of the artificial intelligence chip include image understanding capabilities, natural language understanding capabilities, and speech recognition capabilities. The artificial intelligence chip can enable the courseware generation device based on the language model to have better performance, longer battery life, and better security and privacy. For example, if the courseware generation device based on the language model processes data through the cloud, the data needs to be uploaded and processed before returning the result, which is very inefficient under the existing technical conditions. If the local end of the courseware generation device based on the language model has a strong AI learning ability, then the courseware generation device based on the language model does not need to upload the data to the cloud, and can be directly processed on the local end, thereby improving the security and privacy of the data while improving the processing efficiency.

[0199] The memory 22 is used to store applications and data. The processor 21 executes various functions and data processing of the courseware generation device based on the language model by running the applications and data stored in the memory 22. The memory 22 mainly includes a program storage area and a data storage area, wherein the program storage area can store an operating system and an application required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area can store data created when using the courseware generation device based on the language model (such as audio data, video data, etc.). In addition, the memory 22 can include a high-speed random access memory 22, and can also include a non-volatile memory 22, such as a disk memory 22, a flash memory device or other non-volatile solid-state memory 22, etc. The memory 22 can store various operating systems, such as an operating system developed by Apple, an operating system developed by Microsoft, etc.

[0200] The display screen 24 is used to display images, videos, etc. The display screen 24 can be a touch screen. In some embodiments, the courseware generation device based on the language model may include 1 or N display screens 24, where N is a positive integer greater than 1. The processor 21 may include one or more graphics processing units (GPUs) that execute program instructions to generate or change display information. The courseware generation device based on the language model implements the display function through the GPU and the display screen 24. The GPU is used to perform mathematical and geometric calculations and is used for graphics rendering.

[0201] The Wi-Fi device 23 is used to provide network access that complies with Wi-Fi related standard protocols for the courseware generation device based on the language model. The courseware generation device based on the language model can access the Wi-Fi access point through the Wi-Fi device 23, thereby helping to browse the web and access streaming media, etc., and it provides users with wireless broadband Internet access. The courseware generation device based on the language model can also establish a Wi-Fi connection with a terminal device connected to the Wi-Fi access point through the Wi-Fi device 23 and the Wi-Fi access point for mutual data transmission. In some other embodiments, the Wi-Fi device 23 can also be used as a Wi-Fi wireless access point, which can provide Wi-Fi network access for other electronic devices. Exemplarily, the Wi-Fi device 23 includes at least one wireless network card.

[0202] The Bluetooth device 25 is used to implement data exchange between the courseware generation device based on the language model and other short-range electronic devices (such as terminals, smart watches, etc.) The Bluetooth device 25 in the embodiment of the present invention can be an integrated circuit or a Bluetooth chip.

[0203] The audio circuit 26, the speaker, and the microphone can provide an audio interface between the user and the courseware generation device based on the language model. The audio circuit 26 can transmit the electrical signal converted from the received audio data to the speaker, which is converted into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 26 and converted into audio data, and then the audio data is sent to the terminal via the Internet, Wi-Fi network, or Bluetooth, or the audio data is output to the memory 22 for further processing.

[0204] The power system 27 is used to charge various components of the courseware generation device based on the language model. The power system 27 may include a battery and a power management module. The battery may be logically connected to the processor 21 through a power management chip, so that the power system 27 can manage charging, discharging, and power consumption.

[0205] The peripheral interface 28 is used to provide various interfaces for external input / output devices (external display, external memory 22, user identification module card, etc.). For example, by connecting to the external memory 22 through the external memory 22 interface, such as a MicroSD card, the storage capacity of the courseware generation device based on the language model can be expanded. The peripheral interface 28 can be used to couple the above-mentioned external input / output peripheral devices to the processor 21 and the memory 22.

[0206] The sensor module 29 may include at least one sensor. For example, a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor. Among them, the ambient light sensor can adjust the brightness of the display screen 24 according to the brightness of the ambient light. As a kind of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in each direction (generally three axes), and can detect the magnitude and direction of gravity when stationary, which can be used to identify the application of the posture of the courseware generation device based on the language model (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), etc. Of course, according to actual needs, the sensor module 29 can also include any other feasible sensors.

[0207] The data conversion module 30 may include a digital-to-analog converter and an analog-to-digital converter. The relevant functions of the digital-to-analog converter and the analog-to-digital converter can be referred to the corresponding explanations in the above-mentioned related technical terms, which will not be repeated here.

[0208] An embodiment of the present invention further provides a non-volatile computer-readable storage medium, which stores computer-executable instructions. The computer-executable instructions are executed by one or more processors, for example, to execute the method steps of the language model-based courseware generation method described above.

[0209] An embodiment of the present invention also provides a computer program product, including a computer program stored on a non-volatile computer-readable storage medium, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the language model-based courseware generation method in any of the above-mentioned method embodiments, for example, executes the method steps of the language model-based courseware generation method described above.

[0210] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0211] Through the description of the above embodiments, it is clear to those skilled in the art that each embodiment can be implemented by means of software plus a general hardware platform, or by hardware. It is understood by those skilled in the art that all or part of the processes in the above embodiment methods can be completed by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, which, when executed, can include the processes of the embodiments of the above methods. Among them, the storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.

[0212] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Under the concept of the present invention, the technical features in the above embodiments or different embodiments may also be combined, the steps may be implemented in any order, and there are many other changes in different aspects of the present invention as described above, which are not provided in detail for the sake of simplicity. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A courseware generation method based on a language model, characterized in that: include: Get multiple courseware page titles based on chapter names; Acquire the courseware page content corresponding to each courseware page title through the target language model; Combining each of the courseware page titles with the corresponding courseware page content into one courseware page to obtain multiple courseware pages; Each of the courseware pages is converted into a demonstration page to obtain multiple demonstration pages; and the multiple demonstration pages are combined to obtain a courseware file.

2. The courseware generation method based on the language model according to claim 1 is characterized in that: The method of obtaining multiple courseware page titles according to the chapter name includes: Acquire the teaching link corresponding to the chapter name from a classroom teaching link library according to the chapter name; Acquiring the teaching objective corresponding to the chapter name from a chapter teaching objective library according to the chapter name; The chapter title, the teaching link and the teaching goal constitute a first input text; The first input text is learned and analyzed by the target language model to obtain a corresponding courseware page title.

3. The courseware generation method based on the language model according to claim 2 is characterized in that: The learning analysis of the first input text by the target language model to obtain the corresponding courseware page title includes: Converting the first input text into a vector sequence prefix by using the target language model; Predicting, according to the vector sequence prefix, a word vector that appears after the vector sequence prefix; Transform all predicted word vectors to obtain the courseware page title.

4. The courseware generation method based on the language model according to claim 3 is characterized in that: The converting the first input text into a vector sequence prefix by the target language model comprises: Segmenting the first input text using the target language model to obtain a plurality of subwords and position information of each of the subwords; Obtain the word vector corresponding to each subword from the word vector matrix, and obtain the position vector corresponding to each position information from the position vector matrix; The word vector and the position vector corresponding to each of the sub-words are combined together to form a vector representation of each of the sub-words to obtain multiple vector representations; and the multiple vector representations are formed into a vector sequence prefix.

5. The courseware generation method based on language model according to claim 3 is characterized in that: The predicting, according to the vector sequence prefix, a word vector that appears after the vector sequence prefix comprises: Predicting the occurrence probability of each word vector in the word vector matrix according to the vector sequence prefix, taking the word vector with the highest occurrence probability as the next word vector, and determining the position vector of the next word vector according to the position vector represented by the last vector in the vector sequence prefix; Combining the next word vector and the position vector of the next word vector to form a next vector representation, and adding the next vector representation to the vector sequence prefix to update the vector sequence prefix; The updated vector sequence prefix is ​​re-input into the target language model, and the occurrence probability of each word vector in the word vector matrix is ​​predicted again according to the updated vector sequence prefix to obtain the next vector representation, and then the vector sequence prefix is ​​updated until the end identifier is received or the predicted word vector reaches a predetermined length to obtain all word vectors appearing after the vector sequence prefix.

6. The courseware generation method based on language model according to claim 3 is characterized in that: The above-mentioned conversion of all predicted word vectors to obtain the courseware page title includes: All predicted word vectors are converted into corresponding sub-words, the order of all sub-words is determined according to the position vectors of all predicted word vectors, and all sub-words are arranged according to the order to obtain the courseware page title.

7. The courseware generation method based on language model according to claim 1 is characterized in that: The step of obtaining the courseware page content corresponding to each courseware page title through the target language model includes: Obtain the chapter textbook content, chapter supplementary teaching content, chapter exercise content and teaching suggestions corresponding to the chapter name; The chapter name, the courseware page title, the chapter teaching material content, the chapter supplementary teaching content, the chapter exercise content and the teaching suggestion constitute a second input text; The second input text is learned and analyzed by the target language model to obtain the courseware page content corresponding to the courseware page title.

8. The courseware generation method based on language model according to claim 7 is characterized in that: The learning and analyzing of the second input text by the target language model to obtain the courseware page content corresponding to the courseware page title includes: Converting the second input text into a vector sequence prefix by using the target language model; Predicting, according to the vector sequence prefix, a word vector that appears after the vector sequence prefix; Convert all predicted word vectors into courseware page content.

9. The method for generating courseware based on a language model according to any one of claims 1 to 8, characterized in that: Converting each page of the courseware into a demonstration page to obtain multiple pages of demonstration pages; Combining the multiple demonstration pages to obtain a courseware file includes: Filling the courseware page title and courseware page content of each courseware page into a predefined demonstration template to obtain a process demonstration page; By adjusting the attribute tags of the process demonstration page, the font, font size and / or color are adjusted to obtain a final demonstration page, so as to obtain multiple pages of demonstration pages; The multiple demonstration pages are combined according to the demonstration order to obtain a courseware file.

10. The method for generating courseware based on a language model according to any one of claims 1 to 8, characterized in that: The method for constructing the target language model includes: Establish pre-training data based on textbooks, teacher's books, supplementary teaching books, exercise books and existing courseware files; Train the initial language model using pre-training data; Establishing instruction fine-tuning data according to a preset input text and an expected output text corresponding to the preset input text; The trained initial language model is trained using the preset input text in the instruction fine-tuning data to obtain a predicted output text; The cross entropy loss is calculated according to the expected output text and the predicted output text in the instruction fine-tuning data, and the parameters of the initial language model are fine-tuned according to the cross entropy loss until the cross entropy loss is less than a set loss threshold, so as to obtain the target language model.

11. A courseware generation device based on a language model, characterized in that: include: at least one processor; as well as, a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the language model-based courseware generation method as described in any one of claims 1-10.

Citation Information

Cited By

  • Intelligent enterprise training courseware generation method, device and equipment based on artificial intelligence

    CN121659907A

  • AI-based intelligent generation method, device, and equipment for enterprise training courseware

    CN121659907B