Intelligent tutoring method and device based on online learning platform and storage medium
Patent Information
- Application Number
- CN202610485508.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-14
- Publication Date
- 2026-08-18
AI Technical Summary
[0005]有鉴于此,本申请实施例提供了一种基于线上学习平台的智能助教方法、装置及存储介质,以解决现有技术存在的异构学习资料利用不足、智能助教协同处理能力弱、试题生成与管理割裂的问题
通过获取由用户输入的助教配置数据以及目标业务请求,其中,助教配置数据包括角色设定信息、知识来源信息和任务编排信息;基于知识来源信息获取对应的异构学习资料,并对异构学习资料执行版面识别、内容提取和语义分块处理,得到结构化知识片段集合;基于结构化知识片段集合提取知识实体及实体关系,构建与角色设定信息关联的知识体系;根据目标业务请求和任务编排信息,从知识体系中确定目标知识片段集合,并生成与目标业务请求对应的助教处理上下文;基于助教处理上下文执行智能助教处理,生成与目标业务请求对应的目标处理结果,其中,目标处理结果包括问答结果、总结结果、创作结果、会议处理结果、日程处理结果和试题结果中的至少一种;将目标处理结果输出至线上学习平台,其中,在目标处理结果为试题结果时,将试题结果关联存入目标题库。本申请能够提高异构资料处理效率、提升智能助教协同服务能力、提高试题生成与管理效率。
Smart Images

Figure CN122598502A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to an intelligent teaching assistant method, device and storage medium based on an online learning platform. Background Technology
[0002] With the continuous application of technologies such as artificial intelligence, big data, and virtual reality in corporate training and online learning scenarios, online learning platforms are gradually evolving from traditional course management tools into comprehensive platforms integrating knowledge management, learning support, training organization, and process collaboration. Especially in the process of building learning organizations, training materials are becoming increasingly diverse, and the learning process involves various tasks such as document reading, Q&A, content creation, meeting minutes, and assessments, placing higher demands on the platform's intelligence, scenario-based features, and collaborative capabilities.
[0003] Most existing online learning platforms were deployed using standardized products or software-as-a-service models in their initial development, resulting in relatively fixed functional structures. Their main functions include course publishing, learning records, test management, and basic examinations. While some platforms have incorporated artificial intelligence capabilities, these typically remain at the level of general question answering, simple summarizing, or generating single pieces of content. They lack sufficient support for in-depth analysis of learning materials, knowledge system construction, teaching assistant role configuration, streamlined task arrangement, and intelligent question generation. Furthermore, existing platforms often lack unified processing and comprehensive utilization mechanisms when dealing with heterogeneous materials such as documents, images, and audio / video, leading to weak knowledge accumulation capabilities.
[0004] Based on this, existing technologies have at least the following problems: the platform has a low level of intelligence, making it difficult to provide integrated teaching assistant support around training and learning scenarios; it lacks the ability to uniformly analyze and organize heterogeneous learning materials, making it difficult to support high-quality question answering, summarizing, and creating; the automation level of the question generation process is insufficient, with question generation, management, and subsequent analysis being disconnected, making it difficult to meet the application needs of current online learning platforms for flexible configuration, intelligent assistance, and efficient operation. Summary of the Invention
[0005] In view of this, embodiments of this application provide an intelligent teaching assistant method, device, and storage medium based on an online learning platform to solve the problems of insufficient utilization of heterogeneous learning materials, weak collaborative processing capabilities of intelligent teaching assistants, and fragmentation of test question generation and management in the prior art.
[0006] A first aspect of this application provides an intelligent teaching assistant method based on an online learning platform, comprising: acquiring teaching assistant configuration data input by a user and a target business request, wherein the teaching assistant configuration data includes role setting information, knowledge source information, and task arrangement information; acquiring corresponding heterogeneous learning materials based on the knowledge source information, and performing layout recognition, content extraction, and semantic segmentation processing on the heterogeneous learning materials to obtain a set of structured knowledge fragments; extracting knowledge entities and entity relationships based on the set of structured knowledge fragments to construct a knowledge system associated with the role setting information; determining a target knowledge fragment set from the knowledge system according to the target business request and task arrangement information, and generating a teaching assistant processing context corresponding to the target business request; performing intelligent teaching assistant processing based on the teaching assistant processing context to generate a target processing result corresponding to the target business request, wherein the target processing result includes at least one of question-and-answer results, summary results, creation results, meeting processing results, schedule processing results, and test results; and outputting the target processing result to the online learning platform, wherein when the target processing result is a test result, the test result is associated and stored in a target question bank.
[0007] A second aspect of this application provides an intelligent teaching assistant device based on an online learning platform, comprising: an acquisition module for acquiring teaching assistant configuration data and a target business request input by a user, wherein the teaching assistant configuration data includes role setting information, knowledge source information, and task arrangement information; a processing module for acquiring corresponding heterogeneous learning materials based on the knowledge source information, and performing layout recognition, content extraction, and semantic segmentation processing on the heterogeneous learning materials to obtain a set of structured knowledge fragments; and an extraction module for extracting knowledge entities and entity relationships based on the set of structured knowledge fragments to construct a knowledge system associated with the role setting information. The determination module is used to determine the set of target knowledge fragments from the knowledge system based on the target business request and task arrangement information, and generate a teaching assistant processing context corresponding to the target business request; the generation module is used to execute intelligent teaching assistant processing based on the teaching assistant processing context, and generate the target processing result corresponding to the target business request, wherein the target processing result includes at least one of question and answer results, summary results, creation results, meeting processing results, schedule processing results, and test results; the output module is used to output the target processing result to the online learning platform, wherein when the target processing result is a test result, the test result is associated and stored in the target question bank.
[0008] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0009] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described method.
[0010] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: By acquiring user-inputted teaching assistant configuration data and target business requests, where the teaching assistant configuration data includes role setting information, knowledge source information, and task arrangement information; acquiring corresponding heterogeneous learning materials based on the knowledge source information, and performing layout recognition, content extraction, and semantic segmentation on the heterogeneous learning materials to obtain a set of structured knowledge fragments; extracting knowledge entities and entity relationships based on the set of structured knowledge fragments to construct a knowledge system associated with the role setting information; determining the target knowledge fragment set from the knowledge system according to the target business request and task arrangement information, and generating a teaching assistant processing context corresponding to the target business request; performing intelligent teaching assistant processing based on the teaching assistant processing context to generate a target processing result corresponding to the target business request, where the target processing result includes at least one of question-and-answer results, summary results, creation results, meeting processing results, schedule processing results, and test results; and outputting the target processing result to an online learning platform, wherein when the target processing result is a test result, the test result is associated and stored in the target question bank. This application can improve the efficiency of heterogeneous data processing, enhance the intelligent teaching assistant collaborative service capability, and improve the efficiency of test generation and management. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating the intelligent teaching assistant method based on an online learning platform provided in an embodiment of this application. Figure 2 This is a schematic diagram of the structure of the intelligent teaching assistant device based on an online learning platform provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0014] With the development and application of technologies such as artificial intelligence and virtual reality, online learning platforms are undergoing rapid technological changes. The rapid development of technologies such as artificial intelligence and big data has provided more possibilities and opportunities for learning and training businesses, and has provided new impetus for building learning organizations and creating a new pattern of knowledge co-creation and sharing across the entire enterprise.
[0015] In the realm of related technologies, online learning platforms initially deployed using the SaaS model as standardized products are outdated in technology, functionally inferior, and suffer from severely lagging iteration compared to current artificial intelligence products, thus failing to meet the needs of future talent development. Furthermore, current online learning platforms exhibit problems such as a simplistic operational model, low levels of intelligence, insufficient data analysis and application capabilities, and high costs and time commitments for customized development.
[0016] In view of the problems existing in the prior art, this application provides an intelligent teaching assistant method for an online learning platform. The teaching assistant of the online learning platform of this application includes AI assistant functions and AI question generation functions.
[0017] The AI assistant performs the following functions: 1) Customize the AI assistant's character style and set up the character; All organization members can create AI assistants, possessing their own powerful, personalized, and context-specific intelligent entities.
[0018] For example, we could create a think tank-based assistant called "Xiao Zhi," or an intelligent assistant that can handle different roles in training scenarios, such as lesson preparation.
[0019] 2) Establish a knowledge system for the intelligent assistant by uploading local files, knowledge bases, and other documents, giving the AI assistant a unique knowledge base. Advanced knowledge settings allow for more personalized responses, and usage data helps optimize the original knowledge text. The process of building a knowledge system: Employing a multimodal document parsing engine, combined with adaptive layout recognition algorithms and semantic segmentation technology, it performs unified structural processing on heterogeneous documents such as PDFs, Word documents, and scanned images. By extracting key entities and relationships through a pre-trained language model, it constructs knowledge graph nodes and automatically removes duplicates and links them, achieving efficient information extraction and integration across document formats to build a knowledge system.
[0020] 3) Customize and orchestrate AI assistant workflows.
[0021] By orchestrating workflows, AI assistants can perform a series of tasks in sequence. 4) Custom anthropomorphic operation RPA.
[0022] Let the AI assistant learn and imitate your usual operating procedures, and converse with the AI assistant in natural language to help you complete common operations.
[0023] 5) Intelligent communication.
[0024] It enables intelligent summarization of messages, documents, and daily learning tasks during routine management, training organization, and learning processes. This may include: It can provide intelligent summaries based on conversation, time, and speaker messages.
[0025] The study / work group can be configured for group Q&A; it can quickly read documents, files such as PDFs and docxes, and images, and intelligently generate summaries of the corresponding content; It can summarize and display the student's daily schedule, to-do items, and the number of chat messages received, so as to quickly view the day's work overview.
[0026] For example, a. Message Summary - Say Goodbye to Climbing Towers Generate multi-topic summaries with one click to quickly understand the chat background and review each chat session with the AI assistant.
[0027] b. Intelligent Q&A - Efficient and saves manpower in answering questions The intelligent assistant creates and trains a question-and-answer robot within the group, eliminating the need to repeatedly answer similar business questions. It is applicable to scenarios such as customer service and Q&A.
[0028] c. Speed Reading - Quickly gain access to a vast amount of knowledge Want to read a vast amount of knowledge but don't have time? Quick reading is here! Get a quick overview of articles and links with a single click, and get recommended questions automatically. You can also learn more about the article content through dialogue.
[0029] d. Work Overview: Generate your work summary with one click. To review your work, there's no need to frantically search through chat logs. Just ask me "What happened today?" or "What important messages did I receive?" to easily summarize your work.
[0030] 6) Intelligent creation.
[0031] It can intelligently generate various document types and transform content into multiple expression methods. It can generate images based on keywords or descriptions of specific scenarios, and create visual data tables based on instructions, listing information for comparison in tables.
[0032] For example, a. Document creation - quickly get started with multi-scenario creation. Use the AI assistant to quickly generate rich content such as marketing plans, creative stories, promotional copy, summary reports, and competitor analysis, allowing AI to comprehensively improve work efficiency.
[0033] b. Whiteboard Collaboration - Generate PPT with a single sentence Creating PPT presentations for summaries and reports takes a lot of time and effort; let an AI assistant quickly generate PPTs on a whiteboard.
[0034] c. Tabular Data Analysis - Making Data Processing Time-Saving and Effortless Let AI assistants gain insights into data and intelligently generate pivot tables and charts, enabling data calculations to be completed without writing functions.
[0035] 7) Smart Meetings.
[0036] It has the ability to transcribe speech into text in real time during meetings, and can intelligently generate summaries based on the meeting content to quickly understand the key points. It also features speech-to-text transcription, intelligent segmentation of meeting content into chapters and sections, intelligent minutes, and task recognition capabilities.
[0037] For example, a. Virtual backgrounds - say goodbye to boring meetings. Use the " / " key to describe the desired background in text, and the AI assistant will design a virtual background image based on your needs, completely unleashing your imagination.
[0038] b. Smart Minutes - Meeting Highlights at a Glance Intelligent summaries are generated based on meeting content, allowing you to quickly understand the key points of the meeting; even those who join later can see what was discussed earlier; and meeting minutes are automatically generated after the meeting to summarize the content of the meeting.
[0039] 8) Smart scheduling.
[0040] It can automatically create events based on the description of the event's theme, time, participants, and location; it can also intelligently generate calendar posters based on the event's theme.
[0041] For example, a. AI assistant creates a new schedule This to-do list management assistant allows you to create, complete, and query to-do items at any time.
[0042] b. Schedule Poster Create a smart poster with AI-enhanced content and one-click text poster generation; intelligently cut out images from photos to easily create portrait posters, making schedule dissemination vivid and effortless.
[0043] The implementation scheme for the AI-generated question function is as follows: The AI-powered question generation function is implemented using large language model technology. First, text content is extracted from audio, video, and document content. Then, the text content is segmented, and finally, prompt words are optimized to generate test questions.
[0044] AI-generated question scenarios require generating one or more types of questions based on a given text. This scenario can be broken down into multiple prompts, with each prompt requiring the large model to perform a relatively simple task. For example, given a prompt template, the large model could generate a single-choice question based on the provided text, returning four options and the correct answer. The returned results should be organized in JSON format for easier subsequent processing.
[0045] After generating various types of test questions with question-answer pairs, save them to different question banks for easy management. Additionally, provide test question editing functionality to correct errors in the results returned by the large model.
[0046] After students complete their answers, automated answer checking can improve grading efficiency, allowing employees to quickly obtain their exam results.
[0047] Furthermore, the answers can be used for subsequent analysis and processing. By analyzing the difficulty of the questions, the difficulty labels of the questions can be optimized and adjusted, which can then be used to process questions based on their difficulty level in future question creation.
[0048] During the above process, training managers / instructors upload reference materials and input question requirements. AI helps users quickly and automatically generate questions with reasonable format and difficulty distribution, which are then added to the question bank.
[0049] AI-generated questions can generate test questions quickly based on text descriptions, video content, etc., helping users create a large number of high-quality test questions and saving users time in test question creation. This can be summarized as follows: 1) Add / select the materials for the questions. You can choose directly from the material library or upload them from your local computer.
[0050] 2) Setting Question Requirements: The AI-generated question generator automatically provides a question distribution based on the provided materials, and adjusts the question type distribution accordingly. Manual adjustment of the question type distribution is also supported. Once completed, different archived question banks can be selected.
[0051] 3) AI-generated questions: The system supports generating questions that meet the format requirements, including the question stem, the answer, the answer explanation, the knowledge points, etc. It can also check, adjust and optimize the questions.
[0052] The technical solution of this application will now be described in detail with reference to the accompanying drawings and specific embodiments.
[0053] Figure 1 This is a flowchart illustrating the intelligent teaching assistant method based on an online learning platform provided in an embodiment of this application. Figure 1 As shown, the method may specifically include: S101, Obtain teaching assistant configuration data and target business request input by the user, wherein the teaching assistant configuration data includes role setting information, knowledge source information and task arrangement information; S102: Based on the knowledge source information, obtain the corresponding heterogeneous learning materials, and perform layout recognition, content extraction and semantic segmentation on the heterogeneous learning materials to obtain a set of structured knowledge fragments; S103, extract knowledge entities and entity relationships based on a set of structured knowledge fragments, and construct a knowledge system associated with role setting information; S104, Based on the target business request and task orchestration information, determine the target knowledge fragment set from the knowledge system and generate the teaching assistant processing context corresponding to the target business request; S105, based on the teaching assistant processing context, execute intelligent teaching assistant processing to generate target processing results corresponding to the target business request, wherein the target processing results include at least one of question and answer results, summary results, creation results, meeting processing results, schedule processing results and test results; S106, output the target processing result to the online learning platform. When the target processing result is a test result, the test result is associated and stored in the target question bank.
[0054] In some embodiments, obtaining teaching assistant configuration data input by the user and a target business request includes: The system receives user input regarding the role configuration, knowledge configuration, and task configuration for the target intelligent teaching assistant. The role configuration represents the target intelligent teaching assistant's scenario identity and response style, the knowledge configuration represents the scope of knowledge sources and corresponding materials, and the task configuration represents the processing flow and result type. The role configuration content, knowledge configuration content, and task configuration content are structured and assembled to generate teaching assistant configuration data corresponding to the target intelligent teaching assistant; Receive business processing instructions initiated by users based on the online learning platform, extract the task description information, input object information and result requirement information corresponding to the business processing instructions, and generate the target business request.
[0055] Specifically, acquiring user-inputted teaching assistant configuration data and target business requests mainly revolves around two stages: establishing the pre-configuration of the target intelligent teaching assistant and standardizing the generation of business requests. The former stage transforms the role attributes, knowledge boundaries, and task processes configured by users for different training scenarios into teaching assistant configuration data that the platform can recognize and call. The latter stage transforms the business processing instructions initiated by users during actual use of the online learning platform into the target business requests required for subsequent intelligent teaching assistant processing.
[0056] Through the above processing, the target intelligent teaching assistant already has a clear scene identity, available knowledge sources, processing flow constraints, and result output requirements before entering the processing stages such as question answering, summarizing, creating, meeting processing, or test question generation. This provides a unified input foundation for subsequent knowledge retrieval, context construction, and result generation.
[0057] In practical implementation, role configuration primarily defines the target intelligent teaching assistant's identity and response style within the current application scenario. For example, in a training organization scenario, the target intelligent teaching assistant can be configured as a course assistant for instructors preparing lessons, a learning assistant for answering students' questions, or an operations assistant for training administrators. Besides scenario-specific identity information, different role configurations can also include style information such as response approach, language style, level of detail in content, and whether to prioritize outputting summaries, key points, or structured results.
[0058] The knowledge configuration primarily defines the scope of knowledge sources and corresponding materials that the target intelligent teaching assistant can access. These materials can come from locally uploaded files, the platform's resource library, knowledge base documents, historical training materials, meeting minutes, and other authorized resource collections. The task configuration primarily defines the task flow and result type adopted by the target intelligent teaching assistant when processing business requests. For example, it might involve first retrieving materials, then extracting key points, and finally generating results; or first extracting knowledge points, then organizing the question stem, and finally generating answer explanations. By uniformly organizing these three types of configuration content, problems such as unclear role boundaries, ambiguous knowledge access scope, and inconsistent result formats can be avoided during subsequent intelligent teaching assistant processing.
[0059] Furthermore, when performing structured assembly of role configuration content, knowledge configuration content, and task configuration content, teaching assistant configuration data can be generated using field mapping, configuration classification, and association binding. Specifically, role configuration content can be mapped to role identifier fields and style constraint fields; knowledge configuration content can be mapped to knowledge source fields and resource object fields; and task configuration content can be mapped to process node fields and result type fields. Furthermore, an association relationship can be established between the target intelligent teaching assistant identifier and the aforementioned fields. Through this structured assembly process, discrete configuration information originally entered by users in natural language, checkbox settings, or interface forms can be uniformly transformed into teaching assistant configuration data that can be directly parsed and invoked by subsequent modules of the platform.
[0060] In a specific example, a corporate training instructor needs to build a target-oriented intelligent teaching assistant for pre-class preparation and in-class Q&A. The instructor first creates the target intelligent teaching assistant on the platform, entering information such as "Course Assistant," "for training preparation and Q&A," and "formal answering style with priority given to key points" in the role configuration. In the knowledge configuration, the instructor selects training materials, policy documents, and past training Q&A records from the platform's resource library, while also uploading local course handouts and several scanned learning materials. In the task configuration, the instructor sets that upon receiving a document summary request, it first locates the relevant materials, then extracts key points, and finally outputs a summary and recommended follow-up questions. Upon receiving a question request, it first extracts knowledge points, then generates multiple-choice and true / false questions, returning the results in a structured format.
[0061] Furthermore, after receiving the above configuration content, the platform identifies "Course Assistant" and "Training Preparation and Q&A" as scenario identity information, "Formal" and "Point-by-Point Output" as response style information, courseware, policy documents, Q&A records, lecture notes, and scanned materials as knowledge source scope and corresponding material objects, and "Material Location - Key Point Extraction - Summary Output" and "Knowledge Point Extraction - Question Generation - Structured Return" as task flow and result type information. Subsequently, the platform assembles the above information into teaching assistant configuration data according to the preset field model and binds and saves it with the teaching assistant identifier of the target intelligent teaching assistant.
[0062] After the aforementioned intelligent teaching assistant is created, users can initiate business processing commands through the online learning platform. These commands can be natural language instructions in the chat window, or function requests triggered on the materials page, group chat page, meeting page, or question-generating page. For example, an instructor might enter "Based on the newly uploaded policy materials, help me compile a pre-training introductory summary," or "Generate five multiple-choice questions with answer explanations based on the course materials." Upon receiving the command, the platform parses the command content, extracting the corresponding task description information, input object information, and result requirement information.
[0063] The task description information represents the type of business the user wants the target intelligent teaching assistant to complete, such as summarizing, answering questions, creating content, or generating questions. The input object information represents the data object, course object, conversation object, or meeting object to be processed. The result requirement information represents constraints such as output format, question type distribution, word count requirement, whether to include explanations, and whether to return a structured result. Taking the aforementioned "generating five multiple-choice questions with answer explanations" as an example, the platform can identify the generated questions as task description information, the current course materials as input object information, and the five multiple-choice questions with answer explanations as result requirement information. Based on the explanation results, it generates a target business request for subsequent knowledge fragment retrieval, teaching assistant processing context generation, and intelligent teaching assistant processing module calls.
[0064] For example, in a learning group scenario, the administrator can pre-create a group Q&A-oriented intelligent teaching assistant, setting its role as a learning Q&A assistant in the role configuration, associating it with frequently asked questions, policy manuals, and business specifications within the group in the knowledge configuration, and configuring it to receive questions from the group, first match the relevant knowledge materials, then generate a concise answer and return the associated source in the task configuration. When a student asks in the group what materials are needed for reimbursement of training expenses, the platform can recognize this question as a business processing instruction, extracting the Q&A task description information, the input object information corresponding to the reimbursement of training expenses, and the result requirements information for returning the concise answer and the source, thereby generating a target business request. Since this target business request is already coordinated with the aforementioned teaching assistant configuration data, subsequent modules can execute processing within the established role and knowledge boundaries, without calling irrelevant materials or outputting results that deviate from the group Q&A scenario.
[0065] Therefore, this embodiment establishes the configuration foundation for the target intelligent teaching assistant by receiving and structurally assembling role configuration content, knowledge configuration content, and task configuration content; and generates target business requests corresponding to specific business scenarios by parsing business processing instructions. This implementation method can improve the standardization of the target intelligent teaching assistant's pre-configuration, improve the accuracy of business request parsing, and provide a unified and stable data input foundation for subsequent knowledge invocation, context construction, and result generation, thereby improving the adaptability and processing efficiency of intelligent teaching assistant processing in online learning platforms.
[0066] In some embodiments, corresponding heterogeneous learning materials are obtained based on knowledge source information, and layout recognition, content extraction, and semantic segmentation are performed on the heterogeneous learning materials to obtain a set of structured knowledge fragments, including: Based on the knowledge source information, determine the target data object set and obtain the heterogeneous learning data corresponding to each data object in the target data object set; For each data object in heterogeneous learning materials, identify the page structure and content distribution relationship of the data object, and determine the page type information corresponding to different content areas; Based on the page layout type information, the text content, image content and page layout attribute content of each content area are extracted and uniformly transformed to generate structured content units corresponding to each data object; The structured content units are segmented according to semantic relationships to obtain multiple knowledge fragments. Source identifiers, location identifiers, and content category identifiers are associated with each knowledge fragment to generate a set of structured knowledge fragments.
[0067] Specifically, in this embodiment, heterogeneous learning materials are obtained based on knowledge source information, and layout recognition, content extraction, and semantic segmentation are performed on these materials. The core of this process is to transform learning materials with scattered sources, diverse formats, and inconsistent expressions into a set of structured knowledge fragments that can be directly processed by subsequent knowledge entity extraction, knowledge system construction, and intelligent teaching assistant invocation. This process is not simply text extraction, but rather revolves around a processing chain of "material object identification—layout structure recognition—multi-type content extraction—unified conversion—semantic segmentation—fragment identification association," in order to solve the problems of large format differences, unclear content boundaries, and difficulty in unified utilization among courseware, policy documents, scanned materials, image materials, and other learning materials in online learning platforms.
[0068] In practical implementation, the first step is to determine the target set of data objects based on the knowledge source information. This knowledge source information can be generated from the knowledge configuration content during the aforementioned teaching assistant configuration process. It can refer to courseware documents, training policy documents, historical training summaries, and frequently asked questions documents in the platform's resource library, or to lecture notes, scanned textbooks, image-based notifications, and text materials linked after audio / video transcription uploaded by the user locally. After receiving the knowledge source information, the system first determines the target set of data objects corresponding to this processing based on the data source identifier, data type identifier, and data permission identifier, and then retrieves the data content from their respective storage locations.
[0069] In some examples, for materials sourced from a resource library, the original file can be directly retrieved based on the material object identifier; for locally uploaded materials, a corresponding material object identifier can be generated and included in the target material object set after upload; for scanned documents or image materials, their original image data and associated basic metadata can be retrieved simultaneously. Through this process, learning materials scattered in different locations and managed using different storage formats can be uniformly incorporated into a single processing entry point.
[0070] Furthermore, for each data object in the heterogeneous learning materials, the layout structure and content distribution relationship of the data object are identified, and the layout type information corresponding to different content areas is determined. The layout recognition here is not just about identifying page boundaries, but about identifying the arrangement relationship of content areas such as title areas, body text areas, list areas, ideographic image areas, annotation areas, header and footer areas in the document or image, and determining the organization method of the content carried by different areas.
[0071] For example, for training courseware, the system can identify the title text area, key point list area, and image area on each page; for policy documents, it can identify the chapter title area, clause text area, and footnote area; for scanned handouts, it can identify the main text area, page number area, and areas with mixed text and images. The system can combine character density, area position, block spacing, font style features, and text and image arrangement features to determine the page layout type information corresponding to different content areas, thus providing a foundation for the differentiated extraction of different content areas.
[0072] Furthermore, after determining the page layout type information, the system performs content extraction and unified transformation on the text content, image content, and page layout attribute content in each content area based on the page layout type information, generating structured content units corresponding to each document object. Specifically, text content extraction is mainly used to obtain textual information such as body text, titles, list items, and footnotes; image content extraction is mainly used to identify illustrative information, captions, or textual expressions in scanned images carried in image areas; and page layout attribute content extraction is mainly used to retain structural attributes such as the hierarchical position, sequence, page number, area type, and relationships between adjacent areas of the content area.
[0073] The purpose of unified conversion processing is to transform the extracted results from different data objects and content areas into a unified data organization format, such as uniformly forming structured content units that include content text, content type, location, hierarchical relationship, and page order. This way, subsequent modules no longer need to adapt courseware, documents, and scanned materials separately, but can directly process structured content units in a unified format.
[0074] Based on this, the system performs block processing on structured content units according to semantic relationships, resulting in multiple knowledge fragments. Semantic block segmentation is not based on a fixed number of words, but rather on a combination of the hierarchical relationship between titles and body text, semantic continuity of paragraphs, the affiliation of list items, the correspondence between images and text descriptions, and the consistency of the context and theme to determine the block boundaries. For example, the title "Expense Reimbursement Process" and the three key steps below it in the same page of courseware can be divided into one knowledge fragment; the title of a chapter in a policy document and its corresponding clause can be divided into one knowledge fragment; the same illustration and its accompanying text description in scanned learning materials can also be divided into one knowledge fragment.
[0075] In some examples, for content that continues across pages but maintains a consistent theme, the system can merge adjacent structured content units into the same knowledge fragment based on the title sequence and content continuity between pages. After segmentation, each knowledge fragment needs to be associated with a source identifier, a location identifier, and a content category identifier. The source identifier indicates which document the knowledge fragment originates from; the location identifier indicates the page number, area, or sequential position of the knowledge fragment in the original document; and the content category identifier indicates whether the knowledge fragment belongs to a policy statement, course explanation, case study, Q&A, meeting minutes, or other type. Ultimately, each knowledge fragment, along with its associated identifiers, forms a set of structured knowledge fragments.
[0076] Taking a corporate training scenario as an example, instructors can select training materials, policy documents, and Q&A records from the platform's resource library for their target intelligent teaching assistants, and upload local course handouts and several scanned learning materials. The system first generates a set of target material objects based on the aforementioned knowledge source information, including courseware files, policy documents, Q&A record documents, handouts, and scanned image files. For training courseware, the system identifies the course title, knowledge point list, and illustration areas on each page; for policy documents, it identifies chapter titles, clause content, and footnote areas; and for scanned learning materials, it identifies areas with mixed text and images and the main description area.
[0077] Subsequently, the system extracts the text and image content from each area and converts page numbers, area order, hierarchical relationships, and other layout attributes into structured content units. For example, the title of the "Reimbursement Application Conditions" section in a policy document, along with the application scope, approval requirements, and required materials below it, will be uniformly organized into multiple structured content units with related hierarchical levels. Then, based on the semantic relationship between the title and the body text, these units will be divided into complete knowledge fragments, and the policy document will be assigned a source identifier, a corresponding page number position identifier, and a "Policy Explanation" content category identifier.
[0078] For example, a page about customer reception standards in a training courseware can have its title, three key operational points, and illustrated diagrams combined into a single knowledge fragment, marked with the courseware source, corresponding page number, and course content category. This way, when performing Q&A, summarizing, or generating questions later, the system can directly retrieve the knowledge fragment matching the business request from the structured knowledge fragment set, without having to re-parse the original material page by page.
[0079] Therefore, this embodiment generates a set of structured knowledge fragments that can be directly called upon by subsequent intelligent teaching assistant processing by identifying target data objects, recognizing page structure, extracting multiple types of content, uniformly converting, and performing block processing based on semantic association of heterogeneous learning materials. This implementation method can improve the efficiency of unified processing of heterogeneous learning materials, enhance the standardization and usability of knowledge fragment generation, and provide a stable data foundation for subsequent knowledge system construction, target knowledge fragment retrieval, and intelligent teaching assistant result generation.
[0080] In some embodiments, knowledge entities and entity relationships are extracted based on a set of structured knowledge fragments to construct a knowledge system associated with role setting information, including: Based on a set of structured knowledge fragments, the target knowledge entities contained in each knowledge fragment and the semantic relationships between the target knowledge entities are identified, and entity relationship data is generated. Based on entity relationship data, deduplication and association linking are performed on target knowledge entities with the same or related semantics to generate a set of knowledge nodes and a set of node relationships. A knowledge graph structure is constructed based on the set of knowledge nodes and the set of node relationships. The knowledge graph structure is then linked with the role setting information to generate a knowledge system corresponding to the target intelligent teaching assistant.
[0081] Specifically, in this embodiment, knowledge entities and entity relationships are extracted based on a structured knowledge fragment set to construct a knowledge system associated with role setting information. This mainly revolves around semantic understanding of knowledge fragments, entity merging and linking, and knowledge organization oriented towards the teaching assistant role. This process does not simply store the previously obtained knowledge fragments, but further extracts target knowledge entities from the knowledge fragments that can represent training business objects, rule content, course topics, operation steps, Q&A items, and schedule elements. It also identifies semantic relationships such as subordination, reference, description, constraint, correspondence, and temporal sequence between each target knowledge entity. On this basis, a set of knowledge nodes, a set of node relationships, and a knowledge graph structure are formed, enabling the content originally scattered in courseware, policy documents, scanned materials, Q&A records, etc., to be organized in a relational way for subsequent Q&A, summarization, creation, meeting processing, schedule processing, and test question generation to be uniformly called upon.
[0082] In its implementation, the system first performs semantic parsing on each knowledge fragment in the structured knowledge fragment set. Since each knowledge fragment has already been associated with a source identifier, location identifier, and content category identifier in the previous stage, the system can identify the target knowledge entities contained within it by combining the main text content, heading level, contextual fragment relationships, and content category. Target knowledge entities can be course names, policy names, business matters, role objects, approval conditions, operational steps, material names, Q&A topics, meeting topics, time nodes, etc.
[0083] Simultaneously, the system further identifies the semantic relationships between target knowledge entities. For example, in policy-related knowledge fragments, it can identify the requirement relationship between reimbursement applications and submitted materials, and the execution relationship between approval processes and approvers; in course-related knowledge fragments, it can identify the explanatory relationship between customer reception guidelines and reception procedures; and in question-and-answer-related knowledge fragments, it can identify the correspondence between frequently asked questions and standard answers. Through the above processing, the system generates entity relationship data, which may include entity identifiers, entity content, entity categories, and the type and direction of relationships between entities.
[0084] Furthermore, based on entity relationship data, the system performs deduplication and linking processing on target knowledge entities with the same or related semantics. This deduplication process is not limited to content with identical wording, but rather considers entity names, contextual constraints, document type, and relationships. For example, phrases like "expense reimbursement process" in training materials, "reimbursement approval process" in policy documents, and "how to process reimbursement" in historical Q&A, despite differences in wording, can be identified as semantically identical or highly related target knowledge entities after considering the course theme and policy scope, and merged into the same knowledge node or linked strongly.
[0085] For example, the phrase "submit attachments" in the courseware and the "list of materials to be submitted" in the policy document can be identified as semantically related entities, and a corresponding relationship can be established through association links. Deduplication avoids repeatedly returning a large amount of semantically overlapping content during subsequent knowledge retrieval; association linking organizes complementary knowledge content scattered across different materials into a coherent knowledge chain. After processing, the system generates a set of knowledge nodes and a set of node relationships. The set of knowledge nodes represents the merged target knowledge entities, and the set of node relationships represents the semantic connections between the knowledge nodes.
[0086] Based on this, the system constructs a knowledge graph structure according to the set of knowledge nodes and the set of node relationships, and associates the knowledge graph structure with role setting information to generate a knowledge system corresponding to the target intelligent teaching assistant. In this embodiment, role setting information is used to limit the perspective and organizational focus of the knowledge system. For example, when the target intelligent teaching assistant is set as a course assistant, the knowledge system can prioritize building knowledge organizational relationships around course topics, knowledge point levels, case studies, and common questions; when the target intelligent teaching assistant is set as a learning Q&A assistant, the knowledge system can prioritize building knowledge organizational relationships around business matters, rules and regulations, operating conditions, and standard answers; when the target intelligent teaching assistant is set as a training operations assistant, the knowledge system can further associate meeting minutes, schedule items, and training notification content. In other words, the same knowledge graph structure can be associated with different role setting information to form knowledge systems oriented towards different scenario identities and response styles, making subsequent calls more consistent with specific business scenarios.
[0087] The following example uses a corporate training scenario. The instructor creates a target-oriented intelligent teaching assistant for lesson preparation and Q&A, linking training courseware, policy documents, past training Q&A records, locally uploaded handouts, and scanned learning materials to the knowledge sources. The system has already generated a set of structured knowledge fragments from these materials in previous steps. At this point, the system identifies target knowledge entities such as customer reception guidelines, reception procedures, and precautions from a courseware knowledge fragment; target knowledge entities such as reimbursement application conditions, submitted materials, and approval processes from a policy document knowledge fragment; and target knowledge entities such as how training expenses are reimbursed and what attachments need to be uploaded, along with standard answers, from the Q&A record knowledge fragment.
[0088] Subsequently, the system identifies the correspondence and explanatory relationships between how training expenses are reimbursed and the reimbursement application conditions, submitted materials, and approval processes. It also identifies the subordinate relationships between customer reception standards and reception steps and precautions, forming entity relationship data. Afterward, the system performs deduplication and merging on descriptions of the reimbursement process, reimbursement approval process, and how to proceed with the reimbursement process. It establishes links between submitted attachments and the list of submitted materials, generating a unified set of knowledge nodes and node relationships.
[0089] Ultimately, the system constructs a knowledge graph structure organized around course themes, rules and regulations, and Q&A items, and links it to the role setting information of the course assistant, generating a knowledge system corresponding to the target intelligent teaching assistant. Thus, when an instructor subsequently initiates a request to compile a pre-training summary based on policy materials, or when a student requests materials for reimbursement of training expenses, the system can quickly locate relevant knowledge nodes and relationships based on this knowledge system, generating a processing result that conforms to the role setting.
[0090] Therefore, this embodiment, by performing target knowledge entity recognition, semantic relationship extraction, entity deduplication, association linking, and knowledge graph construction on a set of structured knowledge fragments, forms a knowledge system that matches the role setting information. This implementation method can improve the organization of knowledge content in heterogeneous learning materials, improve the accuracy of knowledge retrieval and knowledge retrieval, and enhance the continuous processing capabilities of the target intelligent teaching assistant in scenarios such as lesson preparation, Q&A, summarizing, and question generation.
[0091] In some embodiments, based on the target business request and task orchestration information, a set of target knowledge fragments is determined from the knowledge system, and a teaching assistant processing context corresponding to the target business request is generated, including: Based on the target business request, the target task type, target object information, and result constraint information are parsed. Based on the target task type and task orchestration information, determine the target processing flow corresponding to the target business request, and determine the knowledge retrieval conditions corresponding to each processing node in the target processing flow; Based on the knowledge retrieval conditions, candidate knowledge fragments that match the target object information and result constraint information are retrieved from the knowledge system, and the candidate knowledge fragments are filtered for relevance and arranged in order to obtain a set of target knowledge fragments; The target business request, target processing flow, and target knowledge fragment set are associated and assembled to generate a teaching assistant processing context corresponding to the target business request.
[0092] Specifically, in this embodiment, based on the target business request and task orchestration information, a set of target knowledge fragments is determined from the knowledge system, and a teaching assistant processing context corresponding to the target business request is generated. This mainly revolves around semantic parsing of the business request, determination of the processing flow, targeted invocation of knowledge fragments, and context assembly. This process occurs after the knowledge system is built and before the intelligent teaching assistant's processing. Its purpose is to transform the specific business request initiated by the user into a task context that can be directly used for subsequent question answering, summarizing, creation, meeting processing, schedule processing, or test question generation. This avoids the intelligent teaching assistant directly facing an unorganized original request and a large set of knowledge during the processing, thereby improving the relevance and coherence of subsequent processing.
[0093] In its implementation, the system first parses the target task type, target object information, and result constraint information based on the target business request. The target task type characterizes the business category the user hopes to complete in this request, such as document summarization, Q&A, content creation, meeting minutes generation, schedule organization, or test question generation. The target object information characterizes the data object, course object, group chat object, meeting object, or knowledge topic targeted by this request. The result constraint information characterizes restrictions such as output format, output granularity, content length, structural requirements, question type requirements, whether parsing is included, or whether citations are required.
[0094] For example, in some examples, when an instructor enters "Based on the newly uploaded policy materials, help me compile a pre-training introductory summary," the system can identify "compile summary" as the target task type, "newly uploaded policy materials" as the target object information, "pre-training introductory summary" as a usage constraint, and "output in summary format" as a result constraint. As another example, when an instructor enters "Generate five multiple-choice questions with answer explanations based on the course materials," the system can identify "question generation" as the target task type, "course materials" as the target object information, and "five multiple-choice questions" and "with answer explanations" as result constraints.
[0095] Furthermore, after parsing the target business request, the system determines the target processing flow corresponding to the target business request based on the target task type and task orchestration information, and determines the knowledge retrieval conditions corresponding to each processing node in the target processing flow. The task orchestration information here originates from the task configuration content when the target intelligent teaching assistant is created, which essentially limits the processing order and intermediate result organization method to be adopted in different business scenarios.
[0096] For example, in a document summary scenario, the task arrangement information can be limited to "knowledge location—key point extraction—summary organization—result output"; in a group question-and-answer scenario, it can be limited to "question identification—related knowledge matching—standard answer generation—source inclusion"; and in a test question generation scenario, it can be limited to "knowledge point extraction—candidate content screening—question organization—answer parsing generation". Based on the parsed target task type, the system matches the corresponding process template from the task arrangement information and transforms it into a target processing flow specific to this request.
[0097] Furthermore, the system determines the corresponding knowledge retrieval conditions for each processing node in the target processing flow. For example, at the knowledge location node, the target object can be limited to policy documents and lecture materials; at the key extraction node, knowledge fragments with content categories identified as policy descriptions and course explanations can be prioritized; and at the question organization node, knowledge fragments containing clear knowledge points, rule items, and question-answer pairs can be prioritized. In this way, the scope of knowledge retrieved by each processing node and the selection criteria are predetermined.
[0098] Furthermore, based on knowledge retrieval conditions, the system retrieves candidate knowledge fragments from the knowledge system that match the target object information and result constraint information. These candidate knowledge fragments are then filtered for relevance and ordered to obtain a set of target knowledge fragments. The retrieval object here is not the original data, but rather the knowledge nodes, node relationships, and corresponding knowledge fragments within the aforementioned constructed knowledge system associated with role-setting information. The system can first locate relevant data sources, knowledge topics, course subjects, or business matters based on the target object information, and then further filter knowledge fragments that meet the needs of summary usage, Q&A scenarios, question types, or meeting minutes based on the result constraint information.
[0099] For example, in a pre-training summary request, the system can prioritize retrieving knowledge fragments from newly uploaded policy materials categorized as policy descriptions, rule requirements, and process descriptions, while reducing the weight of knowledge fragments related to historical Q&As or irrelevant cases. In a request with five multiple-choice questions and answer explanations, the system can prioritize retrieving knowledge fragments from course materials that have clear knowledge point boundaries, are well-expressed, and suitable for conversion into objective questions. After retrieving candidate knowledge fragments, the system also performs relevance filtering and sequential arrangement. Relevance filtering is used to remove fragments that are less relevant to the current request, have duplicate content, or do not meet the result constraints; sequential arrangement is used to arrange knowledge fragments according to course logic, policy hierarchy, knowledge dependencies, or task flow order to ensure good contextual coherence when generating results later.
[0100] Furthermore, after obtaining the target knowledge fragment set, the system associates and assembles the target business request, the target processing flow, and the target knowledge fragment set to generate a teaching assistant processing context corresponding to the target business request. This teaching assistant processing context is not simply a collection of text, but a structured task context that includes request type, object boundaries, output constraints, process node order, and corresponding knowledge fragment reference relationships. Through this associative assembly process, the subsequent intelligent teaching assistant processing module can directly perform question answering, summarizing, creation, meeting processing, schedule processing, or test question generation based on a unified teaching assistant processing context, without needing to repeatedly parse the user's original request or perform unconstrained searches within a large knowledge system.
[0101] The following example uses a corporate training scenario. The instructor pre-creates a "course assistant" type intelligent teaching assistant and sets the following in the task scheduling information: When a document summary request is received, it first locates the relevant materials, then extracts key points, and finally outputs a summary and recommended follow-up questions; when a question request is received, it first extracts knowledge points, then generates multiple-choice and true / false questions, and returns the results in a structured format. When the instructor enters "Based on the newly uploaded policy materials, help me prepare a pre-training introductory summary" on the platform, the system first parses the target task type as summary processing, the target object information as the newly uploaded policy materials, and the result constraint information as the purpose of the pre-training introductory summary and the summary output format.
[0102] Subsequently, based on the task arrangement information corresponding to the summary processing, the system determines the target processing flow for this request, including the document location node, the key extraction node, and the summary organization node, and configures knowledge retrieval conditions for each node. For example, at the document location node, the knowledge retrieval condition is limited to retrieving knowledge fragments from newly uploaded policy materials; at the key extraction node, knowledge fragments with content categories identified as policy descriptions, process requirements, and submission conditions are prioritized for retrieval.
[0103] Next, the system retrieves candidate knowledge fragments related to reimbursement application conditions, submitted materials, and approval processes from the knowledge system, and arranges them according to the order of the policy chapters to obtain a set of target knowledge fragments. Finally, the system assembles the business request, processing flow, and the sequentially arranged set of target knowledge fragments into a teaching assistant processing context for direct use by the subsequent summary generation module.
[0104] For example, in a learning group scenario, the learning Q&A assistant created by the administrator is configured in the task orchestration information to first match knowledge materials after receiving questions in the group, then generate a concise answer and return the associated source. When a student asks what materials are needed to reimburse training expenses, the system parses the request as a Q&A processing type, identifies the reimbursement of training expenses as the target object information, and identifies the concise answer with the source as the result constraint information. Then, based on the task orchestration information, it determines the target processing flow of knowledge matching—answer generation—source attachment, and accordingly retrieves candidate knowledge fragments related to the reimbursement application conditions, the list of submitted materials, and the approval requirements from the knowledge system. After screening and sequential arrangement, a set of target knowledge fragments is formed, and further assembled into a teaching assistant processing context. Since the boundaries of the Q&A task, the source of knowledge fragments, and the output constraints are clearly defined in this context, the subsequent teaching assistant processing can more stably generate results that conform to the group Q&A scenario.
[0105] Therefore, this embodiment parses the target business request by task type, object information, and result constraints, determines the target processing flow and knowledge retrieval conditions by combining task orchestration information, and then retrieves and orchestrates target knowledge fragments from the knowledge system to ultimately generate the teaching assistant processing context. This implementation method can improve the matching accuracy between business requests and the knowledge system, enhance the context organization capabilities of subsequent intelligent teaching assistant processing, and strengthen the relevance and coherence of result generation in scenarios such as summarizing, answering questions, and generating quizzes in online learning platforms.
[0106] In some embodiments, intelligent teaching assistant processing is performed based on the teaching assistant processing context to generate a target processing result corresponding to the target business request, including: The target processing mode corresponding to the target business request is determined based on the teaching assistant processing context. The target processing mode includes at least one of the following: question and answer processing mode, summary processing mode, creation processing mode, meeting processing mode, schedule processing mode, and test question processing mode. Based on the target processing mode, the set of target knowledge fragments, task description information and result constraint information in the teaching assistant processing context are combined to generate task input data corresponding to the target processing mode. Based on the task input data, perform content generation, content refinement, content transformation, information extraction, or question generation to obtain the initial processing result corresponding to the target business request; The initial processing results are formatted according to the result constraint information to generate the target processing result.
[0107] Specifically, in this embodiment, intelligent teaching assistant processing is performed based on the teaching assistant processing context to generate a target processing result corresponding to the target business request. This mainly revolves around processing mode determination, task input construction, patterned intelligent processing, and result format standardization. This process occurs after the generation of the teaching assistant processing context, and its purpose is to further transform the teaching assistant processing context, which has already completed knowledge positioning, process constraints, and object limitations, into a target processing result that can be directly used in the business scenarios of online learning platforms. Since the preceding steps have already associated the target business request, target processing flow, and target knowledge fragment set in the teaching assistant processing context, the intelligent teaching assistant processing in this embodiment is not boundless generation, but rather a scenario-based processing process executed under a given business scenario, a given knowledge scope, and given output constraints.
[0108] In its implementation, the system first determines the target processing mode corresponding to the target business request based on the teaching assistant's processing context. The target processing mode is used to characterize the processing logic and result organization method that the current request should adopt, and it can include at least one of the following: question-and-answer processing mode, summary processing mode, creation processing mode, meeting processing mode, schedule processing mode, and test question processing mode.
[0109] Specifically, when the task description information carried in the teaching assistant's processing context represents a response to a system issue, course issue, or group question, it can be identified as a question-and-answer processing mode; when the task description information represents summarizing materials, chat logs, or meeting content, it can be identified as a summary processing mode; when the task description information represents generating promotional copy, summary reports, training materials, or creative content, it can be identified as a creation processing mode; when the task description information represents organizing meeting speech-to-text content into chapters, extracting key points, and extracting tasks, it can be identified as a meeting processing mode; when the task description information represents generating a schedule topic, time, participants, and location based on natural language description, it can be identified as a schedule processing mode; and when the task description information represents generating exam questions based on course materials or text content, it can be identified as a test question processing mode. For the same business request, the system can also combine result constraint information to determine whether to adopt a composite processing mode, such as performing summary processing first and then creation processing.
[0110] Furthermore, after determining the target processing mode, the system performs combined processing on the target knowledge fragment set, task description information, and result constraint information in the teaching assistant processing context according to the target processing mode, generating task input data corresponding to the target processing mode. This combined processing is not a simple splicing, but rather a differentiated organization of the input content based on different processing modes. For example, in the question-and-answer processing mode, the system can combine rule explanation fragments, question-and-answer pairs, and case explanation fragments from the target knowledge fragment set with the user's question content to form task input data oriented towards answer generation; in the summary processing mode, the system can combine sequentially arranged knowledge fragments with summary granularity requirements and key point scope requirements to form task input data oriented towards key point extraction.
[0111] In the creation processing mode, knowledge content from the target knowledge fragment set, response style from the role setting information, and stylistic requirements from the result constraint information can be combined. In the meeting processing mode, meeting-related knowledge fragments, speech-to-text, and minutes structure requirements can be combined. In the schedule processing mode, event descriptions from business requests can be combined with existing time and object information. In the test question processing mode, knowledge fragments with clear knowledge points, question type requirements, question quantity requirements, and result field requirements can be combined. Through the above combined processing, task input data can more accurately correspond to the intelligent teaching assistant processing needs in different scenarios.
[0112] Furthermore, based on the task input data, the system performs content generation, content refinement, content conversion, information extraction, or question generation to obtain the initial processing result corresponding to the target business request. Among them, content generation is mainly applicable to question-and-answer, content creation, schedule description completion, and some meeting minutes generation scenarios; content refinement is mainly applicable to document summaries, chat summaries, and meeting key point summaries scenarios.
[0113] Content transformation processing is mainly applicable to converting knowledge content into point-by-point explanations, summary texts, poster copy, or other target expression forms; information extraction processing is mainly applicable to extracting chapters, to-do items, participants, and time nodes from meeting content, or extracting themes, times, and locations from schedule descriptions.
[0114] The question generation process is primarily used to create multiple-choice and true / false questions based on course materials, document content, or transcribed text. The initial processing results typically contain the main content of the target business outcome, but their output structure, field formats, or display methods are not yet fully adapted to the specific business interface of the online learning platform, thus requiring further refinement.
[0115] Furthermore, after obtaining the initial processing result, the system performs formatting processing on the initial processing result according to the result constraint information to generate the target processing result. Formatting processing is mainly used to convert the initial processing result into a standard result format that meets the platform's requirements for calling, displaying, or storing. For example, in the question-and-answer processing mode, the initial answer content can be formatted as "question-answer-source"; in the summary processing mode, it can be formatted as "title-summary-recommended follow-up questions"; in the creation processing mode, it can be formatted as "manuscript title-body content-structure paragraphs"; in the meeting processing mode, it can be formatted as "chapter title-meeting focus-to-do items"; in the schedule processing mode, it can be formatted as "schedule topic-time-participants-location"; and in the test question processing mode, it can be formatted as a structured result containing fields such as question stem, options, correct answer, answer explanation, and knowledge points. Through formatting processing, the target processing result can be directly output to the corresponding chat interface, material interface, meeting interface, schedule interface, or question bank module of the online learning platform.
[0116] The following example uses a corporate training scenario. An instructor creates a course assistant-type intelligent teaching assistant, configured for use in training preparation, in-class Q&A, and test question generation. When the instructor inputs "Based on the newly uploaded policy materials, help me prepare a pre-training introductory summary," the system determines the target processing mode as a summary processing mode based on the generated teaching assistant processing context. It combines the target knowledge fragments corresponding to the policy materials, the summary task description information, and the result constraints of the pre-training introductory summary and point-by-point summarization into task input data. Subsequently, the system performs content extraction and transformation processing on the task input data, extracting key content such as the policy's scope of application, application conditions, required materials, and approval process, forming an initial processing result. Then, based on the result constraints, it is organized into a target processing result including an introductory title, key point summary, and recommended follow-up questions, which the instructor can directly view on the lesson preparation page.
[0117] For example, in a learning group scenario, the administrator pre-configures a learning Q&A assistant to automatically answer frequently asked questions within the group. When a student asks, "What materials are needed for reimbursement of training expenses?", the system determines the target processing mode as a question-and-answer processing mode. It combines target knowledge segments related to reimbursement application conditions, a list of submitted materials, and approval requirements from the knowledge system with the student's question and a concise answer with citations as result constraints to form the task input data. Subsequently, the system performs content generation processing, obtaining an initial processing result containing instructions on submitted materials and citations of relevant regulations, which is then reorganized into a question-and-answer result suitable for display within the group.
[0118] For example, in a question-generating scenario, when the instructor inputs "Generate five multiple-choice questions with answer explanations based on the course materials," the system determines the target processing mode as the question processing mode. It then combines course knowledge segments, the question generation task description, and the result constraints of the five multiple-choice questions with answer explanations to generate task input data. Based on this task input data, the system performs question generation processing, forming an initial processing result containing the question stem, options, correct answer, and explanation. This result is then further standardized according to the question bank requirements to generate the target processing result that can be directly written into the question bank.
[0119] Therefore, this embodiment determines the target processing mode based on the teaching assistant processing context, and combines the target knowledge fragment set, task description information, and result constraint information according to the mode. Then, it performs corresponding content generation, refinement, transformation, extraction, or question generation processes to ultimately obtain the target processing result that meets the platform's usage requirements. This implementation method can improve the matching degree between intelligent teaching assistant processing and specific business scenarios, improve the standardization and usability of result generation, and enhance the intelligent service capabilities of the online learning platform in scenarios such as Q&A, summarizing, creating, meeting processing, schedule processing, and test question generation.
[0120] In some embodiments, when the target processing mode is a question processing mode, question generation processing is performed based on the task input data to obtain an initial processing result corresponding to the target business request, including: Based on the target learning content corresponding to the task input data, the text content is extracted and segmented into segments to obtain multiple candidate question segments. Based on the result constraint information, determine the requirements for question type, number of questions, and result field, and generate corresponding question prompt templates based on the requirements for question type, number of questions, and result field; Based on the question prompt template, the candidate question segments are processed to generate questions, and candidate question data corresponding to each candidate question segment is obtained. The candidate question data includes at least one of the following: question stem content, answer content, explanation content, and knowledge point content. The candidate test question data is structured and organized to generate initial processing results.
[0121] Specifically, in this embodiment, when the target processing mode is the question processing mode, question generation processing is performed based on the task input data. This mainly revolves around the extraction of target learning content, the formation of candidate question segments, the generation of question prompt templates, the generation of candidate questions, and the structured organization of results. This process does not involve directly inputting the entire course material into the intelligent model to generate questions all at once. Instead, it first extracts and segments the target learning content corresponding to the question request, then constructs a question prompt template matching the current question generation task based on the question type, number of questions, and output field requirements. Subsequently, candidate question data is generated for different candidate question segments, and finally, multiple candidate question data are organized into an initial processing result that can be checked, edited, and stored in the database. Through this processing chain, the question generation process can be made consistent with the course material content, question requirements, and platform question bank management requirements.
[0122] In its implementation, the system first extracts text content based on the target learning content corresponding to the task input data, and then performs segmentation processing on the text content to obtain multiple candidate question segments. The target learning content can come from courseware, training materials, policy documents, Q&A materials, or text content transcribed from audio and video materials. For example, training managers or instructors can select course materials from the resource library or upload local training materials and input question requirements; the system then extracts text content corresponding to the course theme from the selected materials.
[0123] In some examples, for courseware materials, page titles, key knowledge points, and explanatory text can be extracted; for document materials, chapter titles, clause content, and case studies can be extracted; for video materials, continuous text can be generated based on the audio-visual transcription results, and then the explanatory content related to the course theme can be extracted. After the text content is extracted, the system does not mechanically segment it according to a fixed length, but rather segments it according to title boundaries, knowledge point boundaries, paragraph semantic continuity, and question-and-answer completeness, so that each candidate question segment corresponds to a relatively complete knowledge point or rule explanation as much as possible. For example, reimbursement application conditions, submission material lists, and approval processes can be generated into different candidate question segments so that the subsequently generated test questions are more focused and clear on the examination points.
[0124] Furthermore, the system determines the requirements for question type, number of questions, and result fields based on the result constraint information, and generates corresponding question prompt templates based on these requirements. The result constraint information here originates from the parsed results of the aforementioned target business request. For example, when the instructor inputs "Generate five multiple-choice questions with answer explanations based on the course materials," the system can identify the five multiple-choice questions as the question quantity and question type requirements, and the answer explanations as the result field requirements. If the user further requests four options to be returned in a structured format, the system can also include the four options and the structured output in the result field requirements.
[0125] The purpose of question prompt templates is to break down complex question-generating scenarios into single, easily executable task constraints for the model, enabling it to generate question content that meets platform requirements based on specified fragments. For example, for multiple-choice questions, the prompt template can be limited to generating a single-choice question based on provided text content, with the returned result including the question stem, four options, the correct answer, and answer explanation. For true / false questions, the prompt template can be limited to generating only the true / false question stem, the standard answer, and explanation. In this way, different question-generating requests can correspond to different templates, rather than using a uniform and coarse-grained question-generating instruction.
[0126] Furthermore, after generating the question prompt template, the system performs question generation processing on each candidate question segment based on the template, obtaining candidate question data corresponding to each segment. This question generation process does not generate questions all at once from the entire batch of course materials, but rather generates questions segment by segment and knowledge point by knowledge point, thereby improving the correspondence between each question and the original knowledge point. Candidate question data may include at least one of the following: question stem, answer, explanation, and knowledge point. In a preferred implementation, it may also include option content, question type identifier, and source segment identifier.
[0127] Taking the aforementioned example of "generating five multiple-choice questions with answer explanations based on the course materials," the system can generate multiple-choice questions examining the prerequisites for reimbursement applications based on candidate question segments related to application conditions; generate multiple-choice questions examining the completeness of materials based on candidate question segments related to the list of submitted materials; and generate multiple-choice questions examining the sequence of procedures based on candidate question segments related to the approval process. Since each candidate question corresponds to a specific candidate question segment, if a discrepancy is found in the wording of a question later, it can be traced back to the original knowledge segment for inspection and adjustment.
[0128] Subsequently, the system performs structured organization on the candidate test question data, generating initial processing results. The purpose of structured organization is to unify and standardize the separately generated candidate test question data into a data set that the platform can continue to process, facilitating subsequent question checking, manual editing, question bank storage, and automatic grading configuration. Specifically, candidate test question data can be merged according to question type identifier, knowledge point identifier, question order, and source fragment order to form an initial processing result containing multiple test question entries. Each test question entry can be uniformly configured with question stem, option, answer, explanation, and knowledge point fields, and can further retain the association information with the original course materials. In this way, in subsequent platform processes, training administrators can either directly write the initial processing results into the corresponding question bank, or enter the test question editing interface to revise individual test questions to correct inaccuracies in the model generation.
[0129] In a real-world example, the instructor selects courseware, policy documents, and video transcription content as the target learning material on the platform. They then input a business processing instruction to generate five multiple-choice questions, return four options, the correct answer, and answer explanations, with the results returned in a structured format. The system first extracts text content corresponding to the course theme from the courseware, policy documents, and video transcription text, then segments it into several candidate question segments based on knowledge point boundaries. Subsequently, the system generates corresponding question prompt templates based on requirements such as multiple-choice questions, five questions, four options, answer explanations, and structured returns, and generates candidate question data for each candidate question segment. After generation, the system organizes all candidate question data into a unified initial processing result, allowing the instructor to preview, modify, and then archive it to a designated question bank.
[0130] Therefore, this embodiment extracts and segments the target learning content into text, generates question prompt templates based on question type requirements, question quantity requirements, and result field requirements, and then generates candidate question data based on each candidate question segment and performs structured organization, forming an initial processing result that is connected to the platform's question bank management process. This implementation method can improve the correspondence between generated questions and course knowledge points, improve the standardization of question output, and improve the efficiency of question generation, checking, and database processing in the online learning platform.
[0131] In some embodiments, outputting the target processing result to an online learning platform includes: Based on the result type corresponding to the target processing result, determine the target display method or target storage method corresponding to the result type; Perform result encapsulation processing on the target processing result to generate result data corresponding to the result type. The result data includes result content, result identifier, and associated business identifier. The results data are sent to the corresponding business interface or business module of the online learning platform for display, access or subsequent processing. When the target processing result is a test question result, the test question result is associated with the target question bank identifier and written to the corresponding question bank storage area.
[0132] Specifically, in this embodiment, the target processing results are output to the online learning platform, mainly focusing on result type determination, result encapsulation, result delivery, and question bank storage. This process occurs after the intelligent teaching assistant processing, and its role is to further convert the previously generated question-and-answer results, summary results, creation results, meeting processing results, schedule processing results, or test results into result data that can be directly received and used by various business interfaces or business modules of the online learning platform. This avoids the mixing of different types of results in terms of display format, calling method, and storage location, ensuring the standardization of subsequent display, retrieval, calling, and management processing on the platform side.
[0133] In practical implementation, the system first determines the target display method or target storage method corresponding to the result type of the target processing result. The result type can be determined based on the aforementioned target processing mode and formatted processing results. For example, when the target processing result is a question-and-answer result, the system can map it to a chat session display method or a group question-and-answer reply method; when the target processing result is a summary result, it can map it to a document summary display method, a message summary display method, or a work overview display method; when the target processing result is a creation result, it can map it to a document editing interface, a whiteboard generation interface, or a content export interface; when the target processing result is a meeting processing result, it can map it to a meeting minutes interface or a to-do list extraction interface; when the target processing result is a schedule processing result, it can map it to a schedule management interface or a poster generation module; and when the target processing result is a test result, it can map it to a test preview interface, a question editing interface, and a question bank storage method. Through this correspondence between result types and display / storage methods, the platform can automatically select the appropriate hosting location for different business results.
[0134] Furthermore, the system performs result encapsulation processing on the target processing result, generating result data corresponding to the result type. This result data may include result content, result identifier, and associated business identifier. Result content carries the actual output response text, summary content, document content, meeting minutes content, schedule information, or test question entries; the result identifier uniquely identifies the output result, facilitating subsequent retrieval, updating, and tracking by the platform; and the associated business identifier indicates the business source corresponding to the result, such as a corresponding chat session identifier, document object identifier, meeting identifier, schedule identifier, or question bank identifier.
[0135] In a preferred implementation, the result data may further include a result type identifier, generation time information, target intelligent teaching assistant identifier, and source object reference information. For example, for summary results, the system can encapsulate the summary text, summary result identifier, and associated business identifier corresponding to the target document or target group chat; for test results, the system can encapsulate the question stem, options, answers, explanations, and other result content, and associate them with the current question request identifier and the target question bank identifier. Through result encapsulation processing, different types of results can form a unified field structure, facilitating subsequent calls by various modules of the platform according to standard interfaces.
[0136] Furthermore, after completing the result encapsulation, the system sends the result data to the corresponding business interface or business module of the online learning platform for display, retrieval, or subsequent processing. This sending is not limited to a single display action; it can enter different business chains based on different result types. For example, when a user requests "Help me compile a pre-training introductory summary based on the newly uploaded policy materials," the system-generated summary result can be encapsulated into result data containing the summary text, result identifier, and document object identifier, and sent to the material reading interface or lesson preparation interface for display. When a student asks a question in a learning group, the system-generated Q&A result can be encapsulated and sent to the group chat interface as an automatic reply message. When the system generates minutes based on meeting content, the corresponding result data can be sent to the meeting minutes module for users to further view and edit. When the system automatically creates a schedule based on natural language descriptions, the corresponding result data can be sent to the schedule management module for subsequent reminders and poster generation.
[0137] When the target processing result is a test question, the system also needs to associate the test question result with the target question bank identifier and write it to the corresponding question bank storage area. The target question bank identifier here can come from the user's selection of the archived question bank in the aforementioned question generation business request, or it can come from the platform's preset binding relationship between the question bank corresponding to the course type, training project, or knowledge topic.
[0138] Before writing the test results to the question bank storage area, the system can first establish an association between each test question entry and the target question bank identifier, the question creation task identifier, and the source material identifier, and then write it to the corresponding question bank storage area. This allows training administrators to manage test questions according to the question bank dimension and also to trace back the course materials or question creation tasks corresponding to a specific test question. In a preferred implementation, after the test results are written to the question bank storage area, a question editing interface can be accessed for manual review and revision, adapting to application scenarios where errors occur in the results returned by the large model.
[0139] Referring to the previous example, after the instructor initiates a business request on the platform to "generate five multiple-choice questions with answer explanations based on the course materials," the system generates the test results and determines the target storage method as either question bank writing or question editing and display based on the result type. Subsequently, the system encapsulates the test results, forming result data containing the question stem, answer content, explanation content, result identifier, and a business identifier associated with the question generation task, and sends it to the test preview interface for the instructor to view. After the instructor confirms that the test results have been archived to the designated question bank, the system associates the test results with the target question bank identifier and writes them to the corresponding question bank storage area for subsequent use by the exam, practice, and test management modules.
[0140] Therefore, this embodiment achieves effective connection between the intelligent teaching assistant results and the online learning platform's business modules by determining the target display method or target storage method based on the result type, and performing result encapsulation, interface delivery, and question bank association storage processing on the target processing results. This implementation method can improve the display and management efficiency of different types of processing results on the platform side, improve the standardization of result retrieval and subsequent processing, and enhance the consistency between test results and question bank management processes.
[0141] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0142] Figure 2 This is a schematic diagram of the structure of the intelligent teaching assistant device based on an online learning platform provided in an embodiment of this application. Figure 2 As shown, the device includes: The acquisition module 201 is used to acquire teaching assistant configuration data and target business requests input by the user. The teaching assistant configuration data includes role setting information, knowledge source information and task arrangement information. The processing module 202 is used to obtain corresponding heterogeneous learning materials based on knowledge source information, and to perform layout recognition, content extraction and semantic segmentation on the heterogeneous learning materials to obtain a set of structured knowledge fragments. Extraction module 203 is used to extract knowledge entities and entity relationships based on a set of structured knowledge fragments, and to construct a knowledge system associated with role setting information; The determination module 204 is used to determine the set of target knowledge fragments from the knowledge system based on the target business request and task orchestration information, and generate a teaching assistant processing context corresponding to the target business request; The generation module 205 is used to perform intelligent teaching assistant processing based on the teaching assistant processing context and generate target processing results corresponding to the target business request. The target processing results include at least one of the following: question and answer results, summary results, creation results, meeting processing results, schedule processing results, and test results. The output module 206 is used to output the target processing result to the online learning platform. When the target processing result is a test result, the test result is associated and stored in the target question bank.
[0143] In some embodiments, Figure 2 The acquisition module 201 receives the role configuration content, knowledge configuration content, and task configuration content input by the user for the target intelligent teaching assistant. Among them, the role configuration content is used to represent the scenario identity and response style of the target intelligent teaching assistant, the knowledge configuration content is used to represent the scope of knowledge sources and corresponding material objects, and the task configuration content is used to represent the processing flow and result type. The module performs structured assembly on the role configuration content, knowledge configuration content, and task configuration content to generate teaching assistant configuration data corresponding to the target intelligent teaching assistant. The module also receives the business processing instructions initiated by the user based on the online learning platform, and extracts the task description information, input object information, and result requirement information corresponding to the business processing instructions to generate the target business request.
[0144] In some embodiments, Figure 2 The processing module 202 determines the target data object set based on knowledge source information and obtains the heterogeneous learning data corresponding to each data object in the target data object set; for each data object in the heterogeneous learning data, it identifies the page structure and content distribution relationship of the data object and determines the page type information corresponding to different content areas; based on the page type information, it performs content extraction and unified transformation on the text content, image content and page attribute content in each content area to generate structured content units corresponding to each data object; it performs block processing on the structured content units according to semantic association relationship to obtain multiple knowledge fragments, and associates source identifiers, location identifiers and content category identifiers with each knowledge fragment to generate a set of structured knowledge fragments.
[0145] In some embodiments, Figure 2The extraction module 203, based on a set of structured knowledge fragments, identifies the target knowledge entities contained in each knowledge fragment and the semantic relationships between the target knowledge entities, generating entity relationship data; based on the entity relationship data, it performs deduplication and association linking on target knowledge entities with the same or related semantics, generating a set of knowledge nodes and a set of node relationships; it constructs a knowledge graph structure based on the set of knowledge nodes and the set of node relationships, and associates the knowledge graph structure with the role setting information to generate a knowledge system corresponding to the target intelligent teaching assistant.
[0146] In some embodiments, Figure 2 The determination module 204 parses the target task type, target object information, and result constraint information based on the target business request; according to the target task type and task arrangement information, it determines the target processing flow corresponding to the target business request, and determines the knowledge call conditions corresponding to each processing node in the target processing flow; based on the knowledge call conditions, it retrieves candidate knowledge fragments that match the target object information and result constraint information from the knowledge system, and performs relevance screening and sequential arrangement on the candidate knowledge fragments to obtain a set of target knowledge fragments; it associates and assembles the target business request, target processing flow, and set of target knowledge fragments to generate a teaching assistant processing context corresponding to the target business request.
[0147] In some embodiments, Figure 2 The generation module 205 determines the target processing mode corresponding to the target business request based on the teaching assistant processing context. The target processing mode includes at least one of the following: question-and-answer processing mode, summary processing mode, creation processing mode, meeting processing mode, schedule processing mode, and test question processing mode. According to the target processing mode, the module performs combined processing on the target knowledge fragment set, task description information, and result constraint information in the teaching assistant processing context to generate task input data corresponding to the target processing mode. Based on the task input data, the module performs content generation, content refinement, content conversion, information extraction, or question generation processing to obtain the initial processing result corresponding to the target business request. The module performs formatting processing on the initial processing result according to the result constraint information to generate the target processing result.
[0148] In some embodiments, Figure 2The generation module 205 extracts text content based on the target learning content corresponding to the task input data and performs segmentation processing on the text content to obtain multiple candidate question segments; it determines the requirements for question type, number of questions, and result fields based on the result constraint information, and generates corresponding question prompt templates based on the requirements for question type, number of questions, and result fields; based on the question prompt templates, it performs question generation processing on the candidate question segments to obtain candidate question data corresponding to each candidate question segment, wherein the candidate question data includes at least one of the following: question stem content, answer content, analysis content, and knowledge point content; it performs structured organization on the candidate question data to generate initial processing results.
[0149] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0150] Figure 3 This is a schematic diagram of the electronic device 3 provided in an embodiment of this application. Figure 3 As shown, the electronic device 3 of this embodiment includes: a processor 301, a memory 302, and a computer program 303 stored in the memory 302 and executable on the processor 301. When the processor 301 executes the computer program 303, it implements the steps in the various method embodiments described above. Alternatively, when the processor 301 executes the computer program 303, it implements the functions of each module / unit in the various device embodiments described above.
[0151] Electronic device 3 can be a desktop computer, laptop, handheld computer, cloud server, or other electronic device. Electronic device 3 may include, but is not limited to, processor 301 and memory 302. Those skilled in the art will understand that... Figure 3 This is merely an example of electronic device 3 and does not constitute a limitation on electronic device 3. It may include more or fewer components than shown, or different components.
[0152] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0153] The memory 302 can be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 302 can also be an external storage device of the electronic device 3, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the electronic device 3. The memory 302 can also include both internal and external storage units of the electronic device 3. The memory 302 is used to store computer programs and other programs and data required by the electronic device.
[0154] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0155] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a readable storage medium (e.g., a computer-readable storage medium). Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program may include computer program code, which may be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable storage medium may include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0156] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A method for providing intelligent teaching assistance based on an online learning platform, characterized in that, include: Obtain teaching assistant configuration data input by the user and target business requests, wherein the teaching assistant configuration data includes role setting information, knowledge source information and task arrangement information; Based on the knowledge source information, corresponding heterogeneous learning materials are obtained, and layout recognition, content extraction and semantic segmentation are performed on the heterogeneous learning materials to obtain a set of structured knowledge fragments. Based on the structured knowledge fragment set, knowledge entities and entity relationships are extracted to construct a knowledge system associated with the role setting information; Based on the target business request and the task orchestration information, a set of target knowledge fragments is determined from the knowledge system, and a teaching assistant processing context corresponding to the target business request is generated; Intelligent teaching assistant processing is performed based on the teaching assistant processing context to generate a target processing result corresponding to the target business request. The target processing result includes at least one of the following: question and answer result, summary result, creation result, meeting processing result, schedule processing result, and test result. The target processing result is output to the online learning platform, wherein when the target processing result is a test result, the test result is associated and stored in the target question bank.
2. The method according to claim 1, characterized in that, The process of obtaining the teaching assistant configuration data input by the user and the target business request includes: The system receives user input of role configuration content, knowledge configuration content, and task configuration content for a target intelligent teaching assistant. The role configuration content is used to represent the target intelligent teaching assistant's scenario identity and response style. The knowledge configuration content is used to represent the scope of knowledge sources and corresponding material objects. The task configuration content is used to represent the processing flow and result type. The role configuration content, the knowledge configuration content, and the task configuration content are structured and assembled to generate teaching assistant configuration data corresponding to the target intelligent teaching assistant; The system receives a business processing instruction initiated by a user based on the online learning platform, extracts the task description information, input object information, and result requirement information corresponding to the business processing instruction, and generates the target business request.
3. The method according to claim 1, characterized in that, The process involves acquiring corresponding heterogeneous learning materials based on the knowledge source information, and performing layout recognition, content extraction, and semantic segmentation on the heterogeneous learning materials to obtain a set of structured knowledge fragments, including: Based on the knowledge source information, a target data object set is determined, and heterogeneous learning data corresponding to each data object in the target data object set is obtained; For each data object in the heterogeneous learning materials, identify the layout structure and content distribution relationship of the data object, and determine the layout type information corresponding to different content areas; Based on the page layout type information, the text content, image content, and page layout attribute content in each content area are extracted and uniformly transformed to generate structured content units corresponding to each data object. The structured content units are segmented according to semantic relationships to obtain multiple knowledge fragments. Source identifiers, location identifiers, and content category identifiers are associated with each knowledge fragment to generate the set of structured knowledge fragments.
4. The method according to claim 1, characterized in that, The step of extracting knowledge entities and entity relationships based on the structured knowledge fragment set and constructing a knowledge system associated with the role setting information includes: Based on the structured knowledge fragment set, the target knowledge entities contained in each knowledge fragment and the semantic relationships between the target knowledge entities are identified, and entity relationship data is generated. Based on the entity relationship data, deduplication and association linking are performed on target knowledge entities with the same or related semantics to generate a set of knowledge nodes and a set of node relationships. A knowledge graph structure is constructed based on the set of knowledge nodes and the set of node relationships. The knowledge graph structure is then associated with the role setting information to generate a knowledge system corresponding to the target intelligent teaching assistant.
5. The method according to claim 1, characterized in that, The step of determining a set of target knowledge fragments from the knowledge system based on the target business request and the task orchestration information, and generating a teaching assistant processing context corresponding to the target business request, includes: Based on the target business request, the target task type, target object information, and result constraint information are parsed. Based on the target task type and the task orchestration information, determine the target processing flow corresponding to the target business request, and determine the knowledge retrieval conditions corresponding to each processing node in the target processing flow; Based on the knowledge retrieval conditions, candidate knowledge fragments that match the target object information and the result constraint information are retrieved from the knowledge system, and the candidate knowledge fragments are subjected to relevance screening and sequential arrangement to obtain the target knowledge fragment set; The target business request, the target processing flow, and the target knowledge fragment set are associated and assembled to generate a teaching assistant processing context corresponding to the target business request.
6. The method according to claim 1, characterized in that, The step of performing intelligent teaching assistant processing based on the teaching assistant processing context to generate a target processing result corresponding to the target business request includes: Based on the teaching assistant processing context, a target processing mode corresponding to the target business request is determined, wherein the target processing mode includes at least one of the following: question and answer processing mode, summary processing mode, creation processing mode, meeting processing mode, schedule processing mode, and test question processing mode. According to the target processing mode, the target knowledge fragment set, task description information and result constraint information in the teaching assistant processing context are combined and processed to generate task input data corresponding to the target processing mode. Based on the task input data, perform content generation, content refinement, content conversion, information extraction, or question generation to obtain the initial processing result corresponding to the target business request; The initial processing result is formatted according to the result constraint information to generate the target processing result.
7. The method according to claim 6, characterized in that, When the target processing mode is the question processing mode, question generation processing is performed based on the task input data to obtain an initial processing result corresponding to the target business request, including: Based on the target learning content corresponding to the task input data, extract the text content and perform segmentation processing on the text content to obtain multiple candidate question segments; Based on the result constraint information, the requirements for question type, number of questions, and result field are determined, and a corresponding question prompt template is generated based on the requirements for question type, number of questions, and result field. Based on the question prompt template, the candidate question segments are processed to generate questions, thereby obtaining candidate question data corresponding to each candidate question segment. The candidate question data includes at least one of the following: question stem content, answer content, explanation content, and knowledge point content. The candidate test question data is structured and organized to generate the initial processing result.
8. An intelligent teaching assistant device based on an online learning platform, characterized in that, include: The acquisition module is used to acquire teaching assistant configuration data and target business requests input by the user, wherein the teaching assistant configuration data includes role setting information, knowledge source information and task arrangement information; The processing module is used to obtain corresponding heterogeneous learning materials based on the knowledge source information, and to perform layout recognition, content extraction and semantic segmentation on the heterogeneous learning materials to obtain a set of structured knowledge fragments. The extraction module is used to extract knowledge entities and entity relationships based on the structured knowledge fragment set, and to construct a knowledge system associated with the role setting information; The determination module is used to determine a set of target knowledge fragments from the knowledge system based on the target business request and the task orchestration information, and generate a teaching assistant processing context corresponding to the target business request; The generation module is used to perform intelligent teaching assistant processing based on the teaching assistant processing context to generate a target processing result corresponding to the target business request, wherein the target processing result includes at least one of question and answer results, summary results, creation results, meeting processing results, schedule processing results, and test results; The output module is used to output the target processing result to the online learning platform, wherein when the target processing result is a test result, the test result is associated and stored in the target question bank.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.