Theoretical-practical integrated teaching agent based on high-quality data set and construction method thereof
By constructing an integrated theory-practice teaching intelligent body based on high-quality datasets, the problems of lagging updates and insufficient flexibility of teaching content in existing teaching systems have been solved. This has enabled the accurate generation and response of personalized teaching content, thereby improving teaching efficiency and quality.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-31
- Publication Date
- 2026-03-10
AI Technical Summary
Existing teaching systems struggle to achieve fine-grained knowledge annotation and semantic content retrieval, making it difficult to accurately locate and respond to learners' personalized needs, resulting in lagging updates and insufficient flexibility in teaching content.
We construct an integrated theory-practice teaching intelligent system based on high-quality datasets, including video indexing units, course units, interactive parsing units, content filtering units, and generation units. Through video segment information storage, knowledge point information management, interactive data parsing, and content filtering, we generate personalized teaching videos and add supplementary audio.
It enables fine-grained organization and improved searchability of teaching content, accurately identifies learners' needs, responds promptly to personalized learning requirements, compensates for deficiencies in theoretical explanations, and improves teaching efficiency and quality.
Smart Images

Figure CN121636749A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information technology, in particular to a theory-practice integrated teaching intelligent agent based on a high-quality data set and a construction method thereof. BACKGROUND
[0002] In current education practice, theoretical teaching and practical operation teaching are usually taught by using pre-recorded complete course videos or standardized teaching materials, which leads to the solidification of teaching content presentation form, the lag of updating, the lack of flexibility and dynamic adaptability. These teaching resources are usually organized in chapters or modules, covering a wide range, but the knowledge granularity is relatively coarse, and cannot be refined to specific knowledge points or skill links. When learners have specific questions in the learning process, the existing teaching system is difficult to accurately locate the content fragments highly related to the problem and respond immediately. Although artificial intelligence technology has made significant progress in natural language processing, computer vision and intelligent recommendation, it provides a technical basis for building an intelligent and personalized teaching system. However, how to effectively integrate massive operation video resources, realize fine-grained knowledge annotation, semantic content retrieval, and automatically generate teaching content with both theoretical depth and practical guidance based on learner needs is still a technical problem to be solved. Therefore, there is an urgent need for a theory-practice integrated teaching intelligent agent based on a high-quality data set and a construction method thereof. SUMMARY
[0003] The embodiments of the present specification describe a theory-practice integrated teaching intelligent agent based on a high-quality data set and a construction method thereof.
[0004] In a first aspect, the embodiments of the present specification provide a theory-practice integrated teaching intelligent agent based on a high-quality data set, comprising a video indexing unit, a course unit, an interaction analysis unit, a content filtering unit and a generation unit,
[0005] The video indexing unit stores a plurality of video segment information, the video segment information comprising a video segment address, a knowledge label, basic information, a content description and a predicted effect description;
[0006] The course unit stores or receives course data, the course data comprising a plurality of knowledge point information, the knowledge point information comprising knowledge content and associated practical operation description;
[0007] The interaction analysis unit interacts with the user and analyzes the interaction data generated by the interaction, and generates a to-be-taught knowledge point according to the interaction data and the course data;
[0008] The content filtering unit retrieves the video segment information stored in the video indexing unit according to the to-be-taught knowledge point and filters to obtain a to-be-fused video segment set;
[0009] The generating unit generates a teaching video according to the set of video clips to be fused, generates supplementary audio according to the teaching video and the knowledge point to be taught, and displays the teaching video after the supplementary audio is superimposed on the teaching video to a user.
[0010] In a second aspect, the embodiments of the present specification provide a construction method of a theory-practice integrated teaching intelligent agent based on a high-quality data set, including the following steps:
[0011] Receiving and storing a plurality of video clip information, the video clip information including video clip address, knowledge label, basic information, content description and estimated effect description, constructing a video index unit according to the video clip information;
[0012] Receiving and storing course data, the course data including a plurality of knowledge point information, the knowledge point information including knowledge content and associated practical operation description, constructing a course unit according to the course data;
[0013] Constructing a first intelligent agent unit, the first intelligent agent unit performing interaction with a user and analyzing interaction data generated by the interaction, generating a knowledge point to be taught according to the interaction data and the course data, constructing an interaction analysis unit according to the first intelligent agent unit;
[0014] Constructing a retrieval function unit, the retrieval function unit performing retrieval of video clip information stored in the video index unit according to the knowledge point to be taught and performing screening, obtaining a set of video clips to be fused, constructing a content screening unit according to the retrieval function unit;
[0015] Constructing a second intelligent agent unit, the second intelligent agent unit performing generation of a teaching video according to the set of video clips to be fused, generating supplementary audio according to the teaching video and the knowledge point to be taught, displaying the teaching video after the supplementary audio is superimposed on the teaching video to a user, constructing a generating unit according to the second intelligent agent unit.
[0016] In a third aspect, the embodiments of the present specification provide an electronic device, including a processor and a memory;
[0017] The processor is connected with the memory;
[0018] The memory is used for storing executable program code;
[0019] The processor runs a program corresponding to the executable program code by reading the executable program code stored in the memory, so as to execute the method of any one of the above aspects.
[0020] In a fourth aspect, the embodiments of the present specification provide a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the method of any one of the above aspects.
[0021] In a fifth aspect, the embodiments of the present specification provide a computer program product, which includes a computer program. The computer program is executed by a processor to implement the method of any one of the above aspects.
[0022] The technical solutions provided by some embodiments of the present specification have at least the following beneficial effects:
[0023] In the embodiments of the present specification, the intelligent agent for theory-practice integrated teaching based on high-quality data set and the construction method thereof are provided. By establishing a multi-dimensional video index unit including knowledge labels, content descriptions, basic information and estimated effect descriptions, and combining with the structured knowledge point information in the course unit, the fine-grained organization of teaching resources is realized, and the retrievability and reusability of the theory-practice integrated teaching content are improved. The interactive analysis unit analyzes the learning intention of the user, and intelligently matches the activity to-be-taught knowledge points in combination with the course data, realizes the accurate identification of personalized learning needs, and through the combination of semantic matching, high-dimensional feature clustering and effect evaluation adopted by the content screening unit, the teaching video segments screened are helpful to ensure the relevance of the teaching content, and timely respond to the personalized learning needs of the user. On the basis of splicing the video segments, the blind area of explanation is automatically identified and the supplementary audio is generated, which makes up for the problem of insufficient theoretical explanation in the original video.
[0024] Other features and advantages of the embodiments of the present specification will be further disclosed in the following specific embodiments and drawings. BRIEF DESCRIPTION OF DRAWINGS
[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present specification, the drawings required to be used in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present specification, and other drawings can be obtained by those skilled in the art without creative labor on the basis of these drawings.
[0026] Figure 1 An application schematic diagram of the theory-practice integrated teaching intelligent agent provided by the present specification.
[0027] Figure 2 A structure schematic diagram of the theory-practice integrated teaching intelligent agent provided by the present specification.
[0028] Figure 3 A method flow schematic diagram of generating video content provided by the present specification.
[0029] Figure 4A flowchart of a method for generating a video description is provided in the specification.
[0030] Figure 5 A flowchart of a method for generating a knowledge label is provided in the specification.
[0031] Figure 6 A flowchart of a method for generating a set of video segments to be fused is provided in the specification.
[0032] Figure 7 A flowchart of a method for obtaining supplemental audio is provided in the specification.
[0033] Figure 8 A schematic diagram of an electronic device is provided in the specification. DETAILED DESCRIPTION
[0034] The technical solutions of the embodiments of the present specification will be explained and described below in combination with the drawings of the embodiments of the present specification. The following embodiments are only preferred embodiments of the present specification, and are not all. Based on the embodiments in the embodiments, other embodiments obtained by those skilled in the art without creative labor belong to the protection scope of the present specification.
[0035] The terms "first", "second", "third", and the like in the specification and claims of the present specification and the above drawings are used to distinguish different objects, and are not used to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.
[0036] In the following description, the appearance of terms such as "inner", "outer", "upper", "lower", "left", "right", etc. is only for the convenience of describing the embodiments and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present specification.
[0037] The data involved in the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of related data comply with relevant laws, regulations and standards of relevant countries and regions.
[0038] Before introducing the technical solutions recorded in the present specification, the application scenarios and related technologies of the technical solutions are introduced.
[0039] In the field of vocational education, skill training, etc., the traditional teaching mode is not synchronized in theory teaching and practical operation, but there is a long interval time. Learners have difficulty in effectively converting abstract knowledge into hands-on ability, and the teaching effect is poor. In addition, the problems of scattered teaching resources, low reuse rate, and insufficient personalized support also limit the improvement of teaching efficiency.
[0040] An agent is a core concept in the field of artificial intelligence, referring to an intelligent system that can perceive the environment, understand the task goal, make autonomous decisions, and take actions to achieve specific goals. An agent has the functions of perception, understanding and reasoning, decision-making and planning, action, learning and adaptation. Agents can process multi-modal data, and agents are applied to personalized learning, intelligent tutoring, learning behavior analysis, and teaching resource recommendation scenarios. Using agents for assisted teaching can break through the time and space limitations of traditional classrooms and provide personalized learning assistance for students. Agents continuously track the learning process through data-driven methods, accurately identify knowledge gaps, and improve learning efficiency. It helps to reduce the repetitive work burden of teachers, so they can focus more on teaching design.
[0041] Using agents for theory-practice integrated teaching can effectively improve the efficiency and quality of teaching. However, there is currently a lack of technical solutions on how to match and associate the vast amount of recorded practical demonstration videos with complex knowledge points, and achieve personalized teaching for users 41. Therefore, the present specification provides a theory-practice integrated teaching agent based on high-quality data sets, please refer to Figure 1 , including a video indexing unit 11, a course unit 12, an interaction analysis unit 31, a content filtering unit 21, and a generation unit 22,
[0042] The video indexing unit 11 stores a plurality of video segment information, the video segment information including video segment addresses, knowledge tags 55, basic information, content descriptions 54, and estimated effect descriptions;
[0043] The course unit 12 stores or receives course data, the course data including a plurality of knowledge point information, the knowledge point information including knowledge content and associated practical operation descriptions;
[0044] The interaction analysis unit 31 interacts with the user 41 and analyzes the interaction data generated by the interaction, and generates a to-be-taught knowledge point according to the interaction data and the course data;
[0045] The content filtering unit 21 retrieves the video segment information stored in the video indexing unit 11 according to the to-be-taught knowledge point and filters to obtain a to-be-fused video segment set;
[0046] The generating unit 22 generates a teaching video according to the set of video clips to be fused, and generates a supplementary audio according to the teaching video and the knowledge points to be taught, and superimposes the supplementary audio on the teaching video to display to the user 41.
[0047] The video indexing unit 11 is responsible for centralized management of all preprocessed operation videos 52. The video clip information stored therein not only records the file link of the video clip 52, but also contains structured data entries of multi-dimensional semantic information. For example, the knowledge label 55 is a fine-grained label generated by automatic semantic analysis, covering the theoretical knowledge points, operation skill points, safety specifications, and tool usage, supporting semantic-based accurate matching. The basic information records the technical data and context metadata of the video clip 52, such as duration, resolution, shooting time, on-screen personnel, and used equipment and materials. The content description 54 is a structured text summary of the operation process shown in the video clip, including operation steps, key action nodes, and technical points, which enhances the system's understanding of the video content 53.
[0048] Please refer to the attached Figure 2 The teaching intelligent agent also includes a video analysis unit that receives historical operation demonstration videos 51, divides the historical operation demonstration videos 51 into multiple video clips and stores them, obtains video clip addresses, analyzes each video clip to obtain video content 53, obtains content descriptions 54 according to the video content 53, generates a number of knowledge labels 55 according to the content descriptions 54, reads the duration, shooting date, involved people and objects of the video clip 52 and generates basic information, identifies the shooting scene information, explanation process information and explanation person information of the video clip 52, generates an estimated effect description according to the shooting scene information, explanation process information and explanation person information, and synchronizes the video clip address, knowledge label 55, basic information, content description 54 and estimated effect description to the video indexing unit 11.
[0049] The function of the video analysis unit is to perform intelligent multi-modal analysis and information extraction on the received historical operation demonstration videos 51, complete the processing from the original video to the teaching available data, and synchronize the structured results to the video indexing unit 11 to provide a data basis for subsequent personalized teaching content generation. A segmentation algorithm based on scene change, action switching or time interval is used to automatically divide long videos into multiple semantically complete video clips. Or other technologies disclosed in the art are used to divide the historical operation demonstration videos 51 to obtain video clips. As a recommended way, each video clip usually corresponds to an independent operation step or skill unit. Each video clip is assigned a unique storage path or identifier, a video clip address is generated, and it is stored in the resource library.
[0050] For each video segment, a key frame sequence is obtained through frame extraction technology. A deep learning model is used to identify objects, actions, tools, and working environments in the key frames, generating visual semantic information. The audio stream in the video is extracted and converted into text through speech recognition technology. Natural language processing techniques are used to understand the content of the explanation, forming speech text information. The visual semantic information and speech text information are aligned and fused to generate complete and accurate video content 53 for the operation process shown in the video segment. Based on the obtained video content 53, content description 54 is generated. Recommendations incorporate domain knowledge and contextual information to improve the professionalism and completeness of the description. For example, a domain knowledge base is introduced to extract relevant background knowledge such as standard procedures and technical specifications related to the current operation. The content of adjacent video segments 52 is also considered to ensure the coherence of the operation description. In combination with pre-set description templates (such as including operation steps, tools used, key nodes, and safety prompts), a language generation model outputs the text description.
[0051] Based on the content description 54, knowledge labels 55 are automatically generated. First, natural language processing techniques are used to perform word segmentation, entity recognition, and keyword extraction on the content description 54, identifying technical terms, operation verbs, and core concepts. These entities are then mapped to a pre-built domain knowledge graph, and the hierarchical relationships in the knowledge graph (such as "belongs to," "premise," and "related") are used to expand related knowledge points. Finally, knowledge labels 55 are generated. Knowledge labels 55 come from a set of fine-grained labels such as theoretical knowledge, skill points, operation specifications, and subject classifications, which are pre-configured by humans or extracted from textbooks or other literature in the field by large language models.
[0052] The metadata and content information of the video file are read to generate basic information. This includes technical parameters such as the duration, resolution, and frame rate of the video segments 52, time and space information such as the shooting date and location, personnel involved (such as the explainer and the operator) determined through face recognition or voice recognition, and objects involved (such as tools, materials, and equipment models) identified through object recognition technology.
[0053] To evaluate the teaching suitability of the video clip 52, a predicted effect description is generated. The predicted effect description is a comprehensive evaluation information generated according to the scene type, picture quality, multi-angle support, explanation completeness, language clarity and fluency, interactivity and prompting, knowledge point concentration, and lecturer identity information of the video clip 52, which is used to screen high-quality teaching resources. The scene type includes work site, training classroom, studio, or virtual environment, and the effect degree decreases in turn. The resolution, stability, lighting conditions, and composition rationality of the video clip 52 are analyzed. The composition rationality is determined according to the visibility of the operation part. The video with shaking picture, dark light, or blocked key operation area has poor predicted effect. The video with clear and stable picture, bright light, and multi-angle switching is helpful for learners to understand the operation details comprehensively, and has better predicted effect. The explanation content and expression method of the video clip 52 are analyzed by combining speech recognition and natural language processing. The explanation completeness is the degree of covering the key steps, technical points, and safety precautions of the operation, and the higher the explanation completeness, the better the predicted effect. The speech clarity, moderate speed, excessive verbal habits, or pauses are evaluated, and the existence of them reduces the predicted effect. The teaching strategies such as questioning, emphasizing, and summarizing in the explanation are identified, and the existence of them improves the predicted effect. The knowledge point concentration represents the number of knowledge points explained in a unit time, and the more the number of knowledge points explained in a unit time, the lower the knowledge point concentration, and the lower the predicted effect. The lecturer identity information is obtained through metadata or speech / visual recognition, and the lecturer is determined to be a professional teacher, an industry technician, or an ordinary student, and the higher the identity authority, the better the predicted effect.
[0054] The fusion analysis is performed to generate the predicted effect description in the form of natural language text. For example, the predicted effect description is: “The picture is clear and stable, and the key steps are displayed in multiple angles, which has strong teaching visibility. The explanation logic is clear, and the terminology is standard, but the speed is too fast, which may be difficult for beginners to follow. The training classroom is shot in a real environment, but the background is cluttered, and some operation details are blocked. The lecturer is professional and authoritative, and the demonstration action is standard, but the safety precautions are not emphasized enough.”
[0055] The video analysis unit analyzes each video clip to obtain the video content 53, please refer to the attached Figure 3 The following steps are performed:
[0056] Step S101) Extract the frame images of the video clip to obtain a key frame sequence.
[0057] Step S102) Use the pre-accessed image recognition model to identify the objects, actions, and scenes of the key frame sequence, and generate visual semantic information according to the identified objects, actions, and scenes.
[0058] Step S103) Extract the speech of the video segment 52 and perform speech recognition to obtain the speech text.
[0059] Step S104) Generate the video content 53 based on the visual semantic information and the speech text.
[0060] Frame extraction is performed on the current video segment, and a key frame extraction method is used to retain only frames where the picture has changed significantly or contains important information. Key frame extraction is performed using techniques known in the art. Suppose a historical practical demonstration video 51 about "car engine oil replacement" is received. The video parsing unit divides it into multiple segments, one of which, about 2 minutes long, shows the operation step of "removing the oil filler cap cover". Key frame 1: the technician stands in front of the vehicle engine compartment, with both hands close to the filler area. Key frame 2: close-up shot, the technician's hand is twisting the oil filler cap cover. Key frame 3: the filler cap cover has been removed, and the technician places it on a clean cloth. The key frame sequence condenses the visual information of the operation process.
[0061] Using a pre-trained deep learning image recognition model, the frames in the key frame sequence are analyzed to identify objects, actions and scenes, and visual semantic information is generated. Key objects such as "oil filler cap cover", "engine compartment", "clean cloth", and "gloves" are identified. Actions such as "hand grip", "twist and rotate", and "place" are identified. The scene is judged to be "car repair workshop", with a clean and well-lit environment. The visual semantic information generated is: "The technician removes the oil filler cap cover from the engine in the repair workshop by rotating it with his hands, and then places it on a clean cloth".
[0062] The audio stream in the video segment is extracted synchronously, and speech recognition technology is used to convert the speech into text. For example, the identified explanation is: "Next, we need to open the oil filler cap cover first, so that we can maintain ventilation inside the engine when we drain the oil, to avoid creating negative pressure that affects the oil draining speed. Note that you should rotate it counterclockwise, with moderate force, and do not use brute force". The visual semantic information and the speech text are aligned and fused, redundancies are eliminated, and details are supplemented, and finally a complete and accurate video content 53 description is generated.
[0063] For example, the visual information of the video segment 52 confirms the action and object of "removing the filler cap cover". The speech text explains the purpose of the operation (maintaining ventilation), the direction (counterclockwise), and the operation points (moderate force). Finally, through natural language generation technology, the two are integrated into a coherent description: "The technician rotates the oil filler cap cover counterclockwise and removes it from the engine in the car repair workshop, and places it on a clean cloth. Open the engine ventilation to ensure smooth oil draining later. Pay attention to the force when operating to avoid damaging the threads".
[0064] The video parsing unit obtains the content description 54 according to the video content 53, please refer to the attached Figure 4 , the following steps are performed:
[0065] Step S201) Read the field of the historical practical demonstration video 51, and obtain the field knowledge base.
[0066] Step S202) Extract the field knowledge associated with the video content 53 from the field knowledge base.
[0067] Step S203) Read the video content 53 of the adjacent video segment 52 as context content.
[0068] Step S204) Read the preset content description 54 prompt template, which includes operation steps, tools used, key action nodes, and precautions.
[0069] Step S205) Generate content description 54 according to the associated field knowledge, context content, video content 53, and prompt template.
[0070] After the video parsing unit obtains the video content 53 of the video segment 52, it needs to further generate a structured and standardized content description 54 to facilitate subsequent knowledge labeling and teaching applications. This process is not simply a repetition, but needs to be combined with field knowledge, context logic, and preset specifications. Identify the professional field to which the video belongs. Through metadata or content analysis, the system determines that the video belongs to the "automobile maintenance and repair" field. Then, the corresponding field knowledge base is called. This knowledge base contains structured knowledge such as automobile structure, maintenance standard process, safety specifications, and tool usage manual. For example, the knowledge base clearly records: "In the engine oil replacement process, opening the filler cap is a standard pre-step to balance the air pressure and prevent slow oil drainage."
[0071] With the current video content 53 as a query condition, semantic matching is performed in the field knowledge base to extract relevant professional knowledge. The following knowledge is matched and extracted: "The function of the engine oil filler cap is to seal the engine crankcase ventilation system", "Before replacing the engine oil, the filler cap must be opened, otherwise it will cause slow oil drainage", "When disassembling, hand force should be used, and tools should not be used to forcibly disassemble to prevent thread stripping". Read the video content 53 of the adjacent video segment 52 as context content, the content of the previous segment is: "The technician locates the engine oil drain bolt under the vehicle chassis", and the content of the next segment is: "The technician uses a wrench to loosen the drain bolt and starts draining old engine oil".
[0072] Load the preset content description 54 prompt template, for example, the template is as follows:
[0073] Operation steps: [briefly describe the operation];
[0074] Tools used: [List tools, if none, write "by hand"];
[0075] Key action node: [Details of key action];
[0076] Notes: [Safety, norms, common mistakes].
[0077] According to the associated domain knowledge, context content, video content 53, and prompt template, the generated content description 54 is: "Operation steps: Before starting to drain old engine oil, manually open the engine oil filler cap to keep the engine interior ventilated. Tools used: by hand. Key action node: Rotate the cap counterclockwise until it is removed and place it on a clean cloth to avoid contamination. Notes: When operating, use even force to rotate, do not use tools to forcibly disassemble, as this may damage the threads; this step is a necessary preparation for oil draining and cannot be omitted."
[0078] When the video analysis unit generates a number of knowledge tags 55 according to the content description 54, please refer to Appendix Figure 5 , the following steps are performed:
[0079] Step S301) Tokenize and perform entity recognition on the content description 54, extracting technical terms, action verbs, and key concepts.
[0080] Step S302) Map the extracted entities to the pre-set knowledge graph nodes.
[0081] Step S303) Extract relevant concept tags and hierarchical relationships from the knowledge graph, generating a number of knowledge tags 55, including theoretical knowledge points, skill points, operation norms, and subject classification.
[0082] After generating the standardized content description 54, the video analysis unit needs to further convert it into machine-understandable and searchable knowledge tags 55. Natural language processing is performed on the content description 54, performing tokenization and named entity recognition to identify the key semantic units. From the content description 54, the following entities are extracted: technical terms: "engine oil filler cap", "engine", "ventilation", "threads", "draining oil"; action verbs: "open", "rotate", "remove", "place", "disassemble"; key concepts: "hand force", "contamination", "necessary preparation", "prohibition of tool use".
[0083] Read the pre-built "automobile maintenance domain knowledge graph". This graph organizes knowledge in a graph structure, with nodes representing knowledge points, skills, tools, norms, etc., and edges representing relationships between them (such as "belongs to", "premise", "related", etc.).
[0084] The extracted entities are matched and aligned with the nodes in the knowledge graph one by one. Exemplary,
[0085] The "oil filler cap" is mapped to the "filler assembly" node under the "engine lubrication system" in the knowledge graph. "Open", "rotate", and "remove" are mapped to the "disassembly operation" skill node under the "basic maintenance skill" category. "Prohibit forced disassembly with tools" is mapped to the "bare-handed operation principle" sub-node under the "safety operation specification" node. "Maintain engine internal ventilation" is mapped to the "engine oil change principle" theoretical knowledge point. Through mapping, unstructured text entities are converted into structured knowledge graphs. Starting from the mapped nodes, relationship expansion is performed in the knowledge graph to extract directly related upper and lower concepts, prerequisite knowledge, related skills, etc., generating a set of multi-dimensional and fine-grained knowledge labels 55. Exemplary, the knowledge labels 55 include: {theoretical knowledge points | engine crankcase ventilation principle, lubrication system sealing mechanism}, {skill points | manual disassembly of sealing cover, standard maintenance pre-operation}, {operation specifications | prohibit disassembly of filler cap with tools, component cleaning and storage specifications}, {discipline classification | automobile maintenance and repair, engine system maintenance}. It is identified that the operation is a pre-step of the "oil change process", so the label "oil change standard process" is automatically associated to ensure that the video segment can be retrieved through more knowledge points.
[0086] The interaction analysis unit 31 interacts with the user 41 and analyzes the interaction data generated by the interaction, and when generating the knowledge points to be taught according to the interaction data and the course data, the following steps are performed:
[0087] Receiving interactive natural language with the user 41;
[0088] Using a pre-established intent recognition model, obtaining the learning goal of the user 41 according to the interactive natural language;
[0089] Semantically matching the learning goal with the knowledge point information in the course data to obtain the knowledge point information with the highest matching degree;
[0090] Obtaining the knowledge points to be taught according to the knowledge point information with the highest matching degree.
[0091] The interaction analysis unit 31 realizes the conversion from the fuzzy user 41 request to the relatively accurate teaching knowledge point through natural language understanding and semantic matching technology. A student majoring in automobile maintenance wants to learn the "oil replacement" related skills through the intelligent teaching system. He inputs by voice or text: "I want to learn how to operate when replacing oil, especially how to open the cap without damaging it?" The original input of the user 41 is received through the user 41 interface (such as a voice input box, a chat window). The statement contains two pieces of information: the main learning goal: "learn the operation of oil replacement"; the specific concern: "how to open the oil filler cap without damaging it".
[0092] Using a pre-established intent recognition model, usually based on BERT, RoBERTa deep learning architecture, the learning goal of the user 41 is obtained according to the interaction natural language. The intent recognition model is trained on a large amount of teaching dialogue data and can identify the learning intent of the user 41, the problem type (such as "learning operation", "troubleshooting", "understanding principle"). The exemplary intent recognition model outputs the learning goal of the user 41 as: "learn the correct disassembly method of the oil filler cap in the oil replacement process, focus on operation skills and precautions to avoid damaging parts".
[0093] The learning goal is semantically matched with the knowledge point information in the course data. For example, the course contains the following related knowledge point information as shown in Table 1.
[0094] Table 1 Knowledge point information contained in the course
[0095]
[0096] The calculation result shows that the knowledge point K001 matching degree: 65% (related to "oil replacement", but not focused on "cap"). Knowledge point K002 matching degree: 40% (related to "disassembly", but the object is "bolt"). Knowledge point K003 matching degree: 92% (explicitly contains keywords such as "filler cap", "disassembly by hand", "prevention of damage", etc.). Knowledge point K004 matching degree: 35% (irrelevant). Therefore, the knowledge point information with the highest matching degree is determined to be knowledge point K003. The matched knowledge point information K003 is set as the teaching knowledge point of this teaching service. It contains complete theoretical knowledge (principle of engine ventilation system) and operation requirements (disassembly by hand, prohibited use of tools).
[0097] The content screening unit 21 retrieves the video segment information stored in the video index unit 11 according to the teaching knowledge point and screens to obtain a set of to-be-fused video segments. Please refer to the accompanying drawings Figure 6 , the following steps are performed:
[0098] Step S401) Compare the knowledge point to be taught with the knowledge tags 55 of the video segment information, and obtain all video segment information that is semantically matched;
[0099] Step S402) Generate a high-dimensional feature vector of the video segment information according to the basic information and content description 54 of the video segment information;
[0100] Step S403) Cluster all video segment information according to the high-dimensional feature vector, obtain a plurality of cluster centers, and obtain video segment information closest to each cluster center as a candidate set, respectively;
[0101] Step S404) Calculate the matching degree of the estimated effect description of the video segment information in the candidate set and the knowledge point to be taught, respectively, and eliminate the video segment information with a matching degree lower than a preset threshold;
[0102] Step S405) Obtain a to-be-fused video segment set according to the remaining video segment information in the candidate set.
[0103] The core concepts in the knowledge point to be taught K003, "oil filler cap", "manual disassembly", "thread protection", and "ventilation system", are taken as query conditions and are semantically matched with the knowledge tags 55 of all video segments 52 in the video index unit 11.
[0104] For example, the following 5 related segments are retrieved.
[0105] Video segment F1: The technician rotates the filler cap counterclockwise by hand to remove it and places it on a clean cloth (knowledge content: manual disassembly, thread protection, and component cleaning).
[0106] Video segment F2: The technician uses a wrench to disassemble the rusted filler cap (knowledge content: tool disassembly and fault handling).
[0107] Video segment F3: A close-up shot shows the internal thread structure of the filler cap (knowledge content: component structure and theoretical explanation).
[0108] Video segment F4: The technician quickly opens the cap by hand without explanation (knowledge content: manual disassembly and operation demonstration).
[0109] Video segment F5: The instructor draws a diagram on the classroom whiteboard to explain the ventilation principle (knowledge content: ventilation principle and theoretical teaching).
[0110] According to the label semantic similarity screening, the video segments F1, F3, F4, and F5 highly related to K003 are retained, and the inconsistent video segment F2 (using a tool, violating the operation specification) is eliminated.
[0111] The basic information and content description 54 are converted into a numerical high-dimensional feature vector. Video segments F1 to F5 are mapped to feature vectors A, B, C, D respectively. All video segment information is clustered according to the high-dimensional feature vector to obtain a plurality of cluster centers, and the video segment information closest to each cluster center is obtained as a candidate set. The feature vectors are clustered and analyzed (such as the K-Means algorithm), with the purpose of avoiding content homogenization and ensuring teaching diversity.
[0112] The clustering results are as follows: cluster 1 (biased towards practical demonstration): video segments F1, F4, select the video segment F1 closest to the center (higher quality); cluster 2 (biased towards theoretical explanation): video segments F3, F5, select the video segment F5 closest to the center (more systematic explanation); cluster 3 (component close-up): video segment F3 (isolated point), select video segment F3. From the 4 candidate segments, 3 most representative segments are selected: video segments F1, F3, F5, to form the candidate set.
[0113] The estimated effect description of each segment in the candidate set is read, and the matching degree with the knowledge point K003 to be taught is calculated. The estimated effect description of video segment F1 is "clear and stable picture, close-up shot highlighting key actions, explanation emphasizing force control, teaching effect optimal". Matching degree calculation: highly relevant to "disassembly by hand" and "prevention of damage", obtaining a matching degree of 95%. The estimated effect description of video segment F3 is "static image display, no operation demonstration, professional explanation but lack of practical operation correlation". Matching degree calculation: although it involves the concept of "thread", there is no operation guidance, obtaining a matching degree of 60% (lower than the preset threshold of 70%). The estimated effect description of video segment F5 is "clear explanation logic, but pure theoretical teaching without practical operation pictures". Matching degree calculation: related to "ventilation principle", but deviates from the focus of "operation skill", obtaining a matching degree of 65% (lower than the threshold). Video segments F3 and F5 are excluded, and only video segment F1 is retained. After clustering optimization and quality filtering, only video segment F1 is left in the candidate set, which is determined as the to-be-fused video segment set.
[0114] The generating unit 22 generates the supplementary audio according to the teaching video and the knowledge point to be taught. Please refer to the attached Figure 7 , the following steps are performed:
[0115] Step S501) Identify the content integrity and rhythm of the original audio in the teaching video for the explanation of the knowledge point to be taught.
[0116] Step S502) Obtain a plurality of knowledge points not fully explained in the teaching video, and generate a plurality of supplementary explanation texts according to the knowledge content stored in the course unit 12.
[0117] Step S503) generating a plurality of voice segments according to the plurality of supplementary explanation texts.
[0118] Step S504) aligning each of the voice segments to a corresponding time position in the teaching video, respectively.
[0119] Step S505) obtaining supplementary audio according to each voice segment and its time position.
[0120] Transcribe the original audio in the video segment F1 into text to obtain "open the filler cap, prepare to drain oil." Then, compare and analyze the text with the complete knowledge content of the knowledge point K003 to be taught.
[0121] Knowledge integrity assessment: the knowledge point K003 to be taught requires mastering "why to open the cap" (balance air pressure, prevent negative pressure), "how to operate correctly" (manually, uniform force), and "precautions" (prohibit using tools, prevent thread damage). However, the original audio only mentions the operation name and does not involve principles, skills, and specifications, with an integrity score of only 30%.
[0122] Rhythm analysis: the original audio statement is short, appearing 2 seconds before the operation starts and lasting 3 seconds, followed by complete silence. It is judged that the explanation rhythm is too fast and the information density is too low, making it difficult for learners to fully absorb.
[0123] It is judged that there are a large number of insufficiently explained knowledge points, and supplementary audio needs to be generated. Identify the key knowledge points missing in the original video and extract the corresponding knowledge content from course unit 12 to generate targeted supplementary explanation texts.
[0124] Missing point 1: operation principle. Supplementary text: "Opening the oil filler cap is to connect the engine interior with the atmosphere to form a ventilation circuit. If not opened, negative pressure will form in the crankcase during oil draining, causing slow or even interrupted old oil discharge". Missing point 2: operation specification. Supplementary text: "Must operate manually, use palm perception to control rotation force. Strictly prohibit using wrenches and other tools to forcibly disassemble, otherwise it will easily cause plastic cover body breakage or metal thread galling, causing maintenance accidents". Missing point 3: operation key points. Supplementary text: "When disassembling, keep your wrist stable and rotate counterclockwise at a uniform speed. If you feel excessive resistance, pause and check for foreign object jamming, do not use brute force". Call the voice synthesis engine to convert the above three supplementary explanation texts into natural and smooth voice segments respectively. The voice synthesis engine uses technology disclosed in the art. Analyze the timeline of the video segment F1, identify the key action nodes, and synchronize and align the voice segments with the actions. Merge the three aligned voice segments in chronological order to generate a continuous supplementary audio. The audio is mixed with the background environment sound (such as slight mechanical sound) of the original video.
[0125] In another aspect, the present specification provides a construction method of a theory-practice integrated teaching intelligent agent based on a high-quality data set, comprising the following steps:
[0126] receiving and storing a plurality of video segment information, the video segment information comprising a video segment address, a knowledge tag 55, basic information, a content description 54 and a predicted effect description, and constructing a video index unit 11 according to the video segment information;
[0127] receiving and storing course data, the course data comprising a plurality of knowledge point information, the knowledge point information comprising knowledge content and associated practical operation description, and constructing a course unit 12 according to the course data;
[0128] constructing a first intelligent agent unit, the first intelligent agent unit performing interaction with a user 41 and analyzing interaction data generated by the interaction, generating a to-be-taught knowledge point according to the interaction data and the course data, and constructing an interaction analysis unit 31 according to the first intelligent agent unit;
[0129] constructing a retrieval function unit, the retrieval function unit performing retrieval of video segment information stored in the video index unit 11 according to the to-be-taught knowledge point and performing screening, obtaining a to-be-fused video segment set, and constructing a content screening unit 21 according to the retrieval function unit;
[0130] constructing a second intelligent agent unit, the second intelligent agent unit performing generation of a teaching video according to the to-be-fused video segment set, and generating a supplementary audio according to the teaching video and the to-be-taught knowledge point, superimposing the supplementary audio onto the teaching video and displaying the teaching video to the user 41, and constructing a generation unit 22 according to the second intelligent agent unit.
[0131] wherein the method for obtaining the video segment information comprises the following steps:
[0132] receiving a historical practical operation demonstration video 51, dividing the historical practical operation demonstration video 51 into a plurality of video segments and storing the video segments, and obtaining a video segment address;
[0133] analyzing each video segment to obtain a video content 53, obtaining a content description 54 according to the video content 53, and generating a plurality of knowledge tags 55 according to the content description 54;
[0134] reading the length, shooting date, involved persons and objects of the video segment 52 and generating basic information;
[0135] identifying shooting scene information, explanation process information and explanation person information of the video segment 52, and generating a predicted effect description according to the shooting scene information, explanation process information and explanation person information;
[0136] The video clip information is generated according to the video clip address, the knowledge tag 55, the basic information, the content description 54 and the estimated effect description.
[0137] Referring to Figure 8 The electronic device provided by the embodiment of the present specification has the structure as shown in the figure.
[0138] As Figure 8 As shown in the figure, the electronic device 1100 can include at least one processor 1101, at least one network interface 1104, a user interface 1103, a memory 1105 and at least one communication bus 1102. The communication bus 1102 can be used to realize the connection and communication of the above-mentioned components. The user interface 1103 can include a key, and the optional user interface can also include a standard wired interface, a wireless interface. The network interface 1104 can include but is not limited to a Bluetooth module, an NFC module, a Wi-Fi module, etc. The processor 1101 can include one or more processing cores. The processor 1101 connects various parts in the electronic device 1100 through various interfaces and lines, executes various functions of the routing device and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 1105, and calling data stored in the memory 1105. Optionally, the processor 1101 can be realized by at least one of DSP, FPGA and PLA. The processor 1101 can integrate CPU, GPU and modem, etc. The CPU is mainly used to process the operating system, user interface and application program, etc.; the GPU is used to render and draw the content to be displayed on the display screen; and the modem is used to process wireless communication.
[0139] It can be understood that the above-mentioned modem can also not be integrated into the processor 1101, but be realized by a separate chip.
[0140] The memory 1105 can include a RAM and can also include a ROM. Optionally, the memory 1105 includes a non-transitory computer-readable medium. The memory 1105 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 1105 can include a program storage area and a data storage area, where the program storage area can store the instructions for implementing the operating system, the instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), the instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store the data involved in the above-mentioned various method embodiments, etc. The memory 1105 can also be at least one storage device located away from the aforementioned processor 1101. The memory 1105, as a computer storage medium, can include an operating system, a network communication module, a user 41 interface module, and an application program. The processor 1101 can be used to invoke the application program stored in the memory 1105 and execute the method in the above-mentioned various embodiments.
[0141] The embodiments of the present specification also provide a computer-readable storage medium, which stores instructions, and when the instructions run on a computer or a processor, the computer or the processor executes the steps in the above-mentioned embodiments. The various constituent modules of the above-mentioned electronic device, if realized in the form of software function units and sold or used as independent products, can be stored in the computer-readable storage medium.
[0142] The embodiments of the present specification also provide a computer program product, which includes a computer program, and the computer program is executed by a processor to realize the steps in the above-mentioned embodiments.
[0143] The technical features in the embodiments and the implementation forms can be combined arbitrarily without conflict.
[0144] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes a plurality of computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present specification are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in or transmitted by a computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center through a wired (such as a coaxial cable, an optical fiber, a digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server, data center, etc. integrated with a plurality of available media. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a digital versatile disc (DVD)), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
[0145] When implemented by hardware or firmware, the foregoing method processes are programmed into a hardware circuit to obtain a corresponding hardware circuit structure, and the corresponding functions are implemented. For example, a programmable logic device (PLD) (such as a field programmable gate array (FPGA)) is an integrated circuit whose logic function is determined by the user 41 programming the device. A digital system is "integrated" on a PLD by the designer himself programming, without having to ask the chip manufacturer to design and manufacture a special integrated circuit chip. Moreover, instead of manually making integrated circuit chips, such programming is now mostly implemented using "logic compiler" software, which is similar to the software compiler used when developing programs, and the original code before compilation must also be written in a specific programming language, which is called a hardware description language (HDL), and there are many HDLs. Those skilled in the art should also be clear that only the method processes need to be logically programmed in the above-mentioned several hardware description languages and programmed into integrated circuits, and the hardware circuit that implements the logical method processes can be easily obtained.
[0146] The above-described embodiments are merely exemplary diagnostic description rather than limitations on the scope of the present specification. Any modifications, equivalent replacements, and improvements made to the technical solutions of the present specification by those of ordinary skill in the art, without departing from the design spirit of the present specification, shall fall within the protective scope of the claims of the present specification.
Claims
1. A high-quality data set-based theory-practice integrated teaching intelligent agent, characterized in that, it comprises a video indexing unit, a course unit, an interaction analysis unit, a content screening unit and a generation unit, the video indexing unit stores a plurality of video segment information, the video segment information comprising a video segment address, a knowledge label, basic information, content description and estimated effect description; the course unit stores or receives course data, the course data comprising a plurality of knowledge point information, the knowledge point information comprising knowledge content and associated practical operation description; the interaction analysis unit interacts with the user and analyzes the interaction data generated by the interaction, generates a knowledge point to be taught according to the interaction data and the course data; the content screening unit screens the video segment information stored in the video indexing unit according to the knowledge point to be taught, and obtains a set of to-be-fused video segments; the generation unit generates a teaching video according to the set of to-be-fused video segments, generates a supplementary audio according to the teaching video and the knowledge point to be taught, and displays the teaching video with the supplementary audio superimposed.
2. The high-quality data set-based theory-practice integrated teaching intelligent agent according to claim 1, characterized in that, it further comprises a video analysis unit, which receives a historical practical operation demonstration video, divides the historical practical operation demonstration video into a plurality of video segments and stores them, obtains a video segment address, analyzes each video segment to obtain video content, obtains content description according to the video content, generates a plurality of knowledge labels according to the content description, reads the duration, shooting date, involved persons and objects of the video segment and generates basic information, identifies the shooting scene information, explanation process information and explanation person information of the video segment, generates estimated effect description according to the shooting scene information, explanation process information and explanation person information, and synchronizes the video segment address, knowledge label, basic information, content description and estimated effect description to the video indexing unit.
3. The high-quality data set-based theory-practice integrated teaching intelligent agent according to claim 2, characterized in that, when the video analysis unit analyzes each video segment to obtain video content, the following steps are performed: extract the frame images of the video segment to obtain a key frame sequence; use a pre-accessed image recognition model to identify objects, actions and scenes in the key frame sequence, and generate visual semantic information according to the identified objects, actions and scenes; extract the speech of the video segment and perform speech recognition to obtain speech text; generate video content according to the visual semantic information and the speech text.
4. The high-quality data set-based theory-practice integrated teaching intelligent agent according to claim 3, characterized in that, when the video analysis unit obtains content description according to the video content, the following steps are performed: read the field of the historical practical operation demonstration video to obtain a field knowledge base; extract the field knowledge associated with the video content from the field knowledge base; read the video content of the adjacent video segment as context content; reading a preset content description prompt template, the prompt template including operation steps, tools used, key action nodes, and tips; generating a content description according to the associated domain knowledge, context content, video content, and prompt template.
5. The intelligent agent for theory-practice integrated teaching based on high-quality data sets according to claim 4, characterized in that, when the video analysis unit generates a plurality of knowledge tags according to the content description, the following steps are performed: performing word segmentation and entity recognition on the content description to extract technical terms, operation verbs, and key concepts; mapping the extracted entities to preset knowledge graph nodes; extracting related concept tags and hierarchical relationships from the knowledge graph to generate a plurality of knowledge tags, including theoretical knowledge points, skill points, operation specifications, and subject classifications.
6. The intelligent agent for theory-practice integrated teaching based on high-quality data sets according to any one of claims 1 to 5, characterized in that, when the interaction analysis unit interacts with the user and analyzes the interaction data generated by the interaction, and generates a knowledge point to be taught according to the interaction data and course data, the following steps are performed: receiving interactive natural language with the user; using a pre-established intent recognition model to obtain the user's learning goal according to the interactive natural language; performing semantic matching on the learning goal and the knowledge point information in the course data to obtain the knowledge point information with the highest matching degree; obtaining the knowledge point to be taught according to the knowledge point information with the highest matching degree.
7. The intelligent agent for theory-practice integrated teaching based on high-quality data sets according to any one of claims 1 to 5, characterized in that, when the content screening unit retrieves video segment information stored in the video index unit according to the knowledge point to be taught and performs screening to obtain a set of video segments to be fused, the following steps are performed: comparing the knowledge point to be taught with the knowledge tags of the video segment information to obtain all video segment information with semantic matching; generating a high-dimensional feature vector of the video segment information according to the basic information and content description of the video segment information; clustering all video segment information according to the high-dimensional feature vector to obtain a plurality of cluster centers, and obtaining the video segment information closest to each cluster center as a candidate set; calculating the matching degree between the estimated effect description of the video segment information in the candidate set and the knowledge point to be taught, and removing the video segment information with a matching degree lower than a preset threshold; obtaining the set of video segments to be fused according to the remaining video segment information in the candidate set.
8. The intelligent agent for theory-practice integrated teaching based on high-quality data sets according to any one of claims 1 to 5, characterized in that, when the generation unit generates supplementary audio according to the teaching video and the knowledge point to be taught, the following steps are performed: identifying the content integrity and rhythm of the original audio in the teaching video for the knowledge point to be taught; obtaining a plurality of knowledge points not fully explained in the teaching video, and generating a plurality of supplementary explanation texts according to the knowledge content stored in the course unit; generating a plurality of voice segments according to the plurality of supplementary explanation texts; aligning each of the voice segments to a corresponding time position in the teaching video respectively; obtaining supplementary audio according to each voice segment and the time position thereof. 9.The construction method of the theory-practice integrated teaching intelligent agent based on a high-quality dataset according to claim 1, characterized in that, comprising the steps of: receiving and storing a plurality of video segment information, the video segment information comprising a video segment address, a knowledge label, basic information, a content description, and an estimated effect description, and constructing a video index unit according to the video segment information; receiving and storing course data, the course data comprising a plurality of knowledge point information, the knowledge point information comprising knowledge content and associated practical operation description, and constructing a course unit according to the course data; constructing a first intelligent agent unit, the first intelligent agent unit performing interaction with a user and analyzing interaction data generated by the interaction, generating a knowledge point to be taught according to the interaction data and the course data, and constructing an interaction analysis unit according to the first intelligent agent unit; constructing a search function unit, the search function unit performing searching and screening of the video segment information stored in the video index unit according to the knowledge point to be taught, obtaining a set of video segments to be fused, and constructing a content screening unit according to the search function unit; constructing a second intelligent agent unit, the second intelligent agent unit performing generation of a teaching video according to the set of video segments to be fused, and generating supplementary audio according to the teaching video and the knowledge point to be taught, superimposing the supplementary audio onto the teaching video, and displaying the teaching video to the user, and constructing a generation unit according to the second intelligent agent unit. 10.The construction method according to claim 9, characterized in that, the method of obtaining the video segment information comprises the steps of: receiving a historical practical operation demonstration video, dividing the historical practical operation demonstration video into a plurality of video segments and storing the video segments, and obtaining a video segment address; analyzing each video segment to obtain video content, obtaining a content description according to the video content, and generating a plurality of knowledge labels according to the content description; reading the length of the video segment, the shooting date, the involved person and the involved article, and generating basic information; identifying shooting scene information, explanation process information, and explanation person information of the video segment, and generating an estimated effect description according to the shooting scene information, the explanation process information, and the explanation person information; generating the video segment information according to the video segment address, the knowledge label, the basic information, the content description, and the estimated effect description.
Citation Information
Patent Citations
Interactive teaching method and system based on teaching video
CN120318743A
Video retrieval based contextualized learning
US20250077859A1