Course arrangement method and device, electronic equipment, storage medium and program product

By choreographing course resource materials on the visual interface and generating course agents, the problem of low efficiency in educational course video production in the existing technology is solved, and fast and efficient course production is achieved.

CN120144022APending Publication Date: 2025-06-13ZHEJIANG PRISM HOLOGRAPHIC TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510230686.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the process of manually recording videos of educational courses, the existing technology requires long-term equipment debugging, recording materials, screening and sorting, and editing clips. Most of the materials need to be re-recorded when teaching objectives or audience preferences change, which is less efficient.

Method used

Through the user's orchestration operation on the visual interface, the course resource materials are arranged, multiple course scenes are generated, and the course agent is generated using pre-made digital human images to achieve multimodal interaction with the teaching object.

Benefits of technology

Reduce dependence on real lecturers, can quickly reuse resource materials and digital human image, significantly shorten the course production cycle, and improve the efficiency of course production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144022A_ABST
    Figure CN120144022A_ABST
Patent Text Reader

Abstract

The invention provides a course arrangement method and device, electronic equipment, a storage medium and a program product, and the method comprises the steps: carrying out the arrangement of course resource materials selected by a user in response to the arrangement operation of the user on a visual interface, and obtaining a plurality of course scenes; and generating a course intelligent agent according to the plurality of course scenes and a pre-made digital human image, wherein the course intelligent agent is used for carrying out multi-modal interaction with the teaching object to realize teaching. In the implementation process of the scheme, the course resource materials are arranged in response to the arrangement operation of the user on the visual interface, the course intelligent agent is generated according to the arranged course scene and the pre-made digital human image, and the course intelligent agent is an education auxiliary intelligent agent integrated with the artificial intelligence technology. According to the method, course making can be completed in a short time only by simply performing arrangement operation on the visual interface, so that the course making period is greatly shortened, and the course making efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of education and artificial intelligence. Specifically, it relates to a course arrangement method, device, electronic device, storage medium, and program product. Background Art

[0002] Currently, in the process of manually recording educational course videos, a series of operations such as equipment debugging, recording materials, screening and sorting, and editing are required to complete the production of educational course videos, which takes a relatively long production time cycle. However, once the teaching objectives or audience preferences change, most of the materials need to be re-recorded and screened, sorted, and edited again, because each video is tailored for a real instructor in a specific scenario, and when it comes to being applied to other scenarios later, it often requires the real instructor to re-record. Therefore, the current method of manually recording and editing courses has low efficiency. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide a course arrangement method, device, electronic device, storage medium, and program product, which are used to improve the problem of low efficiency in making courses.

[0004] The embodiments of this application provide a course arrangement method, including: in response to an arrangement operation by a user on a visual interface, arranging the course resource materials selected by the user to obtain multiple course scenarios; generating a course intelligent agent according to the multiple course scenarios and a pre-made digital human image, and the course intelligent agent is used to interact with the teaching object in a multi-modal manner to achieve teaching. In the implementation process of the above solution, by responding to the arrangement operation by the user on the visual interface, arranging the course resource materials, and generating a course intelligent agent according to the arranged course scenarios and the pre-made digital human image, the course intelligent agent is an educational assistance intelligent agent integrating advanced artificial intelligence (AI) technology in the field of course teaching, so it can interact with the teaching object in a multi-modal manner to achieve teaching. Since these resource materials and digital human images are separated, they are convenient for free combination, and the digital human images that are not real people can be well reused without repeatedly recording videos, effectively reducing the dependence on real instructors. Only by simply performing an arrangement operation on the visual interface can the course production be completed in a short time, greatly shortening the production cycle of the course and improving the efficiency of making courses.

[0005] Optionally, in the embodiments of the present application, arranging the course resource materials selected by the user includes: processing the course resource materials using an artificial intelligence AI control, and arranging the processed resource materials; and / or, arranging the course resource materials into a course scenario, and setting an artificial intelligence AI control in the course scenario, where the AI control is used to interact with the teaching object to obtain an interaction result. In the implementation process of the above solution, the AI control can better support multi-modal interaction, and the use of the AI control to automatically process and arrange the course resource materials reduces the need for manual intervention, greatly accelerating the conversion process from the original materials to the finished course. Since most of the course resource materials can be handed over to the AI control for automatic processing, the course content can be quickly updated according to the latest research results or market demands without having to re-record the entire video, greatly improving the flexibility and response speed.

[0006] Optionally, in the embodiments of the present application, the AI control includes: a text-to-speech TTS module; processing the course resource materials using an artificial intelligence AI control includes: using the TTS module to perform natural speech synthesis on the text files in the course resource materials; and / or, setting an artificial intelligence AI control in the course scenario includes: setting the TTS module in the course scenario, where the TTS module is used to synthesize the text content generated in the interaction result into voice data. In the implementation process of the above solution, by using the TTS module to synthesize the text files into voice data, since most of the voice content is automatically generated, when the course needs to be updated or modified, only the corresponding text files need to be adjusted and the TTS module is re-run to quickly generate a new voice version without having to re-record the entire video, thus greatly reducing the voice recording work that originally required a large amount of time and human resources, enabling the work of synthesizing the text files into voice data to be automatically completed in a short time, and significantly improving the efficiency of course production.

[0007] Optionally, in the embodiments of the present application, the AI control includes: an automatic speech recognition ASR module; processing the course resource materials using an artificial intelligence AI control includes: using the ASR module to perform speech recognition on the audio files in the course resource materials to obtain the recognized ASR text; and / or, using the ASR module to perform speech recognition on the voice data in the video files in the course resource materials to obtain the recognized ASR text; and / or, setting an artificial intelligence AI control in the course scenario includes: setting the ASR module in the course scenario, where the ASR module is used to recognize the interaction content from the voice data of the interaction result and determine whether to switch the scenario according to the interaction content. In the implementation process of the above solution, after setting the ASR module in the course scenario, the ASR module can recognize the students' answers and switch the scenario, thus enhancing the interactivity and sense of participation of the course.

[0008] Optionally, in the embodiments of the present application, the AI control also includes: a natural language understanding (NLU) module; using the AI control to process the course resource materials further includes: using the NLU module to perform natural language understanding on the ASR text or the text files or audio files in the course resource materials; and / or, setting the AI control in the course scenario includes: setting the NLU module in the course scenario, and the NLU module is used to perform semantic understanding on the text content or voice data generated in the interaction result, and determine whether to perform scenario switching according to the result of the semantic understanding. In the implementation process of the above solution, through the NLU module, the answers, questions or discussion contents of the students are deeply understood, the intentions and emotions behind them are identified, and whether to perform scenario switching is determined according to the result of the semantic understanding, so as to provide more personalized feedback and suggestions, which enables each student to obtain the most suitable learning path and support for themselves, thereby enhancing the interactivity and sense of participation of the course. Further, based on the user semantics understood by the NLU module, the course agent can have a more rich and meaningful conversation with the students, enhancing the authenticity of the interaction and significantly improving the interaction quality in the course.

[0009] Optionally, in the embodiments of the present application, after arranging the course resource materials into a course scenario, it further includes: setting the interaction conditions corresponding to the AI control and the jump scenarios corresponding to the interaction conditions in the course scenario, and the interaction conditions are used to switch to the jump scenarios corresponding to the interaction conditions after matching the interaction results obtained by the AI control. In the implementation process of the above solution, by setting the interaction conditions corresponding to the AI control and the jump scenarios corresponding to the interaction conditions in the course scenario, when the answers of the students do not meet the expectations, targeted feedback can be provided immediately, and they can be guided to try again or enter different auxiliary teaching scenarios, ensuring that each student can get timely help and support. Further, when it is necessary to update or optimize the course content, only the corresponding rules need to be adjusted, and the changes can be quickly realized without re-recording the entire video, thereby improving the efficiency of making courses.

[0010] Optionally, in the embodiments of the present application, arranging the course resource materials selected by the user includes: editing the playing time and / or the duration of the course resource materials. In the implementation process of the above solution, by precisely controlling the playing time and the duration, the user can better manage the attention concentration period, avoiding information overload or missing important knowledge points caused by too long or too short content. When it is necessary to modify or optimize the course content, only the corresponding playing time and duration need to be adjusted, and a new course version can be quickly generated without re-recording the entire video, thereby improving the efficiency of making courses.

[0011] Optionally, in the embodiments of the present application, the course resource materials include: video files; arranging the course resource materials selected by the user includes: after the video file is imported, creating a blank scene and setting the timeline of the blank scene to the timeline of the video file; using an AI control to identify natural language text from the video file; marking out a content summary and / or interaction opportunities on the timeline of the video file according to the natural language text, and the interaction opportunities are used to prompt the user to set interaction conditions and the corresponding jump scenes at the marked positions.

[0012] In the implementation process of the above solution, by marking out a content summary and / or interaction opportunities on the timeline of the video file according to the natural language text identified from the video file by the AI control, the previous historical video can be quickly converted into a course intelligent body, improving the interactivity of the course. At the same time, it also avoids the waste of historical video resources and effectively improves the development and production efficiency of the course. In addition, by automatically processing and analyzing the imported video file through the AI module and marking out a content summary and / or interaction opportunities on the timeline of the video file according to the natural language text, since most of the work is completed by the AI, when the course content needs to be updated, only need to re-run the relevant AI module to process the new video file, and the latest version can be quickly generated without repeating the manual editing process, thus significantly improving the development and production efficiency of the course.

[0013] Optionally, in the embodiments of the present application, before generating a course intelligent body according to multiple course scenes and a pre-produced digital human image, it further includes: responding to a debugging request for the current scene and making debugging modifications to the current scene; and / or, responding to a preview request for the current scene and determining the preview effect of the current scene according to all the resource materials in the current scene. In the implementation process of the above solution, by initiating a debugging or preview request at any time during the creation process and immediately seeing the modified effect without waiting to adjust after the entire course arrangement is completed, this fast feedback mechanism greatly shortens the development cycle, thus significantly improving the efficiency of making courses.

[0014] Optionally, in the embodiments of the present application, after generating a course intelligent agent according to multiple course scenarios and a pre-made digital human image, it further includes: in response to the replacement operation of the digital human image, replacing the digital human image in the course intelligent agent with a new digital human image selected by the user. In the implementation process of the above solution, by replacing the digital human image in the course intelligent agent with a new digital human image selected by the user, since both the digital human image and the course resource materials are reusable, and the course intelligent agent embeds the digital human image into the course resource materials without binding the digital human image and the course resource materials. Therefore, the teacher digital human image in the course intelligent agent can be replaced with other digital human images at any time while keeping the course content in the intelligent agent unchanged, so that a large number of personalized course intelligent agents can be generated quickly to meet the preferences of different students, and at the same time, the efficiency of making courses is effectively improved.

[0015] Optionally, in the embodiments of the present application, the course intelligent agent is further used to collect the learning behavior habit data of the teaching object during the multi-modal interaction process, and the learning behavior habit data is used to optimize the course intelligent agent. In the implementation process of the above solution, by continuously collecting the learning behavior habit data of students, the course intelligent agent can more accurately understand the needs and preferences of each student, provide a more personalized learning path and support, thereby improving the personalization degree of the course intelligent agent.

[0016] The embodiments of the present application further provide a course arrangement device, including: a course scenario obtaining module, configured to arrange the course resource materials selected by the user in response to the arrangement operation of the user on the visual interface to obtain multiple course scenarios; a course intelligent agent generating module, configured to generate a course intelligent agent according to the multiple course scenarios and a pre-made digital human image, and the course intelligent agent is used to perform multi-modal interaction with the teaching object to implement teaching.

[0017] Optionally, in the embodiments of the present application, the course scenario obtaining module includes: a course resource arrangement sub-module, configured to process the course resource materials using an artificial intelligence AI control and arrange the processed resource materials; and / or, an AI control setting sub-module, configured to arrange the course resource materials into a course scenario and set an artificial intelligence AI control in the course scenario, and the AI control is used to interact with the teaching object to obtain an interaction result.

[0018] Optionally, in the embodiments of the present application, the AI control includes: a text-to-speech TTS module; the resource material arrangement sub-module includes: a natural speech synthesis unit, configured to perform natural speech synthesis on the text file in the course resource materials using the TTS module; and / or, the AI control setting sub-module includes: a TTS module setting unit, configured to set the TTS module in the course scenario, and the TTS module is used to synthesize the text content generated in the interaction result into voice data.

[0019] Optionally, in the embodiments of the present application, the AI control includes: an automatic speech recognition (ASR) module; a resource material arrangement sub-module, including: a material audio recognition unit for using the ASR module to perform speech recognition on the audio files in the course resource materials to obtain the recognized ASR text; and / or, a video audio recognition unit for using the ASR module to perform speech recognition on the speech data in the video files in the course resource materials to obtain the recognized ASR text; and / or, an AI control setting sub-module, including: an ASR module setting unit for setting the ASR module in the course scenario, where the ASR module is used to recognize the interaction content from the speech data of the interaction result and determine whether to perform a scenario switch according to the interaction content.

[0020] Optionally, in the embodiments of the present application, the AI control further includes: a natural language understanding (NLU) module; a resource material arrangement sub-module, including: a natural language understanding unit for using the NLU module to perform natural language understanding on the ASR text or the text files or audio files in the course resource materials; and / or, an AI control setting sub-module, including: an NLU module setting unit for setting the NLU module in the course scenario, where the NLU module is used to perform semantic understanding on the text content or speech data generated in the interaction result and determine whether to perform a scenario switch according to the result of the semantic understanding.

[0021] Optionally, in the embodiments of the present application, the course arrangement device further includes: an AI interaction setting module for setting the interaction conditions corresponding to the AI control and the jump scenarios corresponding to the interaction conditions in the course scenario, where the interaction conditions are used to switch to the jump scenarios corresponding to the interaction conditions after matching the interaction results obtained by the AI control.

[0022] Optionally, in the embodiments of the present application, the course scenario obtaining module includes: a resource material editing sub-module for editing the playing time and / or the duration of the course resource materials.

[0023] Optionally, in the embodiments of the present application, the course resource materials include: video files; the course scenario obtaining module includes: a scenario timeline setting sub-module for creating a blank scenario after the video files are imported and setting the timeline of the blank scenario as the timeline of the video files; a text recognition sub-module for using the AI control to recognize natural language text from the video files; an interaction opportunity annotation sub-module for annotating the content summary and / or interaction opportunities on the timeline of the video files according to the natural language text, where the interaction opportunities are used to prompt the user to set the interaction conditions and the jump scenarios corresponding to the interaction conditions at the annotation points.

[0024] Optionally, in the embodiments of the present application, the course arrangement device further includes: a scene debugging and modification module, configured to respond to a debugging request for the current scene and perform debugging and modification on the current scene; and / or, a current scene preview module, configured to respond to a preview request for the current scene and determine the preview effect of the current scene according to all resource materials in the current scene.

[0025] Optionally, in the embodiments of the present application, the course arrangement device further includes: a digital human image replacement module, configured to respond to an operation of replacing the digital human image and replace the digital human image in the course intelligent body with a new digital human image selected by the user.

[0026] Optionally, in the embodiments of the present application, the course intelligent body is further configured to collect learning behavior habit data of the teaching object during the multi-modal interaction process, and the learning behavior habit data is used to optimize the course intelligent body.

[0027] The embodiments of the present application further provide an electronic device, including: a processor and a memory, where the memory stores machine-readable instructions executable by the processor, and the machine-readable instructions, when run by the processor, execute the method described above.

[0028] The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored, and the computer program, when run by the processor, executes the method described above.

[0029] The embodiments of the present application further provide a computer program product, including: a computer program or computer instructions, and the computer program or computer instructions, when run by the processor, execute the method described above. Description of the Drawings

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required to be used in the embodiments of the present application. It should be understood that the following drawings only show some embodiments in the embodiments of the present application, and therefore should not be regarded as a limitation on the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0031] Figure 1 A flowchart showing the course arrangement method provided by the embodiments of the present application;

[0032] Figure 2 A schematic diagram showing the arrangement operation of the user on the visualization interface provided by the embodiments of the present application;

[0033] Figure 3 A schematic diagram showing the structure of the course arrangement device provided by the embodiments of the present application;

[0034] Figure 4 Schematic structural diagram of an electronic device provided by an embodiment of the present application shown. Detailed implementation manners

[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. It should be understood that the accompanying drawings in the embodiments of the present application are only for the purposes of illustration and description, and are not used to limit the protection scope of the embodiments of the present application. In addition, it should be understood that the schematic drawings are not drawn to actual scale. The flowcharts used in the embodiments of the present application illustrate operations implemented according to some embodiments of the embodiments of the present application. It should be understood that the operations in the flowcharts may not be implemented in sequence, and steps without logical context relationships may be reversed in order or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flowchart or remove one or more operations from the flowchart under the guidance of the content of the embodiments of the present application.

[0036] In addition, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. The components of the embodiments of the present application usually described and illustrated in the accompanying drawings here may be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the embodiments of the present application to be protected, but only represents the selected embodiments of the present application.

[0037] It can be understood that "first" and "second" in the embodiments of the present application are used to distinguish similar objects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit being different. In the description of the embodiments of the present application, the term "and / or" is only an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the front and rear associated objects. The term "multiple" refers to two or more (including two). Similarly, "multiple groups" refers to two or more groups (including two groups).

[0038] It should be noted that the course scheduling method provided by the embodiments of the present application can be executed by an electronic device. Here, the electronic device refers to a device terminal or a server having the function of executing a computer program. The device terminal is, for example: a smart phone, a personal computer, a tablet computer, a personal digital assistant, or a mobile Internet device, etc. The server refers to a device that provides computing services through a network. The server is, for example: an x86 server and a non-x86 server. The non-x86 server includes: mainframes, minicomputers, and UNIX servers.

[0039] In the related art, usually, creating course videos of educational content relies on manual video recording and editing. For example: First, the teaching syllabus needs to be carefully planned. Then, professional lecturers or teachers conduct teaching recordings at a specific venue. Finally, after the recording is completed, the editing team edits the original materials, including selecting appropriate segments, adding subtitles or other visual aids. The production cycle is relatively long. However, this method has an obvious efficiency bottleneck. From the preliminary preparation to the post-processing, each link requires a large amount of manpower and time investment. Once the teaching objectives or audience preferences change, the entire recording process and editing process need to be repeated. Therefore, in the face of rapidly changing teaching requirements, the efficiency of creating courses in this way of manual recording plus editing is relatively low.

[0040] In response to the above problems, the embodiments of the present application provide a course arrangement method. The main idea is that through the arrangement operations of the user on the visual interface, a course agent is generated based on the arranged course scenes and pre-made digital human images, which can effectively reduce the dependence on real lecturers, and can effectively reuse resource materials and digital human images. Only simple arrangement is required to complete course production in a short time, so as to improve the efficiency of creating courses.

[0041] Please refer to Figure 1 the flow schematic diagram of the course arrangement method provided by the embodiments of the present application shown below; the implementation manner of the above course arrangement method may include:

[0042] Step S110: In response to the arrangement operations of the user on the visual interface, arrange the course resource materials selected by the user to obtain multiple course scenes.

[0043] Please refer to Figure 2Schematic diagram of the arrangement operation of the user provided by the embodiment of the present application on the visualization interface; the visualization interface is a graphical user interface (Graphic User Interface, GUI) of an arrangement tool for providing arrangement operations, and the visualization interface may include a menu bar on the left and an editing area on the right. Among them, the specific course to be arranged can be selected above the menu bar (after the course list pops up), and the course resource materials and / or artificial intelligence (Artificial Intelligence, AI) controls to be dragged into the editing area can be selected below the menu bar. Since there are many types and functions of AI controls, their types and functions will be introduced in detail below. The user can preview the spatial layout of each course resource material above the editing area, and can also edit attributes such as the playback time and / or duration of each course resource material on the timeline below the editing area, and can also edit the processing method of each course resource material by the AI control, as well as the interaction conditions when the scene switches and the jump scene corresponding to the interaction conditions, etc.

[0044] Course resource materials refer to the resource materials used to arrange into course scenes, which may include: files such as videos, audios, pictures, and animations. In addition, for the developers of the arrangement tool, the above-mentioned resource materials may also include custom scripts (not shown in the figure), and the programming language of the custom scripts here can be the development language of the arrangement tool. Among them, the above-mentioned arrangement tool can be a WEB service accessed through a browser, that is, the user can operate the arrangement tool in the browser to complete the course production process; in addition, it can also be an application program (APP) running on a mobile terminal (such as a tablet computer or a smart phone), or an executable program running on a personal computer.

[0045] Step S120: Generate a course intelligent agent according to multiple course scenes and a pre-produced digital human image, and the course intelligent agent is used to interact with the teaching object in a multi-modal manner to realize teaching.

[0046] The digital human image (Digital Human Avatar) refers to a three-dimensional virtual digital human image representation, which can be a teacher digital human image pre-produced by the arrangement tool developer, or a digital human image generated by the teacher user through a deep learning model according to the uploaded multiple pictures, or a digital human image output by the deep learning model when the student user (or teaching object) inputs multiple favorite cartoon pictures or anime pictures to the deep learning model. Among them, the above-mentioned deep learning model can be a pre-trained multi-modal model such as Sora, or an AI product model of WorldLabs for generating a 3D world.

[0047] The Course Intelligent Agent (CIA) is an educational assistance intelligent agent that integrates artificial intelligence (AI) technology in the field of course teaching, aiming to improve teaching quality and learning efficiency through intelligent means. The Course Intelligent Agent can provide customized learning paths and support based on students' learning behaviors, interactive feedback, and personalized needs.

[0048] It can be understood that compared with traditional videos that embed the image of a real teacher into the video, both the digital human image and the course resource materials in the above-mentioned Course Intelligent Agent can be reused. Moreover, the Course Intelligent Agent embeds the digital human image into the course resource materials and does not bind the digital human image and the course resource materials. Therefore, the teacher digital human image in the Course Intelligent Agent can be replaced with other digital human images at any time while keeping the course content in the intelligent agent unchanged, so that a large number of personalized Course Intelligent Agents can be quickly generated to meet the preferences of different students, and at the same time, the efficiency of making courses is effectively improved.

[0049] In the implementation process of the above solution, by responding to the user's arrangement operation on the visual interface, the course resource materials are arranged, and a course intelligent agent is generated according to the arranged course scenario and the pre-made digital human image. The course intelligent agent is an educational assistance intelligent agent that integrates advanced artificial intelligence (AI) technology in the field of course teaching, so it can interact with the teaching object in multiple modalities to achieve teaching. Since these resource materials and digital human images are separated and easy to be freely combined, and the digital human images that are not real people can be well reused without repeatedly recording videos, the dependence on real lecturers is effectively reduced. Only by simply performing arrangement operations on the visual interface, the course production can be completed in a short time, greatly shortening the course production cycle and improving the efficiency of making courses.

[0050] As an optional implementation manner of the above step S110, in the process of arranging the course resource materials selected by the user, artificial intelligence AI controls can be used to arrange the course resource materials. This implementation manner may include:

[0051] Step S111: Use artificial intelligence AI controls to process the course resource materials and arrange the processed resource materials.

[0052] For example, the implementation of the above step S111 is as follows: First, import files such as videos, audios, pictures, and animations related to the course as course resource materials. Then, select the course resource materials from the left menu bar and drag the selected course resource materials into the editing area on the right. Finally, select a specific AI control from the left menu bar and use the AI control to process the course resource materials dragged into the editing area to obtain processed resource materials, and then the processed resource materials can be further arranged.

[0053] And / or, the implementation of arranging the course resource materials selected by the user may include:

[0054] Step S112: Arrange the course resource materials into a course scenario and set an artificial intelligence AI control in the course scenario. The AI control is used to interact with the teaching object to obtain an interaction result.

[0055] The above arrangement of the course resource materials into a course scenario may be to first create a blank scenario, then place the course resource materials in the blank scenario, and arrange the course resource materials in the scenario. After the arrangement in the scenario is completed, the arranged course scenario can be obtained.

[0056] It can be understood that the above interaction result may include two cases: The first case is the processing result of the data of the teaching object during the interaction. For example, if the student is the teaching object and generates voice data when answering a question, then the AI control can recognize the voice data as text data, and this part will be introduced in detail below. The second case is the data output by the AI control itself. For example, some AI controls can output text data, and this part will also be introduced in detail below.

[0057] In the implementation process of the above solution, by using the AI control to automatically process and arrange the course resource materials, the need for manual intervention is reduced, and the conversion process from the original materials to the finished course is greatly accelerated. Since most of the course resource materials can be handed over to the AI control for automatic processing, the course content can be updated quickly according to the latest research results or market demands without having to re-record the entire video, which greatly improves the flexibility and response speed. Further, the artificial intelligence AI control set in the course scenario supports functions such as speech recognition and natural language processing, enabling the virtual instructor to have a nearly real conversation with the students, enhancing the interactivity and immersion. Therefore, it can bring a new teaching interaction method to educational courses. This teaching interaction method is more natural, vivid, and interesting, thus enhancing the interactivity in the teaching process.

[0058] As an alternative implementation of the above step S111, the above AI control can include: a Text To Speech (TTS) module; the implementation of using the artificial intelligence AI control to process the course resource materials can include:

[0059] Step S111a: Use the TTS module to perform natural speech synthesis on the text files in the course resource materials.

[0060] An implementation example of the above step S111a is as follows: After setting the opening audio and video 1 in Scene 1 of the course scenario, assuming the content of the text file in the course resource materials is "Dear students! Are you ready?", a TTS module can be set in the course scenario, and the text file in the course resource materials can be dragged onto the just-set TTS module, or "Dear students! Are you ready?" can be entered in the content attribute of the TTS module, so that the TTS module can perform supernatural speech synthesis on the text file or the input content to obtain an audio file, thereby enabling the teacher to no longer produce course resource materials by recording, effectively saving the time for the teacher to produce audio file materials and improving the efficiency of course production.

[0061] Optionally, in order to make the speech synthesis function of the TTS module more flexible, voice categories (such as bass, baritone, tenor, alto, mezzo-soprano or soprano, etc.), emotional states (such as high-spirited, deep, passionate or optimistic, etc.), and pauses (1 second, 2 seconds or 3 seconds) can also be set in the TTS module. In addition, after setting the TTS module, the teacher user can also perform real-time listening back and modification and debugging on the TTS module, etc., so that the synthesized audio file can meet the satisfaction of the teacher user.

[0062] And / or, as an alternative implementation of the above step S112, the implementation of setting the artificial intelligence AI control in the course scenario can include:

[0063] Step S112a: Set a TTS module in the course scenario, and the TTS module is used to synthesize the text content generated in the interaction result into voice data.

[0064] For example, in the implementation of the above step S112a: after setting up the TTS module in the course scenario, during subsequent interactions, the TTS module can synthesize the text content generated in the interaction results into voice data. Specifically, after the child asks the question "What is a rabbit?", the question content can be input into the pre-trained large language model, and the pre-trained large language model is allowed to answer and output the text data corresponding to the question content. Then, the TTS module is used to synthesize the text data corresponding to the question content into voice data, so as to play the voice data for the child to answer the question raised by the child. Among them, the above pre-trained large language model can be ChatGPT, Alpaca, ChatGLM, Claude, Llama, etc.

[0065] In the implementation process of the above solution, through high-quality TTS technology, the virtual instructor can communicate with students in a natural and fluent language like a real person, making the learning process more vivid and interesting. This natural language expression not only improves the students' participation but also enhances their immersion. Using the TTS module to synthesize text files into voice data greatly reduces the need for manual audio recording, which means that the voice recording work that originally required a large amount of time and human resources can be automatically completed in a short time, thus significantly improving the efficiency of course production.

[0066] As an alternative implementation of the above step S111, the above AI control can include: an Automatic Speech Recognition (ASR) module; the implementation of using the artificial intelligence AI control to process the course resource materials can include:

[0067] Step S111b: Use the ASR module to perform speech recognition on the audio file in the course resource materials to obtain the recognized ASR text.

[0068] For example, in the implementation of the above step S111b: assume that the assignment for an English course is to practice listening, and the course resource materials are the audio files of a certain English radio station. Since there is no corresponding newspaper or manuscript for this English radio station, at this time, the ASR module can be used to perform speech recognition on the audio files of the English radio station to obtain the recognized ASR text.

[0069] And / or, the implementation of using the artificial intelligence AI control to process the course resource materials can include:

[0070] Step S111c: Use the ASR module to perform speech recognition on the voice data in the video file in the course resource materials to obtain the recognized ASR text.

[0071] For example, the implementation of step S111c is as follows: Assume that the assignment for an English course is to practice listening, and the course resource material is a video file of a certain obscure movie. Since no corresponding subtitle text is found for this obscure movie video, the ASR module can be used at this time to perform speech recognition on the speech data in the video file of the obscure movie video to obtain the recognized ASR text. Converting the audio file and the speech data in the video into text through the ASR module not only facilitates text processing and retrieval, but also provides subtitle support for students with hearing impairments, which greatly improves the accessibility of educational resources.

[0072] And / or, as an alternative implementation of step S112, the implementation of setting the artificial intelligence AI control in the course scenario described above may include:

[0073] Step S112b: Set an ASR module in the course scenario. The ASR module is used to identify the interaction content from the speech data of the interaction result and determine whether to perform a scenario switch based on the interaction content.

[0074] For example, the implementation of step S112b is as follows: After setting the ASR module in scenario 1 of the course scenario, during subsequent interactions, the ASR module can identify the interaction content from the speech data of the child's reading. For example, if the child is asked to learn the pronunciation of "rabbit", then the interaction content identified by the ASR module containing the character "rabbit" can be used as the interaction condition for the first option (Option1). When the interaction content identified by the ASR module meets the interaction condition of the first option, it can jump to scenario 2. Here, scenario 2 is the jump scenario, which can be specifically set to prompt correct or give rewards, etc. Similarly, if the child is asked to learn the pronunciation of "rabbit", then the interaction content identified by the ASR module not containing the character "rabbit" can be used as the interaction condition for the second option (Option2). When the interaction content identified by the ASR module meets the interaction condition of the second option, it can jump to scenario 3, and scenario 3 can be set to encourage another try or play a video related to rabbits, etc.

[0075] As an alternative implementation of step S111, the above AI control may include: a Natural Language Understanding (NLU) module; the implementation of using the artificial intelligence AI control to process the course resource material may include:

[0076] Step S111d: Use the NLU module to perform natural language understanding on the ASR text or the text file or audio file in the course resource material.

[0077] For example, the implementation of the above step S111d is as follows: Assume that the assignment for an English course is to practice reading comprehension, and the course resource material is a video file of an obscure movie. Since no corresponding subtitle text can be found for this obscure movie video, the ASR module can be used to perform speech recognition on the speech data in the video file of the obscure movie to obtain the recognized ASR text. Then, the NLU module can be used to perform natural language understanding on the ASR text file to obtain the understanding result of this obscure movie video. If the course resource material is a text file of a popular movie video, then the NLU module can be used to perform natural language understanding on the text file of this popular movie video to obtain the understanding result of this popular movie video; among them, the above understanding result can be a story summary, a central idea, a description of a character's psychology, and so on. Optionally, the above ASR module can also be used to perform natural language understanding on the audio file in the course resource material. For example, using this ASR module to perform natural language understanding on the speech data in an obscure movie video or a popular movie video, or using this ASR module to perform natural language understanding on the audio file of a certain song.

[0078] And / or, as an optional implementation manner of the above step S112, the implementation manner of setting the artificial intelligence AI control in the course scenario may include:

[0079] Step S112c: Set an NLU module in the course scenario. The NLU module is used to perform semantic understanding on the text content or speech data generated in the interaction result, and determine whether to perform a scenario switch according to the result of the semantic understanding.

[0080] For example, in the implementation of the above step S112c: after setting up the NLU module in the course scenario, during subsequent interactions, the NLU module can semantically understand the text content generated in the interaction result. For example, asking the child to tell a story with the theme of "Happy Rabbit". After the course agent obtains the child's voice data, it can first use the ASR module to recognize the text content from the voice data of the interaction result, and then use the NLU module to semantically understand the text content. The interaction condition at this time can be that the text content contains "rabbit", and the "rabbit" in the text content is happy. If the interaction condition is met, it can jump to the correct prompt or reward scenario 2. The above NLU module can be a pre-trained large language model. For example, according to the specific content of the above text content and interaction conditions, a prompt text is generated, and the prompt text is input into the pre-trained large language model to let the pre-trained large language model determine whether the text content meets the interaction conditions. If the interaction conditions are met, it jumps to the correct prompt or reward scenario 2. Among them, the above pre-trained large language model can be ChatGPT, Alpaca, ChatGLM, Claude, Llama, etc.

[0081] Optionally, after the course agent obtains the child's voice data, it can also directly use the NLU module to semantically understand the voice data and determine whether to switch scenarios according to the result of the semantic understanding. It can be understood that in the specific implementation of the above NLU module, the voice data can be first recognized as text content (for example, using the ASR module), and then semantic information is extracted from the text content for understanding. For example, using a pre-trained large language model to extract semantic information from the text content.

[0082] As an alternative implementation of the above step S110, after arranging the course resource materials into a course scenario, it may further include:

[0083] Step S112d: Set the interaction conditions corresponding to the AI control and the jump scenario corresponding to the interaction conditions in the course scenario. The interaction conditions are used to switch to the jump scenario corresponding to the interaction conditions after matching the interaction results obtained by the AI control.

[0084] For example, in the implementation of step S112d above: after setting the ASR module in scene 1 of the course scenario, the interaction conditions for the first option (Option1) corresponding to the ASR module can also be set in scene 1. For example, including the character "rabbit" in the interaction content recognized by the ASR module is set as the interaction condition for the first option (Option1), and scene 2 with correct prompts or rewards can also be set as the jump scene. Similarly, the interaction conditions for the second option (Option2) corresponding to the ASR module can be set in scene 1. For example, the interaction content recognized by the ASR module not including the character "rabbit" is set as the interaction condition for the second option (Option2), and scene 3 that encourages another try or plays a video related to rabbits can also be set as the jump scene.

[0085] As an alternative implementation of the above step S110, the implementation of arranging the course resource materials selected by the user may include:

[0086] Step S113: Edit the playing time and / or duration of the course resource materials.

[0087] For example, in the implementation of step S113 above: assuming the above course resource material is an opening audio, the playing time of the opening audio can be set at the 0th second or the 1st second of scene 1. Specifically, the opening audio can be dragged to the 0th second or the 1st second of scene 1, or the playing start time value of the opening audio can be set at the 0th second or the 1st second by setting the attributes of the opening audio. The time axis of the scene can be displayed on the visual interface to facilitate the user to determine the playing time of the resource material. Similarly, by dragging or setting the attributes, the duration of the opening audio can be edited to 10 seconds or 15 seconds, or cut to 8 seconds, etc.

[0088] In the implementation process of the above solution, by precisely controlling the playing time and duration, users can better manage the attention concentration period, avoid information overload or omission of important knowledge points caused by overly long or short content, and users can clip or extend certain segments according to actual needs to ensure that each part is closely related to the current learning objective, thereby improving the overall learning effect.

[0089] As an alternative implementation of the above step S110, the above course resource materials may include: video files; the implementation of arranging the course resource materials selected by the user may include:

[0090] Step S115: After the video file is imported, create a blank scene and set the time axis of the blank scene to the time axis of the video file.

[0091] Optionally, after the video file is imported into the choreography tool, the content of the video file can be audited first by the background server of the choreography tool, and only after the audit is passed can it be used in the choreography tool. Of course, in the specific practice process, the video file can also not be audited.

[0092] For example, the implementation manner of the above step S115 is as follows: after the video file is imported, a blank scene is created, and the time axis of the blank scene is set to the time axis of the video file. It can be understood that creating a blank scene and setting its time axis to the time axis of the video file simplifies the import and choreography process of video materials, enabling users to avoid manually adjusting the time axis and greatly saving the time for preliminary preparation work.

[0093] Step S116: Use an AI control to identify natural language text from the video file.

[0094] For example, the implementation manner of the above step S116 is as follows: use the automatic speech recognition (ASR) module in the artificial intelligence AI control to identify ASR text from the speech data in the video file. The choreography tool can also use the optical character recognition (OCR) module in the AI control to identify OCR text from the video image frames in the video file. The choreography tool can also use the natural language understanding (NLU) module in the AI control to proofread the content of the ASR text and the OCR text. After the content proofreading of the ASR text and the OCR text passes (for example, the similarity value between the ASR text and the OCR text is greater than a preset similarity threshold), determine the natural language text according to the ASR text and the OCR text. Specifically, it can be to merge the ASR text and the OCR text into the natural language text, or to determine the text with the highest confidence in the ASR text and the OCR text as the natural language text. By proofreading the content of the ASR text and the OCR text through the NLU module, a more accurate and coherent natural language text is obtained. This dual verification mechanism reduces the error rate and ensures the quality of the final output.

[0095] Step S117: Mark the content summary and / or interaction opportunity on the time axis of the video file according to the natural language text. The interaction opportunity is used to prompt the teacher user to set interaction conditions and the corresponding jump scenes at the marked positions.

[0096] For example, the implementation of the above step S117 can be as follows: The above NLU module can be a pre-trained large language model. The orchestration tool can generate prompt engineering text based on the natural language text and preset prompts, and input the prompt engineering text into the pre-trained large language model to make the pre-trained large language model return a content summary and / or interaction opportunities. After receiving the content summary and / or interaction opportunities returned by the NLU module, the orchestration tool can automatically mark the prompt information of the content summary and / or interaction opportunities on the timeline of the video file, so that the teacher user can set interaction conditions and the corresponding jump scenarios at the marked positions according to the prompt information of the content summary and / or interaction opportunities. These interaction opportunities prompt the user to set AI controls, interaction conditions, and corresponding jump scenarios, making the learning process more interactive and targeted. For example, when the student answers correctly, the system can automatically jump to more complex content; when encountering difficulties, additional help or review materials are provided. Marking the content summary and interaction opportunities on the timeline of the video file according to the natural language text provides students with a clear learning path and participation points, and these markings can be dynamically adjusted according to the student's progress and interests to provide personalized learning guidance.

[0097] As an alternative implementation of the above step S120, before generating the course intelligent agent based on multiple course scenarios and pre-made digital human images, it may further include:

[0098] Step S121: In response to a debugging request for the current scenario, debug and modify the current scenario.

[0099] For example, the implementation of the above step S121 can be as follows: After the teacher user sets up the TTS module, the TTS module can be listened to in real time and repeatedly modified and debugged. Specifically, the electronic device responds to a debugging request for a certain TTS module in the current scenario, and plays the synthesized audio effects in multiple different emotional states (such as high-spirited, low-spirited, passionate, or optimistic, etc.), so that the teacher user can debug and modify the current scenario, that is, select one of the multiple different emotional states that is most suitable for the current course scenario, thereby completing the debugging request for the current scenario. The debugging function allows developers to carefully check and optimize the interaction logic, timeline settings, AI control configurations, etc. in each scenario to ensure that all elements work as expected, improving the stability and reliability of the final product.

[0100] And / or, before generating the course intelligent agent based on multiple course scenarios and pre-made digital human images, it may further include:

[0101] Step S122: In response to a preview request for the current scenario, determine the preview effect of the current scenario according to all the resource materials in the current scenario.

[0102] For example, the implementation of the above step S122 is as follows: After the teacher user arranges the current scene or edits an AI control, if they want to view the preview effect of the current scene or the effect of the AI control in the current scene, they can click the "Preview" button on the arrangement tool to initiate a preview request for the current scene. In response to the preview request for the current scene, the arrangement tool determines the preview effect of the current scene based on all the resource materials in the current scene. The preview function allows the teacher user to conduct multiple reviews and improvements to ensure that the content is not only technically correct but also meets high standards in terms of educational value and user experience. Through the preview function, developers can simulate the complete user interaction process in the actual environment and identify potential problem points in advance, such as unreasonable interface design and unnatural speech synthesis, so as to carry out targeted optimization.

[0103] As an alternative implementation of the above step S120, after generating the course intelligent agent according to multiple course scenes and pre-made digital human images, it further includes:

[0104] Step S123: In response to the replacement operation of the digital human image, replace the digital human image in the course intelligent agent with the new digital human image selected by the user.

[0105] For example, the implementation of the above step S123 is as follows: In response to the replacement operation of the digital human image, the arrangement tool obtains the new digital human image selected by the user (for example, the user selects and uploads a new digital human image in the local file system, or the arrangement tool obtains multiple candidate images from the server, and the user selects a new digital human image from the multiple candidate images). Then, the arrangement tool replaces the digital human image in the course intelligent agent with the new digital human image selected by the user. By replacing the digital human image in the course intelligent agent with the new digital human image selected by the user, the user can select different digital human images according to personal preferences or specific teaching needs, such as different genders and clothing styles, making the virtual lecturer closer to the students' backgrounds and preferences. Further, it can respond to the user's replacement request in real time and quickly complete the replacement of the digital human image without regenerating the entire course intelligent agent, which not only improves the operation efficiency but also enhances the system's response ability.

[0106] As an alternative implementation of the above course scheduling method, the above course agent is also used to collect the learning behavior habit data of the teaching object during the multi-modal interaction process. The learning behavior habit data is used to optimize the course agent. The above teaching object can be a teenager. Since the course agent is provided with an AI control, the course agent can have a dialogue interaction with the teenager in multi-modal ways such as voice, click, selection, and text input through the AI control, so as to collect the learning behavior habit data during the dialogue interaction. These learning behavior habit data can be stored in the database of the course agent first, and then the course agent can refine and train from the database, so as to better optimize the course agent. Or the course agent can also output recommendation information, so as to optimize the course agent according to the recommendation information, so as to teach and push more personalized content to teenagers.

[0107] Please refer to Figure 3 the structural schematic diagram of the course scheduling device provided by the embodiment of the present application shown; The embodiment of the present application provides a course scheduling device 200, including:

[0108] A course scenario obtaining module 210, configured to respond to an arrangement operation of a user on a visualization interface, arrange the course resource materials selected by the user, and obtain a plurality of course scenarios.

[0109] A course agent generation module 220, configured to generate a course agent according to a plurality of course scenarios and a pre-made digital human image, and the course agent is used to perform multi-modal interaction with a teaching object to achieve teaching.

[0110] As an alternative implementation of the above device, the course scenario obtaining module includes:

[0111] A course resource arrangement sub-module, configured to process the course resource materials using an artificial intelligence AI control, and arrange the processed resource materials.

[0112] And / or, an AI control setting sub-module, configured to arrange the course resource materials into a course scenario, and set an artificial intelligence AI control in the course scenario, and the AI control is used to interact with a teaching object to obtain an interaction result.

[0113] As an alternative implementation of the above device, the AI control includes: a text-to-speech TTS module; The resource material arrangement sub-module includes:

[0114] A natural speech synthesis unit, configured to perform natural speech synthesis on the text file in the course resource materials using the TTS module.

[0115] And / or, the AI control setting sub-module includes:

[0116] A TTS module setting unit is used to set the TTS module in the course scenario. The TTS module is used to synthesize the text content generated in the interaction result into voice data.

[0117] As an alternative implementation of the above device, the AI control includes: an automatic speech recognition ASR module; a resource material arrangement sub-module, including:

[0118] A material audio recognition unit is used to perform speech recognition on the audio file in the course resource material using the ASR module to obtain the recognized ASR text.

[0119] And / or, a video audio recognition unit is used to perform speech recognition on the voice data in the video file in the course resource material using the ASR module to obtain the recognized ASR text.

[0120] And / or, the AI control setting sub-module includes:

[0121] An ASR module setting unit is used to set the ASR module in the course scenario. The ASR module is used to recognize the interaction content from the voice data of the interaction result and determine whether to perform a scenario switch according to the interaction content.

[0122] As an alternative implementation of the above device, the AI control further includes: a natural language understanding NLU module; a resource material arrangement sub-module, including:

[0123] A natural language understanding unit is used to perform natural language understanding on the ASR text or the text file or audio file in the course resource material using the NLU module.

[0124] And / or, the AI control setting sub-module includes:

[0125] An NLU module setting unit is used to set the NLU module in the course scenario. The NLU module is used to perform semantic understanding on the text content or voice data generated in the interaction result and determine whether to perform a scenario switch according to the result of the semantic understanding.

[0126] As an alternative implementation of the above device, the course arrangement device further includes:

[0127] An AI interaction setting module is used to set the interaction conditions corresponding to the AI control and the jump scenario corresponding to the interaction conditions in the course scenario. The interaction conditions are used to switch to the jump scenario corresponding to the interaction conditions after matching the interaction result obtained by the AI control.

[0128] As an alternative implementation of the above device, the course scenario obtaining module includes:

[0129] The resource material editing submodule is used to edit the playing time and / or duration of the course resource material.

[0130] As an optional implementation of the above device, the course resource material includes: a video file; a course scene acquisition module, including:

[0131] The scene timeline setting submodule is used to create a blank scene after the video file is imported, and set the timeline of the blank scene as the timeline of the video file.

[0132] The text recognition submodule is used to recognize natural language text from video files using AI controls.

[0133] The interaction opportunity marking submodule is used to mark the content overview and / or interaction opportunities in the timeline of the video file based on the natural language text. The interaction opportunities are used to prompt the teacher user to set the interaction conditions and the jump scenes corresponding to the interaction conditions at the marked locations.

[0134] As an optional implementation of the above device, the course arrangement device further includes:

[0135] The scene debugging and modification module is used to debug and modify the current scene in response to a debugging request for the current scene.

[0136] And / or, a current scene preview module is used to respond to a preview request for the current scene and determine a preview effect of the current scene according to all resource materials in the current scene.

[0137] As an optional implementation of the above device, the course arrangement device further includes:

[0138] The digital human image replacement module is used to replace the digital human image in the course intelligent body with a new digital human image selected by the user in response to the digital human image replacement operation.

[0139] As an optional implementation of the above-mentioned device, the course agent is also used to collect learning behavior habit data of the teaching objects in the multimodal interaction process, and the learning behavior habit data is used to optimize the course agent.

[0140] It should be understood that the device corresponds to the above-mentioned course arrangement method embodiment and can execute the various steps involved in the above-mentioned method embodiment. The specific functions of the device can be found in the above description, and the detailed description is appropriately omitted here. The device includes at least one software function module that can be stored in a memory in the form of software or firmware or fixed in the operating system (OS) of the device.

[0141] See also Figure 4Schematic structural diagram of the electronic device provided by the embodiments of the present application. An electronic device 300 provided by the embodiments of the present application includes: a processor 310 and a memory 320. The memory 320 stores machine-readable instructions executable by the processor 310. When the machine-readable instructions are executed by the processor 310, the above method is executed.

[0142] The embodiments of the present application also provide a computer-readable storage medium 330. A computer program is stored on the computer-readable storage medium 330. When the computer program is run by the processor 310, the above method is executed. Among them, the computer-readable storage medium 330 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM for short), electrically erasable programmable read-only memory (EEPROM for short), erasable programmable read-only memory (EPROM for short), programmable read-only memory (PROM for short), read-only memory (ROM for short), magnetic memory, flash memory, magnetic disk or optical disc.

[0143] The embodiments of the present application also provide a computer program product, including: a computer program or computer instructions. When the computer program or computer instructions are run by the processor, the method described above is executed.

[0144] It should be noted that the various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same and similar parts between the various embodiments, reference can be made to each other. For device embodiments, since they are basically similar to method embodiments, the description is relatively simple. For related parts, reference can be made to the partial description of the method embodiments.

[0145] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of devices, methods, and computer program products according to multiple embodiments of the embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may also occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, which mainly depends on the functions involved.

[0146] In addition, the various functional modules in the embodiments of the present application may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part. Furthermore, in the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0147] The above description is only an alternative implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the embodiments of the present application, and all should be covered by the protection scope of the embodiments of the present application.

Claims

1. A course arrangement method, characterized in that: include: In response to the arrangement operation of the user on the visual interface, the course resource materials selected by the user are arranged to obtain multiple course scenes; A course agent is generated according to the multiple course scenes and the pre-made digital human images, and the course agent is used to perform multi-modal interaction with the teaching object to realize teaching.

2. The method according to claim 1, characterized in that: The arranging of the course resource materials selected by the user includes: Using artificial intelligence AI controls to process the course resource materials, and arranging the processed resource materials; and / or, The course resource materials are arranged into course scenes, and artificial intelligence AI controls are set in the course scenes. The AI ​​controls are used to interact with the teaching objects to obtain interaction results.

3. The method according to claim 2, characterized in that The AI ​​control includes: a text-to-speech TTS module; the use of the artificial intelligence AI control to process the course resource material includes: Using the TTS module to perform natural speech synthesis on the text files in the course resource materials; And / or, setting an artificial intelligence AI control in the course scenario, including: The TTS module is set in the course scene, and the TTS module is used to synthesize the text content generated in the interaction result into voice data.

4. The method according to claim 2, characterized in that: The AI ​​control includes: an automatic speech recognition ASR module; the use of the artificial intelligence AI control to process the course resource material includes: Using the ASR module to perform speech recognition on the audio file in the course resource material to obtain recognized ASR text; and / or, using the ASR module to perform speech recognition on the speech data in the video file in the course resource material to obtain recognized ASR text; And / or, setting an artificial intelligence AI control in the course scene includes: The ASR module is set in the course scene, and the ASR module is used to identify the interaction content from the voice data of the interaction result, and determine whether to switch the scene according to the interaction content.

5. The method according to claim 4, characterized in that The AI ​​control also includes: a natural language understanding NLU module; the use of the artificial intelligence AI control to process the course resource material also includes: Use the NLU module to perform natural language understanding on the ASR text or the text file or audio file in the course resource material; And / or, setting an artificial intelligence AI control in the course scene includes: The NLU module is set in the course scene, and the NLU module is used to perform semantic understanding on the text content or voice data generated in the interaction result, and determine whether to switch the scene according to the result of the semantic understanding.

6. The method according to claim 2, characterized in that After arranging the course resource materials into course scenes, the method further includes: The interactive conditions corresponding to the AI ​​control and the jump scenes corresponding to the interactive conditions are set in the course scene, and the interactive conditions are used to switch to the jump scenes corresponding to the interactive conditions after matching the interactive results obtained by the AI ​​control.

7. The method according to claim 1, characterized in that The arranging of the course resource materials selected by the user includes: Edit the playing time and / or duration of the course resource material.

8. The method according to claim 1, characterized in that The course resource material includes: a video file; the arrangement of the course resource material selected by the user includes: After the video file is imported, a blank scene is created, and the timeline of the blank scene is set as the timeline of the video file; recognizing natural language text from the video file using the AI ​​control; A content overview and / or an interaction opportunity is marked in the timeline of the video file according to the natural language text, and the interaction opportunity is used to prompt the user to set an interaction condition and a jump scene corresponding to the interaction condition at the marked location.

9. The method according to any one of claims 1 to 8, characterized in that: Before generating a course intelligent body according to the plurality of course scenes and the pre-made digital human images, the method further includes: In response to a debugging request for a current scene, performing debugging modification on the current scene; And / or, in response to a preview request for a current scene, determining a preview effect of the current scene according to all resource materials in the current scene.

10. The method according to any one of claims 1 to 8, characterized in that: After generating the course intelligent body according to the plurality of course scenes and the pre-made digital human images, the method further includes: In response to the digital human image replacement operation, the digital human image in the course agent is replaced with the new digital human image selected by the user.

11. The method according to any one of claims 1 to 8, characterized in that: The course agent is also used to collect the learning behavior habit data of the teaching object during the multimodal interaction process, and the learning behavior habit data is used to optimize the course agent.

12. A course arrangement device, characterized in that: include: A course scene acquisition module is used to respond to the user's arrangement operation on the visual interface, arrange the course resource materials selected by the user, and obtain multiple course scenes; The course agent generation module is used to generate a course agent based on the multiple course scenes and pre-made digital human images. The course agent is used to perform multi-modal interaction with the teaching object to achieve teaching.

13. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and the machine-readable instructions are executed by the processor to perform any method according to claims 1 to 11.

14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is executed.

15. A computer program product, characterized in that include: A computer program or a computer instruction, wherein when the computer program or the computer instruction is executed by a processor, the method according to any one of claims 1 to 11 is executed.

Citation Information

Cited By

  • Visual construction method and device for interactive AI class based on structured data

    CN121116265A

  • Interactive AI class execution method based on user operation perception, terminal equipment and service system

    CN121116308A