Video processing methods, apparatus, electronic devices and storage media
By showcasing the first action-based instructional video and controls in action-based teaching, recording and analyzing the matching degree between the learned actions and the instructional actions, and providing real-time guidance, the problem of lacking accurate feedback in action-based teaching is solved, thereby improving learning efficiency and user experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-12-02
- Publication Date
- 2026-06-02
AI Technical Summary
Existing AI tools cannot provide accurate and real-time personalized feedback in motion-based teaching. Learners need to manually play video tutorials in slow motion or pay for real teachers to provide motion guidance and correction, resulting in low learning efficiency.
By displaying a first action-based instructional video and a first control, the system responds to the control by recording learning actions during playback, while simultaneously playing the instructional and learning videos. Computer vision technology and artificial intelligence are used to analyze the matching degree between the learning actions and the instructional actions, providing real-time guidance.
It allows learners to record actions while watching videos, enabling them to learn by direct comparison, improving learning efficiency, reducing the effort required to select instructional videos, and enhancing the user experience.
Smart Images

Figure CN122137984A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology in teaching, and more specifically, to a video processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] In teaching settings, when learners encounter unfamiliar knowledge, they manually type questions into a browser or use AI-powered dialogue tools like ChatGPT to obtain answers. While existing AI tools can provide personalized feedback based on user voice input, these methods are limited to language-based teaching and cannot provide accurate and real-time personalized feedback for action-based learning.
[0003] In action-based video tutorials (such as guitar, musical instruments, dance, etc.), learners can only manually play the video tutorials in slow speed to roughly perceive their own level of movement completion, and may even need to pay to enroll in classes to seek help from real teachers for movement guidance and correction. Summary of the Invention
[0004] This application provides a video processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which can solve the above-mentioned problems of the prior art. The technical solution is as follows: According to one aspect of the embodiments of this application, a video processing method is provided, the method comprising: Show the first tutorial video for the first action class and the first control; In response to the first control being triggered, real-time learning actions are recorded during the playback of the first action-type teaching video, so that the first action-type teaching video and the real-time learning video can be played simultaneously. The first action-related instructional video is either a first intelligent instructional video or a first intelligent instructional exclusive video; the first intelligent instructional exclusive video is a video generated based on a second intelligent instructional video and historical learning videos. For any one of the first intelligent teaching video and the second intelligent teaching video, the intelligent teaching video is generated by splicing video segments from at least one second action-type teaching video.
[0005] According to another aspect of the embodiments of this application, a video processing apparatus is provided, the apparatus comprising: The video display module is used to display the first action-based tutorial video and the first control. The synchronous playback module, in response to the first control being triggered, records real-time learning actions while playing the first action-type teaching video, so as to play the first action-type teaching video and the real-time learning video simultaneously. The first action-related instructional video is either a first intelligent instructional video or a first intelligent instructional exclusive video; the first intelligent instructional exclusive video is a video generated based on a second intelligent instructional video and historical learning videos. For any one of the first intelligent teaching video and the second intelligent teaching video, the intelligent teaching video is generated by splicing video segments from at least one second action-type teaching video.
[0006] According to another aspect of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the video processing method described above.
[0007] According to another aspect of the embodiments of this application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the above-described video processing method.
[0008] According to one aspect of the embodiments of this application, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the video processing method described above.
[0009] The beneficial effects of the technical solutions provided in this application are: By displaying a first action-type instructional video and a first control, the first control is used to trigger video recording and provide guidance on actions in the recorded video based on the first action-type instructional video. In response to the first control being triggered, the learner's real-time learning actions are recorded while the first action-type instructional video is playing. Simultaneously, the first action-type instructional video and the real-time learning video are played, allowing learners to view and perceive whether their learning actions match the instructional actions. This embodiment of the application enables learners to record actions while watching videos, allowing for more intuitive comparison and learning. Furthermore, compared to simply displaying public domain instructional videos to learners, the first action-type instructional video displayed in this embodiment is a first intelligent instructional video or a first intelligent instructional exclusive video. The video is generated based on the second intelligent teaching video and historical learning videos. In other words, the first intelligent teaching exclusive video contains not only the content of the second intelligent teaching video, but also the learner's own learning information when studying the second intelligent teaching video. It can quickly recall the learner's memory and help the learner consolidate their knowledge. For any one of the first and second intelligent teaching videos, the intelligent teaching video is generated by splicing video segments from at least one second action-type teaching video. This allows for a wider collection of basic materials for generating the first action-type teaching video, enabling the presentation of the content of multiple teaching videos to the learner at the same time. This significantly reduces the effort learners need to select teaching videos, significantly improves learning efficiency, and enhances the user experience. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.
[0011] Figure 1A A schematic diagram of the interface of language teaching software provided for related technologies; Figure 1B A schematic diagram of a browser-based interface for teaching action-related technologies; Figure 1C A schematic diagram of the interface for teaching movement based on professional movement teaching software, provided for related technologies; Figure 1D A schematic diagram illustrating the implementation environment of the video processing method provided in the embodiments of this application; Figure 2A A flowchart illustrating a video processing method provided in an embodiment of this application; Figure 2B A schematic diagram of an interface for a video processing method provided in an embodiment of this application; Figure 3A schematic diagram of a synthetic intelligent teaching video provided in an embodiment of this application; Figure 4 A schematic diagram of an intelligent practice interface provided in an embodiment of this application; Figure 5 An operation flow diagram for displaying a first action-type instructional video is provided as an embodiment of this application; Figure 6 A schematic diagram of an intelligent practice interface provided in an embodiment of this application; Figure 7 A schematic diagram of the interface for recommending third-action class teaching videos based on overall matching degree, provided in an embodiment of this application; Figure 8A A schematic diagram of a video frame in the intelligent teaching-specific video provided in the embodiments of this application; Figure 8B A schematic diagram illustrating the process of generating video frames in a smart teaching-specific video provided in an embodiment of this application; Figure 8C This is a schematic diagram of the interface for displaying checkpoint indication information provided in an embodiment of this application; Figure 8D A schematic diagram of an interface displaying over-point indication information is provided for another embodiment of this application; Figure 9A A schematic diagram of the interface for triggering the generation of intelligent teaching-specific videos provided in an embodiment of this application; Figure 9B A schematic diagram of the interface of My Smart Course provided for an embodiment of this application; Figure 9C A schematic diagram of the search results interface provided in an embodiment of this application; Figure 9D A schematic diagram of the browser discovery interface provided in this application embodiment; Figure 10A A schematic diagram illustrating the video processing method provided in an embodiment of this application; Figure 10B This is a schematic diagram of the technical framework of the video processing method provided in the embodiments of this application; Figure 11A This is an interactive schematic diagram of the video processing method provided in the embodiments of this application; Figure 11B An interactive schematic diagram of real-time motion guidance provided in an embodiment of this application; Figure 11C An interactive schematic diagram illustrating the generation of intelligent teaching-specific videos provided in an embodiment of this application; Figure 12 This is a schematic diagram of the structure of a video processing apparatus provided in an embodiment of this application; Figure 13This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.
[0013] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.”
[0014] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0015] First, let's introduce and explain several terms used in this application: Motion capture: This refers to using a camera to capture motion videos in real time, then sending the footage to an AI backend to analyze the user's current motion accuracy and provide intelligent teaching feedback (such as real-time adjustment and guidance of the current motion, commenting and scoring, and recommending related videos).
[0016] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in recognizing and measuring targets, and then performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as facial recognition and fingerprint recognition.
[0017] Intelligent teaching videos refer to videos with high viewing value that are analyzed, refined, and synthesized using artificial intelligence based on learners' consumption data of video results (viewing time, viewing frequency, interaction, likes, favorites, etc.) in search scenarios.
[0018] Intelligent teaching videos: Based on intelligent teaching videos, motion capture technology is used to identify the user's perfect actions during the learning process and intelligently integrate the video frames of that moment into the corresponding video frames of the original intelligent teaching video to generate dynamic teaching materials with the learner's own unique mark, just like paper teaching materials with the learner's own hand-drawn marks in reality. This can be used more effectively for future review and reinforcement of knowledge.
[0019] Search results page: This refers to the page that retrieves a list of video results when a learner initiates a search by entering the text "query" into a search engine in their browser.
[0020] Intelligent practice interface: refers to an interface that simultaneously plays instructional videos, learning videos, and action guidance information.
[0021] The Smart Course Accumulation Page is an aggregated page that presents all generated smart teaching videos in a video list format.
[0022] The relevant video teaching solutions include two types: language-based and action-based. For language-based teaching solutions, please refer to... Figure 1A Learners can interact with the AI teacher in real time by sending voice messages (1101) and receive evaluations (1102) and optimization feedback (1103) (mainly feedback on the logic of language expression, such as providing more appropriate responses). However, it cannot provide more precise information on pronunciation correction (such as the coordination of organs such as the tongue, teeth, and throat).
[0023] In action-based teaching programs, there are two main approaches: Please see Figure 1B The example shows a schematic diagram of a browser-based interface for action-based teaching. As shown in the figure, the learner initiates a search for guitar lessons in the browser, and the browser displays relevant teaching videos found through a preset search engine. The learner clicks on the teaching video, and the browser starts playing the teaching video. At the same time, the learner practices in the real world by referring to the teaching video. Please see Figure 1C The example shown illustrates the interface of a guitar practice software. The software displays course content (three courses are shown: legato exercises, rock rhythms, and solo techniques). Learners can choose one course to practice. The tutorial content is presented as animated sheet music, but when learners fall behind in rhythm during practice, they need to manually adjust the playback speed. While professional practice tools offer a better experience than existing browser solutions, they still lack precise corrective feedback and a truly intelligent user experience.
[0024] The video processing methods, apparatus, electronic devices, computer-readable storage media, and computer program products provided in this application aim to address the problem that existing teaching solutions, whether language-based or action-based, cannot provide accurate problem points and optimized feedback from a physical perspective. In particular, in teaching scenarios with strong physical movements, existing technologies fail to meet the demands for real-time teaching adjustments, increasing the learning difficulty for learners.
[0025] The technical solutions of this application and their effects are described below through several exemplary embodiments. It should be noted that the following embodiments can be referenced, borrowed from, or combined with each other. Identical terms, similar features, and similar implementation steps in different embodiments will not be repeated.
[0026] Figure 1D This is a schematic diagram illustrating the implementation environment of the video processing method provided in this application embodiment. The implementation environment may include a server 11 and a terminal 12. A wired or wireless communication connection is established between the server 11 and the terminal 12. Optionally, the server 11 may be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The terminal 12 may be a personal computer (PC), in-vehicle terminal, tablet computer, smartphone, wearable device, intelligent robot, or other terminal with data computing, processing, and storage capabilities.
[0027] In this embodiment, the terminal 12 in the system can be used to obtain learner interaction data on the teaching video and send the interaction data to the server 11. The server 11 can analyze and refine the interaction data based on an artificial intelligence model, synthesize an intelligent teaching video containing video clips of interest to the learner, and return the intelligent teaching video to the terminal 12, so that the terminal 12 can display the intelligent teaching video and the first control. In response to the first control being triggered, the terminal 12 records real-time learning actions during the playback of the intelligent teaching video and sends the learning video to the server 11. The server 11 analyzes the matching degree between each learning action in the learning video and the actions in the intelligent teaching video through an artificial intelligence model, thereby providing action guidance information to help users correct their actions more accurately and efficiently and improve learning efficiency.
[0028] Optionally, the terminal 12 may also store an artificial intelligence model. After acquiring the learner's interaction data with the teaching video, the terminal 12 can directly input the interaction data into the artificial intelligence model. The artificial intelligence model analyzes and refines the interaction data to synthesize an intelligent teaching video containing video clips of interest to the learner. The terminal 12 can display the intelligent teaching video and a first control. In response to the triggering of the first control, the terminal 12 records real-time learning actions while playing the intelligent teaching video, and analyzes the matching degree between each learning action in the learning video and the actions in the intelligent teaching video through the artificial intelligence model, thereby providing action guidance information to help users correct their actions more accurately and efficiently, and improve learning efficiency. Correspondingly, the system may not include the server 11.
[0029] Based on the above-described terminology and application scenarios, the video processing method provided in this application embodiment will be described. This method can be applied to a computer device, which can be... Figure 1D Terminal 12 in the scene shown.
[0030] This application provides a video processing method that can be applied to browsers or professional teaching software, such as... Figure 2A As shown, the method includes: S2101. Display the first action class tutorial video and the first control.
[0031] It is understood that the first action-related teaching video in this application embodiment is a video related to action-related teaching. The teaching items for action-related teaching can be physical action items—such as dance, gymnastics, etc., or finger action items—such as playing guitar, flute, etc., or mouth action items—such as English speaking, beatboxing, etc. This application embodiment does not make specific limitations. For ease of description, subsequent embodiments of this application will be described using guitar fingering teaching as an example.
[0032] When this embodiment of the application is applied to a browser, before step S2101, in some embodiments, since browsers all have search functions, in response to the input of search information related to action-related teaching in the search box provided by the browser, multiple action-related teaching videos found by the search engine are displayed. The learner further selects one action-related teaching video from the multiple action-related teaching videos to play, like, or favorite, and the operated action-related teaching video is designated as the first action-related teaching video. In other embodiments, the browser can provide an information flow interface displaying an information stream. The information in this information stream can be determined based on the learner's interests. Therefore, if the learner has previously searched for or browsed action-related teaching videos, at least one action-related teaching video can be displayed in the information flow interface. When the learner selects an action-related teaching video in the information flow interface to play, the selected action-related teaching video is designated as the first action-related teaching video.
[0033] When the embodiments of this application are applied to professional teaching software, the action-type teaching video that learners click to play on the software can be used as the first action-type teaching video. Considering that professional learning software may also provide a search function, the action-type teaching video that learners select to play, like, or favorite from the search results can be used as the first action-type teaching video.
[0034] Considering that most current chat software has the function of publishing and browsing multimedia information (such as Moments), the embodiments of this application can also be applied to chat software. When a learner publishes a video of a certain action or browses a video of an action published by a friend through chat software, by collecting the learner's interaction data, relevant action teaching videos can be obtained from the Internet based on the interaction data to generate a first action teaching video. By pushing the first action teaching video to the learner, the learner can browse the first action teaching video without actively searching for teaching videos, which can stimulate the learner's learning interest.
[0035] This application embodiment shows a first action-type teaching video, which may be displaying the cover of the first action-type teaching video (the first action-type teaching video is not playing), that is, displaying the cover of the first action-type teaching video and the first control; or it may be playing the first action-type teaching video, that is, displaying the first control while playing the first action-type teaching video. This application embodiment does not make specific limitations.
[0036] The first control in this application embodiment is used to trigger video recording and to provide guidance on the actions in the recorded video based on the first action class teaching video.
[0037] The first action-based instructional video in this application embodiment can be a first intelligent instructional video. The first intelligent instructional video is generated by splicing video segments from at least one second action-based instructional video. In some embodiments, when learners browse a series of action-based instructional videos, some videos may only focus on one problem, while others systematically explain several steps—even professional instructional software often exhibits this characteristic. This leads to learners needing to spend a significant amount of time searching and filtering to find videos that match their learning stage, ability, and needs in existing action-based instructional scenarios, which greatly affects their learning enthusiasm and efficiency.
[0038] Therefore, the first intelligent teaching video presented in this application is not simply based on teaching videos found by a search engine, but is a composite of video clips from one or more videos. It embodies the culmination of teaching content from multiple videos and can significantly reduce the effort learners spend searching for teaching videos.
[0039] In some embodiments, the second action-type teaching video of this application includes at least one of an interactive action-type teaching video and a teaching video related to the interactive action-type teaching video. Since the second action-type teaching video originates from the video that the learner has interacted with, it is more in line with the learner's current learning needs.
[0040] In some embodiments, the video clips of this application include at least one of a segment of interest and a segment for illustrating the first target action. That is, the video clips of this application correspond to two dimensions: one is the segment of interest to the learner, which can record the learner's interaction data with each video in the background (including likes, favorites, playback counts of each frame in the video, etc.) and transmit this interaction data to artificial intelligence for filtering to obtain the learner's segment of interest for each video. It can be understood that each segment of interest is composed of multiple video frames, and by orderly combining these segments of interest, the first action-type teaching video can be obtained; the other is the segment illustrating the first target action. The first target action in this application is a key action determined based on historical learning videos. That is, the key action can be either a learning action shown in historical learning videos or a guiding action corresponding to a learning action.
[0041] The key actions in the embodiments of this application can be determined mainly according to the evaluation method of Functional Movement Screen (FMS). Specifically, representative actions that can reflect the learning theme and learning process can be selected from various actions: The learning theme can be the theme of the action-based instruction video corresponding to the learning video, or it can be the theme of the learning action itself. For example, if the theme of the action-based instruction video corresponding to the learning video is "6 commonly used chords", then the actions of these 6 chords can be used as key actions. Or, for example, if a learner performs many actions while learning the F chord fingering, the most standard action can be used as the key action.
[0042] Specifically, in this embodiment, computer vision technology and natural language processing technology are used to perform semantic analysis on the teaching video corresponding to the learning video to determine the theme of the teaching video. Pose estimation algorithms (OpenPose, PoseNet, AlphaPose, etc.) are used to identify the actions in the learning video. The identified actions are iterated to see if they match the teaching theme, and the matching actions are taken as key actions. In some embodiments, pose estimation algorithms can also be used to identify the actions in the teaching video and the learning video, and the action in the learning video that best matches the action described in the teaching video can be taken as the key action. In addition, the identified actions in the teaching video can also be directly taken as key actions.
[0043] Actions that embody the learning process usually refer to actions that learners practice repeatedly. For example, a recorded learning video may involve five actions. Most of these actions are mastered relatively smoothly by the learner, but one action is practiced repeatedly for a long time. Therefore, this action can be regarded as the key action. Of course, actions that embody the learning process are not limited to those that have been mastered. Actions that have not yet been mastered and need to be learned or overcome (also known as "bottlenecks") can also be used to embody the learning process.
[0044] Since the recording of the learning videos is accompanied by the simultaneous playback of the first action-type instructional video, and learners often repeatedly play corresponding segments of the first action-type instructional video during practice, this application embodiment can statistically analyze the video segments repeatedly played by learners and perform action recognition on these segments. The identified actions are then designated as key actions. Similarly, actions in the unplayed portions of the instructional video are considered actions to be learned. In some embodiments, key points can be identified in the actions within the video. By comparing the degree of matching between the key points of the learning actions and the instructional actions, it can be determined whether the learner has mastered the corresponding actions. Actions identified as not mastered in the repeatedly played video segments are then designated as actions to be tackled.
[0045] The embodiments of this application do not specifically limit the way the first target action is described. For example, it can be a direct demonstration of the correct action, or it can be described through text, voice, or drawing.
[0046] In some embodiments, the first action-type teaching video in this application can also be a first intelligent teaching-specific video. The first intelligent teaching-specific video is generated from a second intelligent teaching video and historical learning videos. It is understood that the second intelligent teaching video can be the same as the first intelligent teaching video or a different video. When the second intelligent teaching video and the first intelligent teaching video are the same, it means that before the current display of the first intelligent teaching-specific video, the first intelligent teaching video has already been shown to the learner and the learner's learning video has been recorded. This is equivalent to providing the learner with an opportunity to review their previous learning of the first intelligent teaching video.
[0047] The present application embodiment provides a solution for generating intelligent teaching-specific videos based on a second intelligent teaching video and historical learning videos. This solution can be to merge the second intelligent teaching video with historical learning videos to synchronously display the guided actions in the second intelligent teaching video and the learner's own learning actions; alternatively, it can be to extract the learner's learning process from the learning video, obtain summary information, and add it to the intelligent teaching video as an intelligent teaching-specific video; or it can be a combination of the above two solutions. The present application embodiment does not impose any specific limitations.
[0048] The summary information of the embodiments of this application may include at least one of the following: It indicates to learners the key moments in the learning process, such as highlights, bottlenecks, and breakthroughs. The learner's learning action at a critical moment and the corresponding instruction action for the learning action; The duration of practice for at least one learning action (which may be limited to key actions); The number of times at least one learning action (which may be limited to key actions) is learned; and At least one learning action (which may be limited to key actions) requires motion instruction information. It should be understood that the generation logic of the second intelligent teaching video in this application embodiment is consistent with the generation logic of the first intelligent teaching video, and will not be repeated here.
[0049] S2102. In response to the first control being triggered, record real-time learning actions while playing the first action-type teaching video, so as to play the first action-type teaching video and the real-time learning video simultaneously.
[0050] This application does not specifically limit the way the first control is triggered, such as single click, double click, long press, sliding a preset distance, etc.
[0051] In some embodiments, if step S2101 shows that the first action-type teaching video is not playing, then after the first control is triggered, it will record real-time learning actions when the first action-type teaching video starts playing. Furthermore, in this application embodiment, a separate playback control can be set for the first action-type teaching video, so that the first action-type teaching video is played after the learner triggers a preset operation on the playback control. In some embodiments, the first control can also be used to play the first action-type teaching video at the same time, that is, when the first control is triggered, the first action-type teaching video starts playing and real-time learning actions start recording.
[0052] In some embodiments, if step S2101, displaying the first action-type teaching video, refers to playing the first action-type teaching video, then in response to the triggering of the first control, the terminal will immediately begin recording the learner's learning actions and simultaneously play the recorded learning video and the first action-type teaching video. This helps learners to intuitively view and compare the teaching actions in the first action-type teaching video with the real-time learning actions in the learning video, thereby improving learning efficiency. For ease of description, the interface for playing the first action-type teaching video and the learning video will be referred to as the intelligent practice interface in the following embodiments of this application.
[0053] Generally, when learners use mobile phones, tablets, or other terminals for action-based learning, to simultaneously play the first action-based instructional video and the learning video, the learner's actions can be recorded by turning on the terminal's front-facing camera. Considering that the terminal's rear-facing camera usually produces better results than the front-facing camera, in practical applications, learners can also turn on the rear-facing camera and record their actions based on the principle of specular reflection. Since this method is widely used in current live streaming scenarios, it will not be described in detail in this application's embodiments.
[0054] It should be noted that in this embodiment of the application, when simultaneously playing the first action-type instructional video and the learning video, the size of the two videos in the display area is not limited. For example, the first action-type instructional video can be larger than the learning video, or smaller than the learning video, or each video can occupy half of the display area. In some embodiments, the first action-type instructional video can be on the same layer as the learning video, thus the two videos are separated by a dividing line. In other embodiments, the first action-type instructional video and the learning video can be on different layers, for example, the first action-type instructional video is on the bottom layer and the learning video is on the top layer, or the first action-type instructional video is on the top layer and the learning video is on the bottom layer. It is understood that since the video on the top layer will obscure the video on the bottom layer to some extent, this embodiment of the application can set the video on the top layer to a relatively small size and display it in a corner of the display area. For example, if the display size of the video on the bottom layer is the size of the terminal display screen, then the display size (including horizontal and vertical) of the video on the top layer can be 1 / 3 of the size of the terminal display screen.
[0055] The beneficial effects of the technical solution provided in this application embodiment are as follows: By displaying a first action-type teaching video and a first control, the first control is used to trigger video recording and guide the actions in the recorded video based on the first action-type teaching video. In response to the first control being triggered, the learner's real-time learning actions are recorded during the playback of the first action-type teaching video. On the one hand, the first action-type teaching video and the real-time learning video are played simultaneously, allowing the learner to watch in real time and perceive whether their learning actions match the teaching actions. This application embodiment enables learners to record actions while watching videos, which allows for more intuitive comparison and learning. Furthermore, compared to simply showing learners public domain teaching videos, the first action-type teaching video displayed in this application embodiment is a first intelligent teaching video or a first intelligent teaching exclusive video. The first intelligent teaching exclusive video is a video generated by merging a second intelligent teaching video and historical learning videos. That is to say, the first intelligent teaching exclusive video includes both the content of the second intelligent teaching video and the content of the learner's own learning video when studying the second intelligent teaching video, which can quickly recall the learner's memory and help the learner consolidate their knowledge.
[0056] In some embodiments, for any one of the first and second intelligent teaching videos, the intelligent teaching video is generated by splicing together segments of interest from at least one second action-type teaching video. The second action-type teaching video may include action-type videos in which the learner has previously interacted, and may also include teaching videos related to the action-type teaching videos in which interactions have occurred, thereby collecting a wider range of basic materials for generating the first action-type teaching video. In this application embodiment, the teaching videos related to the action-type teaching videos in which interactions have occurred may include teaching videos with the same or similar teaching themes as the action-type teaching videos in which interactions have occurred, or teaching videos whose teaching themes are chronologically related to the action-type teaching videos in which interactions have occurred, or other action-type teaching videos published by the publisher of the action-type teaching videos in which interactions have occurred. This application embodiment does not specifically limit the scope of these videos.
[0057] Furthermore, by extracting video clips from various second-action category instructional videos and splicing them together, first-action category instructional videos can be obtained. These video clips can include segments of interest or key and difficult movements discovered during the recording of learners' learning videos in the past. This allows for the presentation of the essence of multiple instructional videos to learners at once, significantly reducing the effort learners need to select instructional videos, significantly improving learning efficiency, and enhancing the user experience.
[0058] Based on the above embodiments, as an optional embodiment, after the first control is triggered, this application embodiment will also display action guidance information for the learning video in real time. Since the learner's learning actions are recorded in real time, this application embodiment can use artificial intelligence technology to compare the real-time recorded learning actions with the teaching actions in the first action teaching video, and provide action guidance information by analyzing the differences between the learning actions and the teaching actions.
[0059] The action guidance information provided in this application embodiment can be displayed in any one or a combination of voice, text, images and animation, and this application embodiment does not make specific limitations.
[0060] Please see Figure 2B The figure illustrates an exemplary interface diagram of the video processing method provided in this application embodiment. Figure 2BThe interface 2210 on the left side displays the first action-based tutorial video 2201 and the first control 2202. At this moment, the first action-based tutorial video 2201 is demonstrating a specific finger position, while the first control 2202 prompts the learner to "Enable AI motion correction." To facilitate understanding, a functional overview of the first control 2202 is provided—try watching the video while simultaneously receiving real-time motion analysis. Learners will then understand that when the first control 2202 is triggered, they can simultaneously watch the first action-based tutorial video 2201 and also see their own real-time learning motion. Figure 2B The interface on the right is the intelligent practice interface 2220 after the second control is triggered. It can be seen that most of the area of the intelligent practice interface 2220 is used to display the learning video 2203, while the first action-based instructional video 2201, which was originally displayed in full screen, is now shown in the upper left corner of the interface, allowing learners to more clearly see their finger placement. Figure 2B As can be seen, when the first action-type instructional video and the learning video are played simultaneously, the first action-type instructional video is located on the upper layer and is superimposed on the learning video located on the lower layer. Figure 2B The document also displays action guidance information 2204 for the learning movements. It can be seen that the action guidance information is in the form of text, such as "put your thumb back and bend your index finger". Through the real-time display of action guidance information, learners can adjust their learning movements in a timely manner, thereby completing the learning more efficiently.
[0061] This application embodiment displays real-time action guidance information for learning videos, allowing learners to adjust their movements promptly through intuitive action guidance. Compared to related technologies that only provide instructional videos or offer feedback only after the learner completes the action, this application embodiment enables learners to record their actions while watching the video and receive real-time corrective feedback.
[0062] For example, in the search scenario of "how to practice left-hand dexterity on guitar", please refer to [link / reference]. Figure 3 On the search results page, the first video that learners actively viewed multiple times was about hand gestures (Topic A), the second was about basic methods of graphing (Topic B), the third was about advanced methods of musical notation (Topic C), the fourth was a combination of A and B, and the fifth was a combination of A and C. Users found the A parts of the fourth and fifth videos particularly well-explained and interesting, repeatedly dragging the progress bar to view the corresponding content. By using artificial intelligence to identify learners' preferences for different parts of these videos, the first video (A), the second video (B), the third video (C), the fourth video (A), and the fifth video (A) were refined and combined using big data analysis to create an ordered information arrangement. This resulted in an intelligent teaching video that integrates artificial intelligence and user filtering, significantly improving the efficiency of users revisiting their preferred historical content.
[0063] The second action-based instructional video in this application embodiment may include action-based instructional videos in which interactive operations have occurred. This application embodiment does not limit the specific type of interactive operation; for example, it may be a play operation, a favorite operation, a like operation, a positive comment operation, etc. A positive comment operation refers to a comment containing positive words, such as "This video is great," "Very valuable for learning," "Worth sharing with more people," etc. In some embodiments, the second action-based instructional video in this application embodiment may also include instructional videos with the same teaching theme as the action-based instructional videos in which interactive operations have occurred. For example, if a learner has watched one or more videos on the theme of "left-hand fingering techniques," then videos on that theme that the learner has not yet watched can also be included as second action-based instructional videos, thereby achieving a more comprehensive acquisition of materials for generating the first action-based instructional video.
[0064] Based on the above embodiments, as an optional embodiment, the action guidance information is used to instruct at least one of the following: The learning videos demonstrate the methods for correcting learning actions; Feature information of the learning actions shown in the learning video; The characteristic information of the teaching actions shown in the first action-type teaching video.
[0065] The correction method for the learning actions shown in the learning video described in this application embodiment can be referred to... Figure 2B The movement guidance information 2204 in the text helps learners adjust their posture and movement patterns through instruction and correction methods, so as to match the teaching movements or reduce the deviation between them and the teaching movements.
[0066] The characteristic information of an action (learning action and teaching action) is information that can basically or obviously reflect the characteristics of the action. In the embodiments of this application, the characteristic information of the action can be represented by one or more combinations of text, pictures, and animation.
[0067] Taking guitar fingering instruction as an example, the characteristic information of the movements is represented by lines connecting various key points of the fingers. Please refer to [link / reference needed]. Figure 4The illustration shows a schematic diagram of the intelligent practice interface provided in the embodiment of this application. As shown in the figure, the interface simultaneously plays a first action-type teaching video and a learning video. However, it should be noted that the interface guides learners in two ways. At the bottom of the interface, the text "Put your thumb back, bend your index finger, and make sure the other fingers are in good form" displays a type of action guidance information 3101. At the same time, the multiple fingers shown in the first action-type teaching video and the learning video are visually displayed with the connection lines 3102 on the key points of the fingers, so that learners can more intuitively adjust their gestures by comparing the shape of the connection lines 3102 on each finger in the two video frames.
[0068] Based on the above embodiments, as an optional embodiment, the segment of interest in this application may include a video segment in the corresponding second action-type teaching video whose number of replays exceeds the first count threshold. This application does not specifically limit the size of the first count threshold; for example, it may be 2 times.
[0069] In some embodiments, the segments of interest may further include segments of interest predicted from the corresponding second action class instructional video based on an interest recognition model.
[0070] In some embodiments, the interest recognition model is a model trained in a supervised or unsupervised manner on sample segments of interest, used to identify the learner’s level of interest in at least one of the themes, styles, and content of instructional videos.
[0071] The training samples for the interest recognition model in this application embodiment are sample teaching videos including at least one video segment. The annotation information of the training samples is the video segments of interest in the sample teaching video or the degree of interest of each video segment. It should be understood that when the annotation information of the training samples is the video segments of interest in the sample teaching video, the trained interest recognition model has the ability to input a teaching video and output the segments of interest in the teaching video. Thus, after inputting a second action-type teaching video into the interest recognition model, the interest recognition model can directly output the segments of interest. When the annotation information is the degree of interest of each video segment, the trained interest recognition model has the ability to input a teaching video and output the degree of interest of each video segment in the teaching video. Thus, after inputting a second action-type teaching video into the interest recognition model, the interest recognition model can output the degree of interest of each video segment in the second action-type teaching video. Furthermore, this application embodiment can, based on a preset interest threshold, identify video segments with a degree of interest higher than the interest threshold as segments of interest.
[0072] This application embodiment can obtain the learner's second action-type teaching video by browsing history, then analyze the frame sequence of the interest segments browsed by the learner through data analysis, and then use OpenCV, NLP, AI-driven video editing algorithms to splice multiple interest segments to achieve video fusion, so as to obtain intelligent teaching video and help learners find satisfactory video tutorial results more quickly.
[0073] In some embodiments, after splicing the various video segments, the present application can further use Automatic Speech Recognition (ASR) technology to optimize the audio effect, thereby obtaining a smart teaching video with lower noise and higher speech quality. In addition to optimizing the audio effect, another core aspect of ASR technology is its ability to extract acoustic features from the speech in the video, find the most likely word sequence, thereby recognizing the text content, and performing post-processing operations such as error correction and formatting on the recognized text to obtain more accurate text content. This text content can be used as subtitles to supplement the first action-based teaching video, which is more helpful for learners to understand the content of the video.
[0074] Based on the above embodiments, the embodiments of this application can also utilize computer vision technology to optimize video effects, such as adjusting the brightness and contrast of the video, performing color correction on the video, etc. Considering that the resolution of different video segments may be inconsistent, and the allocation rate of some segments of interest may be low, the embodiments of this application can utilize computer vision technology, especially intelligent super-resolution technology, to efficiently improve the video resolution and unify the resolution of all video segments, thereby making the video viewing effect more natural.
[0075] In some embodiments, considering the differences in content between different video segments, in order to play different video segments more naturally and smoothly, the embodiments of this application can also use computer vision calculations to perform intelligent frame interpolation, which not only improves the video frame rate, but also makes the picture smoother.
[0076] This application embodiment uses ASR and computer vision technology to optimize audio and video effects, generating high-quality intelligent teaching videos, which can effectively reduce the learning cost for learners and improve learner satisfaction.
[0077] Based on the above embodiments, as an optional embodiment, this application embodiment demonstrates a first action-type teaching video and a first control, including: In response to the first search operation for action-related teaching, if there is no intelligent teaching video or intelligent teaching exclusive video related to the first search operation, the first search result is displayed, which includes multiple action-related teaching videos. In response to at least one action-related instructional video in the first search result being interacted with, a second control is displayed; In response to the second control being triggered, the first action-type tutorial video and the first control are displayed.
[0078] This application provides a search function. Learners can enter the information they want to search for through the search box. When searching for information related to action-based teaching, a preset search engine will be invoked. It will first check if any related intelligent teaching videos or dedicated intelligent teaching videos have been generated previously. If no results are found, it means that no intelligent teaching videos or dedicated intelligent teaching videos related to this action-based teaching have been generated or shown to the learner before. Then, search results will be obtained from the internet. It is understood that the search results include multiple action-based teaching videos. When at least one action-based teaching video in the search results is played, a second control will be displayed. In this application embodiment, the second control is used to display the first action-based teaching video and the second control. By triggering the second control, the learner can display an intelligent teaching video spliced from video clips of the played action-based teaching videos. It is understood that since this is the first time an intelligent teaching video related to the first search operation is being shown to the learner, and the learner's learning process has not yet been recorded, there is naturally no dedicated intelligent teaching video related to the first search operation. Therefore, the first action-based teaching video displayed at this time is the first intelligent teaching video, not the first dedicated intelligent teaching video.
[0079] In some embodiments, a first playback threshold can be set for displaying the second control. That is, the second control will only be displayed after the number of playbacks of the action-type teaching video reaches the first threshold. The purpose of this operation is to obtain a first type of action-type teaching video that is more accurate and matches the user's current learning ability based on more playback records of the user.
[0080] It should be understood that due to the existence of the first count threshold, the number of times a learner plays videos in a single search may not be enough to display the second control. For example, if the first count threshold for "guitar lessons" is set to 5 times, if a learner searches for "guitar lessons" for the first time and plays 3 of the video lessons in the search results, the second control will not be displayed because the first count threshold has not yet been reached. When the learner searches for "guitar lessons" again and plays 2 of the video lessons in the search results, the second control will be displayed because the number of video lessons played at this time has reached the first count threshold.
[0081] This application does not limit the size of the first count threshold, for example, it can be 1. It should be noted that when the first count threshold is small, learners can see the intelligent teaching video immediately after each search and browsing of the teaching video, without having to wait for the next search operation.
[0082] As can be seen from the above embodiments, at least one second action-type teaching video in the embodiments of this application includes at least one action-type teaching video played in the search results, that is, the action-type teaching video played in the current search results. In some cases, it may also include action-type teaching videos played in historical search results.
[0083] Please see Figure 5 The example illustrates the operation flow interface diagram of the first action-based teaching video in this application embodiment. It should be noted that this application embodiment is based on the scenario of action-based teaching through search operation, thereby reducing the requirements of action-based teaching for the audience. The audience does not need to download professional teaching software, but can achieve action-based teaching through an ordinary browser, thus lowering the learning threshold for learners.
[0084] like Figure 5 As shown, firstly, the search interface 5110 provides a search box 5111 for learners to input the information they want to search for. To facilitate repeated searches, this embodiment also provides search history for learners to view and select. Learners can select search terms from the search history to search, thereby improving search efficiency. When a learner enters a new search term in the input box 5111: "How to fret the strings with the left hand on a guitar," and the learner clicks the confirmation control 5112, they are redirected to the search results interface 5120, which displays the search results. The search results interface 5120 displays guitar teaching videos found by the search engine from different platforms across the internet. Learners can view these videos from the interface 5111. In step 20, the learner selects any video to play. The playback interface 5130 displays one of the search results. It's understood that during video playback, the terminal records interaction data, such as playback duration, number of times each frame is played, and interaction records (likes, favorites, comments, etc.). Based on the recorded interaction data, the backend automatically generates a first action-based teaching video (intelligent teaching video). In this embodiment, the second control can be displayed on the playback interface 5130 of the currently playing teaching video after the intelligent teaching video is generated, or it can be displayed when the learner returns to the search results interface 5120 after watching a video. Figure 5 As can be seen, in this embodiment of the application, when the learner is watching the teaching video on the interface 5120, a prompt message 5131 and a second control 5132 are displayed. The learner triggers the second control 5132 to display the interface 5140. The interface 5140 simultaneously plays the spliced intelligent teaching video 5141 and the first control 5142.
[0085] Based on the above embodiments, as an optional embodiment, this application embodiment further includes: adjusting the playback speed of the first action-type teaching video according to the overall matching degree determined in real time between the learning video and the first action-type teaching video.
[0086] This application embodiment utilizes artificial intelligence technology to determine the overall matching degree between learning videos and teaching videos in real time. Based on this overall matching degree, it intelligently adjusts the playback speed of the first action-based teaching video. A higher overall matching degree indicates that the learner is currently keeping up with the teaching progress of the first action-based teaching video, and the playback speed can be appropriately increased. Conversely, a lower overall matching degree indicates that the learner is currently not keeping up with the teaching progress of the first action-based teaching video, and the playback speed can be appropriately decreased. This addresses the pain point in related technologies where learners can only learn according to a fixed playback speed of the teaching video. This application embodiment can automatically reduce the playback speed of the first action-based teaching video when the learner cannot keep up with the teaching pace, allowing for more precise adjustments by the learner. When the learner can keep up with the teaching pace, it automatically increases the playback speed of the first action-based teaching video, without wasting the learner's time, improving learning efficiency, reducing learning costs, and bringing a more intelligent experience.
[0087] Please see Figure 6 The example shown is a schematic diagram of the interface of the intelligent practice video according to an embodiment of this application, compared to Figure 2B In this embodiment, the window displaying the intelligent teaching video provides a playback rate indicator 6101. As shown in the figure, the playback rate is currently 1.5, indicating that the learner is keeping up with the teaching progress, hence the playback rate is greater than 1. Simultaneously, the interface also displays indicator 6102, reminding the learner that the video playback rate has been automatically increased because the learner is currently proficient in the learning actions. In some embodiments, the learner can also trigger a preset operation on the playback rate, further displaying a playback rate adjustment interface (not shown in the figure). The learner can also manually adjust the playback rate through this interface.
[0088] The overall matching degree in this application embodiment, as the name suggests, is the real-time matching degree between the learning video and the first action-type teaching video. Considering that both videos can be regarded as a set of actions, the overall matching degree can be understood as the matching degree between each action in the recorded learning video and each action in the played segment of the first action-type teaching video. For example, if a learner studies for one minute using the video processing method provided in this application embodiment, that is, the first action-type teaching video plays for one minute and shows three actions in sequence, and at the same time, the learning video is also recorded for one minute, recording the three actions imitated by the learner, then the two videos can be motion-captured by artificial intelligence, and the matching degree of these three pairs of actions can be compared. Furthermore, the real-time overall matching degree of the two videos can be obtained by weighted averaging, cumulative methods, etc. For example, if the matching degree of action 1 in the two videos is 99%, the matching degree of action 2 is 100%, and the matching degree of action 3 is 98%, then the real-time overall matching degree of the two videos is calculated as 99% by averaging. In some embodiments, different weights can be set for different actions, so that after obtaining the matching degree of each action, the overall matching degree can be determined by weighted average.
[0089] In addition to the above embodiments, as an optional embodiment, it further includes: In response to the overall matching degree being lower than the first matching degree threshold, at least one third action class instructional video is recommended.
[0090] The third type of instructional video in this application embodiment is a learning video with a lower learning difficulty than the first type of instructional video, or an instructional video used to teach a second target action.
[0091] The second target action in this application embodiment is the guidance action in the first action class teaching video that causes the overall matching degree to be lower than the first matching degree threshold.
[0092] In this embodiment, a first matching degree threshold can be set for the overall matching degree. When the overall matching degree is lower than the first threshold, it is considered that the learner is currently unable to keep up with the teaching progress of the first action-type teaching video. At least one third action-type teaching video is then displayed. The learning difficulty of the third action-type teaching video is lower than that of the first action-type teaching video. By learning more basic and easier courses, learners can find their own learning path.
[0093] In some embodiments, considering that instructional videos usually include audio or text narration, this application embodiment can use artificial intelligence to perform semantic recognition on the first action-type instructional video to determine the teaching stage of the first action-type instructional video. By obtaining instructional videos below the teaching stage in the cloud, the third action-type instructional video can be recommended to learners. For example, if the first action-type instructional video is the video of the 5th lesson in a guitar instructional video series, and the learner cannot keep up with the pace of the video, the video of the 3rd lesson can be used as the third action-type instructional video.
[0094] In other embodiments, instructional videos at the same teaching stage, but with more accessible explanations and lower learning difficulty, can also be obtained from the cloud as third-action instructional videos. For example, there are often multiple versions of guitar tabs for the same song, some suitable for beginners and others for experienced learners. If the first-action instructional video is an advanced version of guitar playing for experienced learners, the lower-level version of guitar playing can be recommended to learners as the third-action instructional video.
[0095] Please see Figure 7 The figure illustrates an exemplary interface diagram of a third action category instructional video recommended based on overall matching degree, as provided in an embodiment of this application. Figure 7 The intelligent practice interface 7110 on the left side shows that the playback rate of the first action-type teaching video has dropped to 0.5. Since the overall matching degree is detected to be lower than the first matching degree threshold, the playback of the first action-type teaching video is paused. At this time, the interface provides several third action-type teaching videos 7111 with lower learning difficulty. Learners can browse more third action-type teaching videos by operating the sliding control 7112. When learners click on any of the third action-type teaching videos, that third action-type teaching video will be played. It should be understood that when the third action-type teaching video is replayed, the terminal will also re-record the learner's learning actions and simultaneously play the third action-type teaching video and re-record the learning video in real time.
[0096] Based on the above embodiments, as an optional embodiment, the method further includes: In response to the overall matching degree being no less than the second matching degree threshold when the playback progress is no less than the progress threshold, a second intelligent teaching exclusive video is generated. The second intelligent teaching exclusive video is a video generated by fusing the first action-type teaching video and the real-time learning video. It should be understood that the first matching degree threshold is lower than the second matching degree threshold.
[0097] To encourage users to keep up with the teaching progress over a longer period of time, this application embodiment pre-sets a progress threshold. When a learner starts playing the first action-based teaching video, the playback progress of the first action-based teaching video is recorded. Since the overall matching degree of this application is counted in real time, when the playback progress exceeds the progress threshold, if the overall matching degree is not lower than the second matching degree threshold, a second intelligent teaching exclusive video will be generated. The second intelligent teaching exclusive video is a video generated based on the real-time learning video and the first action-based teaching video. It should be understood that the method of generating the second intelligent teaching exclusive video in this application embodiment is the same as the logic of generating the first intelligent teaching exclusive video in the above embodiment, and will not be repeated in this application embodiment.
[0098] In some embodiments, this application combines learners' own learning videos with teaching videos to synchronize the video feeds of teachers and students, thereby enhancing learners' sense of accomplishment and participation.
[0099] In this application embodiment, the size of the progress threshold is not specifically limited. For example, it can be 50% or 100%. When the progress threshold is 100%, it means that the learner's complete learning process and the first action-type teaching video are integrated.
[0100] Understandably, video fusion technology is a technique that integrates and synthesizes information from multiple cameras or video sources, with the aim of generating a coherent, seamless, and more complete video by effectively integrating multiple video streams.
[0101] Based on the above embodiments, as an optional embodiment, for any one of the first and second smart teaching exclusive videos, at least one video frame in the smart teaching exclusive video includes a first image and a second image. The first image is determined based on the first video frame in the corresponding action-type instructional video; The second image is the foreground image in the second video frame of the corresponding learning video.
[0102] In some embodiments, the first image is a keyframe in the first action-type teaching video. For example, in a video explaining the decomposed action, a static finger position may be shown through multiple video frames. In this case, it is only necessary to select one video frame from these multiple video frames. In this way, the duration of the generated smart teaching video can be equal to the duration of all keyframe video frames.
[0103] In this embodiment, the second image is a foreground image in the second video frame of the learning video, and the frame matching degree between the second video frame and the first video frame is higher than a third preset threshold. Frame matching degree refers to the matching degree between the guided action displayed in the first video frame and the learning action displayed in the second video frame; in other words, frame matching degree also refers to the degree of action completion of the learning action displayed in the second video frame relative to the guided action displayed in the first video frame. This embodiment can use artificial intelligence to capture the actions displayed in the video frames and analyze the feature information of the actions, using the similarity between the feature information of two actions as the frame matching degree.
[0104] Based on the above embodiments, as an optional embodiment, when the first action-type teaching video is a first intelligent teaching video, the first image in the currently generated second intelligent teaching exclusive video is based on the first video frame in the first intelligent teaching video.
[0105] When the first action-oriented teaching video is a first intelligent teaching-specific video, since the first intelligent teaching-specific video has already incorporated historical learning videos, if the currently generated second intelligent teaching-specific video were to further incorporate the currently recorded learning video, the viewing experience might be affected due to display size limitations. Therefore, in this embodiment, the first image in the second intelligent teaching-specific video is the first video frame in the second intelligent teaching video, that is, the historical learning video is removed, and the second intelligent teaching video and the currently recorded learning video are merged. In some embodiments, considering that large-size display devices are gradually becoming more common, the first image in the second intelligent teaching-specific video can also be the first video frame in the first intelligent teaching-specific video.
[0106] Please see Figure 8A The illustration shows a schematic diagram of a video frame in the intelligent teaching video provided in the embodiment of this application. It can be clearly seen from the figure that the video frame includes two parts - a first image 8011 of the video frame in the intelligent teaching video as the background and a second image 8012 of the video frame in the learning video as the foreground. Moreover, the size of the second image is not a regular rectangular size like that of a normal video frame, but only the foreground image of the video frame in the learning video. Since the teaching scene in the embodiment of this application is a guitar teaching scene, the foreground image is the part of the learner playing the guitar.
[0107] Accordingly, the overall matching degree in this application embodiment is obtained through the frame matching degree of each video frame pair between the first action teaching video and the learning video. The overall matching degree in this application embodiment is positively correlated with the number of video frame pairs with a frame matching degree higher than a third preset threshold. Each video frame pair includes video frames from the first action teaching video and the learning video, respectively. In other words, when measuring the overall matching degree, this application embodiment considers that in actual applications, the learner's actions and the progress of the first action teaching video cannot be completely synchronized. Therefore, the backend often does not know which frame in the teaching video to compare with the video frame of the real-time learning video to determine the most accurate match. Thus, it is necessary to compare the video frame of the real-time learning video with multiple video frames already played in the teaching video to obtain the video frame with the highest frame matching degree and the best match from the teaching video.
[0108] This application embodiment counts the number of video frame pairs with a frame matching degree higher than a third preset threshold. The higher the number, the better the learner is in sync with the teaching video; the lower the number, the worse the learner is in sync with the teaching video.
[0109] Please see Figure 8B The illustration shows a flowchart of generating a smart teaching video according to an embodiment of this application. The illustration shows two video frame sequences: a video frame sequence for a learning video and a video frame sequence for a smart teaching video. The video frame sequence for the smart teaching video includes multiple first video frames, and the video frame sequence for the learning video includes multiple second video frames. Whenever a new second video frame is recorded, it is compared with multiple first video frames that have already been played in the smart teaching video. In the illustration, a second video frame 3 has just been recorded. The second video frame 3 is compared with the three first video frames that have already been played. It is found that the frame matching degree between the first video frame 3 and the second video frame 3 reaches 100%, which is higher than the third preset threshold (98%). Therefore, the two video frames are merged into one video frame in the smart teaching video.
[0110] This application embodiment takes into account that learners go through a process from unfamiliarity to proficiency during learning. In this process, there are often some actions that learners need to practice repeatedly or that require a lot of effort to learn. Therefore, in the process of learners referring to the first type of action teaching video, for the learners' mismatched actions (that is, actions for which action guidance information is provided), this application embodiment can count the number of times action guidance information is provided. For those teaching actions whose guidance number exceeds the second threshold, they are displayed in the intelligent teaching exclusive video as checkpoint indication information. In this way, when learners watch the intelligent teaching exclusive video later, they can review those actions that they have repeatedly practiced to master, quickly recall the learners' past learning experience, and help learners review and consolidate the actions.
[0111] The intelligent teaching video in this application embodiment also includes checkpoint indication information. This checkpoint indication information is used to indicate teaching actions where the number of guidance attempts exceeds a threshold for the second attempt during the learning process. Please refer to [link to relevant documentation]. Figure 8C The exemplary embodiment is a schematic diagram of the interface for displaying checkpoint indication information provided in the embodiments of this application, compared to Figure 8A The video frame shown has added a beat indication information 8301. When this video frame is played, the learner will know that this action was a repetitive action that has been trained for a long time during the previous learning process.
[0112] Figure 8D The following is a schematic diagram of the interface for displaying checkpoint indication information provided in another embodiment of this application. As shown in the figure, when playing the intelligent teaching exclusive video, for video frames used to demonstrate teaching actions that have exceeded the second threshold, checkpoint indication information 8402 is displayed on the progress bar 8401. Learners can click on the checkpoint indication information 8402 to directly jump to the corresponding video frame for playback, so that learners can quickly browse the actions that were repeatedly trained and improve learning efficiency.
[0113] Generally, if a learner has not mastered a movement in a first-action instruction video, they will often control the playback progress of the video and repeatedly play the video clips related to that movement. Each time the video is played, if the learner does not perform the movement correctly during practice, the movement instruction information for that movement will be displayed. Therefore, in this embodiment, the movement corresponding to the movement instruction information can be identified each time the movement instruction information is displayed, and the movement instruction information displayed for that movement can be counted. Movements whose count exceeds the second threshold are designated as key movements.
[0114] When generating a smart teaching video based on the first type of action-based instructional video and learning video, the actions displayed in each frame of the smart teaching video are identified. When key actions are identified, they can be processed using the above methods. Figure 8C and Figure 8D The display method shown presents checkpoint instructions to learners.
[0115] Based on the above embodiments, the third action-type teaching video can also be a teaching video for teaching a second target action, whereby the second target action is the guiding action in the first action-type teaching video that causes the overall matching degree to be lower than the first matching degree threshold.
[0116] Since the overall matching degree is determined based on the matching degree of each action key, the guided actions that cause the overall matching degree to be lower than the first matching degree threshold must be actions that the learner has not mastered. Therefore, in this embodiment, those guided actions that cause the overall matching degree to be lower than the first matching degree threshold can also be regarded as the second target actions that the learner needs to focus on learning. The teaching videos of teaching the second target actions can be recommended to the learner to help the learner to carry out targeted training.
[0117] Based on the above embodiments, as an optional embodiment, a second intelligent teaching-specific video is generated, which further includes: Display a third control; In response to the triggering of the third control, at least one of the second intelligent teaching exclusive video and the first action-type teaching video is displayed in a preset information flow interface.
[0118] Please see Figure 9A The illustration shows an example of an interface diagram for triggering the generation of a smart teaching video according to an embodiment of this application. As shown in the figure, since the learner's proficiency in learning actions is detected to be high (playback rate is 2), a prompt message 9111 and a save control (third control) 9112 are displayed below the smart practice interface 9110. When the learner triggers a preset operation on the save control 9112, a second smart teaching video will be generated and saved. In this embodiment of the application, the second smart teaching video can be saved locally on the terminal or in the cloud. It should be understood that the save control 9112 can also be used to save the first action-type teaching video after triggering, or to save both the second smart teaching video and the first action-type teaching video after triggering. The display effect of the interface is the same as... Figure 9A Similarly, the embodiments in this application will not be described in detail.
[0119] Figure 9B The "My Smart Courses" page shown is a feed interface displaying multimedia information posted by learner accounts. When the user clicks... Figure 9AAfter saving using the control shown, the smart teaching-specific video will first be stored on the "My Smart Courses" page 9210. The "My Smart Courses" page will record the name of the smart teaching-specific video. This name can be edited by the learner when saving, or it can be automatically generated based on the theme of the smart teaching video. The page will also record the generation time of the smart teaching-specific video. It is understood that all smart teaching-specific videos generated through this application embodiment will be stored on this page, so learners can browse this page to review their learning experience, increasing learners' engagement with the browser. The smart teaching-specific videos on this page not only support sharing but will also appear as search results. That is, when a user searches for related topics in the future, if relevant past first-action teaching videos and smart teaching-specific videos are identified, they will be actively displayed on the search results page.
[0120] It should be noted that, in this embodiment of the application, when storing the smart teaching-specific video to the "My Smart Courses" page, the corresponding smart teaching video can also be stored, for example, in... Figure 9B The system stores both the intelligent teaching exclusive video 9211 and the intelligent teaching video 9212, allowing learners to watch both pure teaching videos and intelligent teaching exclusive videos that combine teaching videos with their own learning videos.
[0121] Based on the above embodiments, as an optional embodiment, a first action-type instructional video and a first control are shown, including at least one of the following: In response to a first search operation targeting action-related instruction, if a first action-related instructional video exists that is relevant to the first search operation, then the first action-related instructional video will be displayed as a priority search result. In other words, this embodiment combines the learner's search behavior with their learning outcomes. After generating a personalized intelligent instructional video for each learner, when the learner subsequently searches for teaching content related to the personalized intelligent instructional video (e.g., with the same or similar teaching themes), at least one of the first action-related instructional video and the second personalized intelligent instructional video can be displayed as a priority second search result. This allows learners to quickly retrieve previously practiced skills from their previous learning progress.
[0122] In some embodiments, when generating smart teaching videos, semantic analysis can be performed on the videos to determine the teaching theme. For example, if a teaching video involves the explanation and teaching of several common chords such as C, D, and E, the teaching theme could be "common chords & C major & D major & E major & teaching". The smart teaching videos and smart teaching videos are tagged with the teaching theme in the background. When learners subsequently search for action-related teaching methods, the keywords in the search information are matched against the tags of the saved smart teaching videos and smart teaching videos. If a matching smart teaching video or smart teaching video is found, it will be displayed as a priority search result. For example, if a learner later searches for "a complete list of common guitar chords", since smart teaching videos and smart teaching videos with the teaching theme "common chords & C major & D major & E major & teaching" have been previously stored, these two videos can be prioritized as search results.
[0123] In some embodiments, after determining the teaching theme of a video, this application generates a video title based on the teaching theme. Subsequently, when the learner's search operation for action-related teaching is obtained, the keywords in the search information are matched with the video titles of various saved smart teaching videos and smart teaching-specific videos. If a matching smart teaching video or smart teaching-specific video is found, it will be displayed as the priority search result. The information flow interface in this application embodiment can be a search results interface that displays search results. Please refer to [link to relevant documentation]. Figure 9C In this application, a learner searches for guitar-related keywords on search interface 9310. For the learner's search term "how to fret the guitar strings with the left hand," the search results retrieve a smart teaching video related to left-hand fingering that the learner had previously saved on their multimedia information sharing site. This smart teaching video is then displayed as a search result on search results interface 9320. In this embodiment, the smart teaching video is given higher priority than videos found on the internet. When displaying the smart teaching video on the search results interface, this application embodiment also displays the time the video was saved, helping learners organize their learning process and strengthen their memory of the knowledge.
[0124] Based on the above embodiments, as an optional embodiment, this application embodiment demonstrates a first action-type teaching video and a first control, including: In response to opening a preset information flow interface, if it is determined that the time elapsed since the generation of the first action-type instructional video or the time elapsed since its last display meets preset requirements, the first action-type instructional video is preferentially displayed on the information flow interface. The information flow interface in this application embodiment can be an information flow interface that displays trending information or information recommended based on interests, such as... Figure 9D As shown, this example illustrates a schematic diagram of the discovery interface provided in this application embodiment. As shown, the browser's homepage provides multiple interfaces, such as a video interface specifically for playing short videos, a discovery interface, a screening room interface specifically for playing movies and TV dramas, a game interface providing web games, etc. Browsers typically set up a "discovery" interface to recommend trending news or content of interest to users. This application embodiment formulates a knowledge review strategy based on the Ebbinghaus forgetting curve method, displaying the intelligent teaching exclusive video 9410 in a prominent position on the browser's discovery interface at a specific time, enabling learners to discover the intelligent teaching exclusive video and achieve the goal of reviewing and learning new material. It should be understood that... Figure 9D The smart teaching videos shown are just examples; you can also display smart teaching videos in practice.
[0125] Please see Figure 10A The figure illustrates a schematic diagram of the video processing method provided in this application embodiment. As shown, the method mainly includes three steps: generating intelligent teaching videos, generating action guidance information, and generating intelligent teaching-specific videos. Specifically: Generate intelligent teaching videos When learners search for information related to action guidance, the search engine's feedback may sometimes explain only one action, while other times it may systematically explain multiple actions. This application's embodiments record learners' interaction data with search results (e.g., likes, favorites, play counts, operation methods). By extracting content of interest to learners from the accumulated interaction data over a period of time and systematically piecing it together, intelligent teaching videos can be obtained. This application's embodiments prioritize recommending intelligent teaching videos to learners when they search for similar information again. Alternatively, after showing learners search results, intelligent teaching videos can be generated and displayed based on learners' interaction data with the search results, greatly improving the efficiency of learners revisiting their preferred historical content.
[0126] Generate motion guidance information In response to the first control being triggered, the camera is activated to record the learner's real-time learning movements. Simultaneously, intelligent teaching videos and real-time learning videos are played. Artificial intelligence performs real-time motion capture and analysis, providing learners with correction methods (such as finger placement, pressure, and joint bending feedback) through voice and text dialogue, and providing a performance score. The AI also recognizes the learner's current progress. If a beginner is struggling to keep up with the pace, the playback speed will be automatically reduced to help them follow along. However, if the AI detects that the learner is consistently having difficulty correcting their movements (possibly because the video is not suitable for their current level), it will proactively suggest that the learner first learn more basic and simpler course videos to help them find their own learning path.
[0127] This application embodiment captures the actions shown in the learning video, constructs an action recognition model using motion capture and deep learning methods, and combines it with a pose estimation algorithm to extract and compare the user's action completion level in real time, thereby providing real-time scoring and optimization feedback. It helps users correct their actions more accurately and efficiently and improves learning efficiency through text and voice guidance, automatic adjustment of playback speed, and recommendation of targeted videos.
[0128] Generate personalized intelligent teaching videos If a learning video with a high overall match to the learner is identified, the learner will be guided to upgrade to a personalized intelligent teaching video. This involves the user clicking the upgrade control (a third control), and the AI intelligently extracting the foreground video from the learning video (essentially showing only the learner and guitar). This foreground video is then integrated into the intelligent teaching video, creating a personalized intelligent teaching video that combines both instructional and learning elements. Furthermore, this personalized intelligent teaching video can overwrite the original intelligent teaching video and be stored on the "My Intelligent Courses" page. This content not only supports sharing but also appears as search results. When learners search for related topics in the future, if relevant past personalized intelligent teaching videos are identified, they will be automatically displayed in the search results. For action-based courses, repeated review and practice are often necessary to achieve mastery. Based on the Ebbinghaus forgetting curve, a customized review strategy is implemented, pushing intelligent teaching videos and personalized intelligent teaching videos to learners in various scenarios at specific times to reinforce learning. Therefore, the intelligent teaching videos generated and multi-scene reviews provided in this application can not only accumulate the learner's past learning costs, but also effectively recall the learner's learning memories by showing the learner's own image, allowing for continuous review and learning, which is equivalent to a high-quality electronic textbook that can be continuously updated and iterated.
[0129] This application embodiment realizes a learning experience that integrates "viewing + practice + correction + review", thereby helping learners overcome the difficulties of online self-study of action-related courses and learn practical knowledge more accurately and effectively.
[0130] Please see Figure 10B It exemplarily illustrates a schematic diagram of the technical framework of the video processing method provided in the embodiments of this application, corresponding to Figure 10A The technical framework of this application includes three stages: generating intelligent teaching videos, generating action guidance information, and generating intelligent teaching-specific videos. Each stage depends on the completion of the previous stage. The intelligent teaching videos generated in stage one are combined with the learner's real-time learning actions to generate action guidance information in stage two. The intelligent teaching-specific videos upgraded in stage three depend on the learning videos with high overall matching degree in stage two and the intelligent teaching videos in stage one. Specifically: The implementation means of stage 1 in this application mainly include data mining, AI analysis and AI synthesis. The underlying operation objects of stage 1 are the videos and interaction data that the learner has interacted with. By analyzing the videos and interaction data that the learner has interacted with, AI is used to mine the segments of interest in each video that has been interacted with. Furthermore, the AI-driven video editing algorithm is used to splice and synthesize the segments of interest to obtain intelligent teaching videos. The implementation means of stage 2 in this application embodiment mainly include real-time video recording, AI motion capture, and AI analysis. The underlying operation objects of stage 2 are the learning video of the learner's real-time learning behavior and the intelligent teaching video. When the learner triggers the first control, the recording of the learning video will be triggered. The AI will perform motion capture and analysis to determine whether the learner's learning action is correct or incorrect, and generate action guidance information based on the error type to help the learner correct specific types of errors. In addition, the playback speed of the intelligent teaching video can be intelligently adjusted and lower difficulty teaching videos can be provided according to whether the learner's learning ability matches the intelligent teaching video. The implementation of stage 3 in this application embodiment is basically achieved through AI, mainly including AI action analysis, AI image matting, and AI video fusion. The underlying operation objects of stage 3 are learning videos with high overall matching degree of learners, action guidance information, and intelligent teaching videos. Through AI, the video frames in the learning videos and intelligent teaching videos are matched to find the corresponding video frames in the intelligent teaching videos and learning videos. Furthermore, the foreground image of the video frame in the learning video is fused with the corresponding video frame in the intelligent teaching video to obtain the video frame of the intelligent teaching exclusive video. This can enhance the learner's learning interest, and the generated intelligent teaching exclusive video can be reviewed in multiple scenarios through AI. It can not only accumulate the user's past learning costs, but also effectively outline and awaken the user's learning memories at that time, and continuously review and learn new things.
[0131] Based on the above embodiments, as an optional embodiment, the first intelligent teaching video of this application embodiment is generated in the following way: Obtain at least one second-action tutorial video based on historical interaction records; For each second action category instructional video, obtain the segments of interest in the second action category instructional video based on the pre-trained interest recognition model; Based on the semantic correlation of each segment of interest, the splicing order of each segment of interest is determined, and the segments of interest are spliced together according to the splicing order to obtain the first intelligent teaching video.
[0132] This application embodiment reads learners' search behavior records and playback records by using a plugin built into the software or backend logs to obtain information such as the videos viewed, viewing duration, number of repeated views, number of likes, number of favorites, and comment content, thereby obtaining at least one second action-type instructional video. When the plugin or backend logs also record learners' operation records on the viewed videos, the segments of interest in the second action-type instructional videos can also be obtained directly based on the operation records. However, considering that learners may not have watched all the instructional videos, in order to fully explore the segments of interest in the second action-type instructional videos, this application embodiment can pre-train an interest recognition model to obtain the segments of interest in each second action-type instructional video.
[0133] In some embodiments, the interest recognition model is trained in a supervised or unsupervised manner on sample segments of interest. Specifically, the collected sample segments of interest are cleaned and preprocessed, including removing duplicate data, handling missing values, and data normalization. Then, meaningful features are extracted from the data, and the data is transformed into feature vectors that can better describe learners' interests and behaviors. Based on the extracted feature vectors, a supervised or unsupervised machine learning model (such as decision trees, support vector machines, neural networks, etc.) is trained. The machine learning model is evaluated and optimized using methods such as cross-validation and hyperparameter tuning, thereby obtaining a neural network model that can identify learners' interest in the theme, style, or content elements of a specific teaching video.
[0134] In this embodiment, for each second action-type teaching video, the second action-type teaching video is input into an interest recognition model to obtain the interest recognition model's degree of interest in each video frame or video segment in the second action-type teaching video. Furthermore, video frames or video segments with an interest degree greater than the interest threshold are regarded as segments of interest.
[0135] In some embodiments, the interest recognition model of this application can also be trained by supervised learning using multiple training samples with labeled information. The training samples are multiple sample teaching videos, and the labeled information is the segments of interest in the sample teaching videos. Thus, the interest recognition model has the ability to take a teaching video as input and output segments of interest in the teaching video.
[0136] This application embodiment can analyze the semantic information and semantic relevance of each segment of interest using AI, and use AI-driven video editing algorithms to automatically identify scene changes in different segments of interest (by analyzing visual differences between consecutive frames to detect scene switching), identify and track key objects in multiple segments of interest, and determine the splicing order (also known as script) of each segment of interest by combining the semantic correlation between segments of interest. The splicing is performed based on the splicing order, and at the same time, artificial intelligence generative adversarial networks are used to generate some new visual elements (such as images, animations, etc.) to achieve a more natural fusion and transition with each segment of interest, thereby obtaining intelligent teaching videos.
[0137] Furthermore, this application embodiment uses ASR technology to input the intelligent teaching video into a pre-trained speech recognition model (such as a Hidden Markov Model, Deep Neural Network, Convolutional Neural Network, Long Short-Term Memory Network, etc.). The speech recognition model performs feature extraction, acoustic model mapping, and decoding processing on the audio signals in the intelligent teaching video, and outputs the corresponding text sequence. The text sequence may include specific terms, important prompts, etc. Based on the script obtained above, the timing of each piece of information in the text sequence appearing in the intelligent teaching video is determined, and speech synthesis technology (such as WaveNet, Tacotron, etc.) is used to generate a more natural and coherent audio result for the intelligent teaching video.
[0138] Furthermore, embodiments of this application can use AI and machine learning technologies to optimize the generated intelligent teaching video tutorials. Specifically, reinforcement learning algorithms are used to optimize the content and structure of the video, and computer vision technology is used to optimize the visual effects of the video (such as color correction, image enhancement, etc.), ultimately resulting in more refined intelligent video courses.
[0139] Based on the above embodiments, as an optional embodiment, the action guidance information of this application embodiment is generated in the following way: For each video frame in the learning video, the video frame is classified by a pre-built action classification model to obtain a classification result. The classification result is used to indicate whether the learning action in the video frame is correct or the learning action is incorrect. The action guidance information is generated based on the classification results.
[0140] The action classification model in this application is a classification model that classifies each video frame in the learning video and determines the classification result. The classification result in this application includes a classification indicating that the learning action of the video frame is correct and a classification indicating the error type of the learning action. This application does not specifically limit the error type; for example, it can be improper thumb position, uneven plucking force, or excessive plucking amplitude, etc.
[0141] The training samples of the action classification model in this application include both positive and negative samples. Both positive and negative samples are video frames, and each carries annotation information. The annotation information of positive samples indicates that the action displayed in the corresponding video frame is correct; the annotation information of negative samples indicates the type of error in the action displayed in the corresponding video frame. It should be noted that the reason this application requires the use of negative samples, besides considering model training for classification purposes, is that in action teaching, it may be difficult to determine the accuracy of an action based solely on correct actions, but introducing incorrect actions makes it easier to determine whether an action is incorrect.
[0142] Positive samples in this application embodiment can come from keyframes in intelligent teaching videos. By extracting keyframes from the intelligent teaching video and performing pose analysis on the keyframes using a pose estimation model (such as OpenPose or PoseNet), the annotation information of the keyframes is obtained, thereby obtaining positive samples. Considering that intelligent teaching videos generally demonstrate correct actions, this application embodiment can supplement the data by crawling videos containing incorrect actions from the Internet to obtain negative samples.
[0143] The embodiments of this application can perform data preprocessing (such as denoising, normalization, etc.) on positive and negative samples before data labeling to distinguish between correct and incorrect actions, which can be used as a supervision signal for model training.
[0144] In some embodiments, key point data can be extracted from each sample and added to the annotation information, which helps to further improve the model's understanding of the training samples.
[0145] This application uses a deep learning method (a combination of convolutional neural networks and long short-term memory networks) to construct an initial model. The convolutional neural network extracts spatial features from keypoint data, while the long short-term memory network captures action patterns over time. The initial model is trained using a sample set, aiming to teach it how to judge the correctness of actions based on keypoint data. After training, an action classification model is obtained, capable of scoring, providing feedback, and categorizing actions. Taking guitar playing tutorials as an example, the action recognition model will be trained to classify error types including incorrect finger positions, inaccurate string pressing, and incorrect plucking techniques.
[0146] As learners watch and follow along with intelligent instructional videos, their movements are captured via camera. Pose estimation algorithms (OpenPose, PoseNet, AlphaPose, etc.) are used to extract the learned actions from the video stream in real time, feeding them into a trained action classification model to determine the accuracy of the learner's actions.
[0147] Based on the above embodiments, as an optional embodiment, the frame matching degree of the video frame pair is determined in the following way: Motion capture is performed on two video frames in the video frame pair to obtain the motion key points of each of the two video frames. The frame matching degree of the video frame pair is obtained based on the distance between the action key points of the two video frames.
[0148] It should be noted that the embodiments of this application can extract the position information of skeletal key points (in the context of guitar teaching, skeletal key points are used as motion key points) from video frames based on computer vision technology, thereby achieving motion capture of the human body. After obtaining the motion key points of two video frames, the distance between the motion key points of each video frame is calculated. The larger the distance, the greater the difference between the actions displayed in the two video frames, and the lower the frame matching degree of the two video frames.
[0149] This application embodiment provides real-time corrective feedback to learners based on the frame matching degree between video frames in the real-time calculated motion video and video frames in the intelligent teaching video, combined with the teaching video, including presenting text and voice guidance, automatically adjusting playback speed, and recommending targeted videos.
[0150] In some embodiments, this application can set a threshold. When the real-time accumulated frame matching degree is lower than this threshold, it is considered that the learner's current progress is not suitable for the teaching video. At this time, the action classification model will be combined to further analyze in which aspects the user has a large number of errors. Based on the proportion of the analyzed error types, the user can be helped to correct high-frequency errors of specific types. Taking guitar playing tutorials as an example, targeted tutorial recommendations such as finger position training and string pressing techniques can more effectively and accurately correct wrong movements and improve learning results.
[0151] Please see Figure 11A The figure illustrates an interactive schematic diagram of the video processing method provided in this application embodiment. As shown, the interaction in this application embodiment involves a learner, a client, and a server. Specifically: S1101. The user opens a browser, searches for information related to motion instruction in the search box, and initiates a search; S1102. The client responds to the search operation by obtaining the search request and uploading it to the server. S1103. The server extracts search information and related parameters from the search request, performs a query based on the search engine, and retrieves a series of web pages according to the preset index and ranking algorithm. S1104. The server sorts the results based on relevance and weight, and returns the search results to the client. S1105. The client retrieves the search results and presents them visually. S1106. Learners view search results and engage in interactive behaviors such as commenting on, (repeatedly) browsing, and saving videos that interest them. S1107. The client captures the learner's interactive behavior and uploads it to the server; S1108: The server uses AI to analyze interactive behavior, obtains segments that learners are interested in, and uses OpenCV, NLP, and AI-driven video editing algorithms to generate a fusion of video segments; S1109: The server optimizes audio and video effects through technologies such as Automatic Speech Recognition (ASR) and computer vision to generate high-quality intelligent teaching videos, which are then sent back to the client. S1110: The client obtains intelligent teaching videos and displays them in the search results to enhance viewing guidance; S1111 Learners see intelligent teaching videos in the search results and click to view them; S1112: The client plays intelligent teaching videos and provides real-time AI action guidance; S1113. The server will archive the played smart teaching videos to the My Smart Courses interface. S1115. The client will visualize the intelligent teaching videos in my intelligent course interface. S1116. Learners can open the My Smart Courses interface to see all the smart teaching videos generated in the past.
[0152] Please see Figure 11B The figure exemplifies an interactive diagram of real-time action guidance provided in an embodiment of this application. As shown, the interaction in this embodiment involves a learner, a client, and a server. Specifically: S1201, The learner triggers a preset operation on the first control; S1202: The client obtains the camera access command and plays the intelligent teaching video; S1203. Learners can learn by turning on the camera and referring to the intelligent teaching videos. S1204. The client collects learners' real-time learning behavior and uploads it to the server; S1205. The server acquires learning videos, uses motion capture technology and deep learning methods to build an action recognition model, and uses a pose estimation algorithm to judge the matching degree between the learned action and the guidance action in real time, and provides scoring and behavior guidance information. S1206: The client obtains and displays the data recalled by the server in real time, including presenting text and voice guidance, automatically adjusting the playback speed, and recommending targeted videos. S1207. Learners adjust their posture based on real-time motion guidance information, gradually learn the teaching content, and provide excellent feedback. S1208: The client records the learning videos that achieve a perfect score and sends them back to the server, where they are parsed and archived.
[0153] Please see Figure 11C The figure exemplifies an interactive diagram illustrating the generation of intelligent teaching-specific videos according to an embodiment of this application. As shown, the interaction in this embodiment involves learners, clients, and servers. Specifically: S1301. When the learner completes the full score recording (when the playback progress is not less than the progress threshold, the overall matching degree of the learning video is not less than the second matching degree threshold), the system will instruct the learner to upgrade the smart teaching exclusive video. S1302. The client receives the upgrade instruction and uploads the learning video to the server. S1303 The server performs AI background removal on the learning video and synchronizes the video footage of the intelligent teaching video and the learning video through key point matching and video fusion technology. Then, based on the finger key points provided during the recording process, the intelligent teaching video is optimized to obtain an exclusive intelligent teaching video. S1304. Server-side AI semantic analysis intelligent teaching exclusive videos, summarize and refine the video titles (no more than 10 characters). S1305. The server will synchronize the intelligent teaching videos to the search link and the information flow link. S1306. The client receives the smart teaching exclusive video returned by the server and displays the smart teaching exclusive video on the My Smart Courses page. S1307. Learners can access the "My Smart Courses" page to view exclusive smart teaching videos and quickly review their lessons. S1308. When learners forget the relevant knowledge later, a search will be initiated again; S1309. The client obtains the search request, extracts the search information from the search request, retrieves the search information in the historical search records, and if it is determined that the same search has been performed in the past, the search information is marked with information from the historical smart records. S1310. The server marks historical intelligent records according to the search request and identifies relevant intelligent teaching videos or intelligent teaching exclusive videos from the database. S1311. The server will also search for the latest relevant content through the search engine, but will mark the smart teaching video or smart teaching exclusive video at the top of the search results (e.g., first), and sort other relevant content based on relative order, and then return the search results. S1312. The client displays search results visually, such as the title, subtitle, cover image, webpage link, etc. of each video, and adds tags to smart teaching videos or smart teaching exclusive videos to improve recognition. S1313 Learners see intelligent teaching videos or exclusive intelligent teaching videos at the top of the search results, allowing them to quickly review and study. S1314. Learners can randomly view exclusive smart teaching videos on the homepage discovery page and other information flow interfaces. S1305: Based on the Ebbinghaus forgetting curve and the generation time of intelligent teaching videos, the client calculates the time learners spend reviewing knowledge and randomly displays it in the front row of the information flow interface.
[0154] This application embodiment targets learners with self-learning needs related to action-based learning. It combines AI technology to provide a highly personalized and intelligent online learning experience. Compared to related technologies, this application embodiment offers advantages such as stronger personalized recommendations, real-time feedback and guidance, integration of online learning and real-time practice, adaptive learning speed, and supplementary knowledge recommendations. Specifically, it has the following characteristics in terms of product and user value: 1. Personalized Recommendation: This application's embodiments generate high-quality intelligent teaching videos based on learners' in-depth interaction data in search results, achieving stronger personalized recommendations that can greatly improve learners' learning efficiency and satisfaction. Furthermore, the upgraded intelligent teaching videos are integrated with the learner's own learning process, achieving synchronized practice between teachers and students. This content is not only searchable in corresponding search results but also proactively recommended in multiple scenarios at specific times based on the Ebbinghaus forgetting curve method, achieving review and reinforcement, thereby improving learners' learning outcomes.
[0155] 2. Real-time Feedback and Guidance: This application's embodiments combine AI motion capture to detect and correct learners' erroneous movements in real time. Learners can identify and correct their movements more quickly and accurately, avoiding the formation of incorrect movement habits and enhancing their learning interest and motivation. For example, when a learner is learning to play a chord on the guitar, if their fingers are in the wrong position, the AI will analyze the learner's hand gestures, detect and point out the error in real time, and provide correct optimization instructions, allowing the learner to correct it immediately, greatly improving learning efficiency.
[0156] 3. Integrating Online Learning and Real-Time Practice: Learners can record and practice in real-time while watching video tutorials. This learning model allows learners to immediately put what they've learned into practice, deepening their understanding and memory. For example, when learning fingerstyle guitar, learners can watch online demonstrations and simultaneously imitate and practice in real-world settings. Furthermore, real-time recordings can be saved, allowing learners to review and summarize their learning journey, continuously improving. Learners can also share their practice results, learning from, exchanging ideas with, and progressing alongside other learners, increasing the enjoyment and motivation of learning.
[0157] 4. Adaptive Playback Rate: This embodiment of the application intelligently adjusts the video playback speed by recognizing the learner's learning stage, achieving adaptive learning. Learners can learn at a speed that suits them. For example, a beginner learning a complex chord change may need to slow down the video to clearly see every detail; while for parts already mastered, learners can choose to speed up the video to save learning time, ensuring that each type of learner can learn at their most comfortable speed.
[0158] 5. Intelligent Supplementary Knowledge Recommendation: This embodiment of the application can also automatically identify knowledge points that need supplementation during the learner's learning process and recommend relevant videos in real time. This supplementary knowledge recommendation function can help learners improve their knowledge system and enhance learning effectiveness. For example, when a learner is learning to play a song, if the system detects that the learner is having difficulty with a certain chord transition, the system will recommend some videos that specifically teach this chord transition to help the learner better master this skill. In this way, learners do not need to search for supplementary materials elsewhere, saving a lot of time and energy and improving learning efficiency.
[0159] This application provides a video processing apparatus, such as... Figure 12 As shown, the video processing device may include: a video display module 1201 and a synchronous playback module 1202, wherein, Video display module 1201 is used to display the first action-type teaching video and the first control; The synchronous playback module 1202, in response to the first control being triggered, records real-time learning actions during the playback of the first action-type teaching video, so as to play the first action-type teaching video and the real-time learning video simultaneously.
[0160] The first action-related instructional video is either a first intelligent instructional video or a first intelligent instructional exclusive video; the first intelligent instructional exclusive video is a video generated by fusing a second intelligent instructional video and historical learning videos. For any one of the first intelligent teaching video and the second intelligent teaching video, the intelligent teaching video is generated by splicing video segments from at least one second action-type teaching video.
[0161] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.
[0162] This application provides an electronic device, including a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of a video processing method. Compared with related technologies, it can achieve the following: by displaying a first action-type teaching video and a first control, the first control is used to trigger video recording and guide the actions in the recorded video based on the first action-type teaching video. In response to the first control being triggered, the learner's real-time learning actions are recorded during the playback of the first action-type teaching video. On the one hand, the first action-type teaching video and the real-time learning video are played simultaneously, allowing the learner to watch in real time and perceive whether their learning actions match the teaching actions. This application embodiment enables learners to record actions while watching videos, allowing for more intuitive comparison and learning. Furthermore, compared to simply showing learners public domain teaching videos, the first action-type teaching video displayed in this application embodiment is a first intelligent teaching video or... The first intelligent teaching video is generated by merging the second intelligent teaching video and historical learning videos. In other words, the first intelligent teaching video includes content from both the second intelligent teaching video and the learner's own learning videos from the second intelligent teaching video. This can quickly recall the learner's memories and help them consolidate their knowledge. For any one of the first and second intelligent teaching videos, the intelligent teaching video is generated by splicing video clips from at least one second action-type teaching video, thus collecting a wider range of basic materials for generating the first action-type teaching video. Furthermore, by extracting video clips from various second action-type teaching videos and splicing them together, the essence of multiple teaching videos can be presented to the learner at once. This significantly reduces the effort learners need to spend selecting teaching videos, significantly improves learning efficiency, and enhances the user experience.
[0163] In one alternative embodiment, an electronic device is provided, such as Figure 13 As shown, Figure 13The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.
[0164] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0165] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, bus 4002 is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus.
[0166] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.
[0167] The memory 4003 stores computer programs that execute embodiments of this application, and its execution is controlled by the processor 4001. The processor 4001 executes the computer programs stored in the memory 4003 to implement the steps shown in the foregoing method embodiments.
[0168] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.
[0169] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.
[0170] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the illustrations or text descriptions.
[0171] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.
[0172] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.
Claims
1. A video processing method, characterized in that, include: Show the first tutorial video for the first action class and the first control; In response to the first control being triggered, real-time learning actions are recorded during the playback of the first action-type teaching video, so that the first action-type teaching video and the real-time learning video can be played simultaneously. The first action-related instructional video is either a first intelligent instructional video or a first intelligent instructional exclusive video; the first intelligent instructional exclusive video is a video generated based on a second intelligent instructional video and historical learning videos. For any one of the first intelligent teaching video and the second intelligent teaching video, the intelligent teaching video is generated by splicing video segments from at least one second action-type teaching video.
2. The method according to claim 1, characterized in that, The second action-related instructional video includes at least one of an interactive action-related instructional video and an instructional video related to the interactive action-related instructional video; The video clips include at least one of a clip of interest and a clip illustrating a first target action, the first target action being a key action determined based on historical learning videos.
3. The method according to claim 1, characterized in that, The first action-type tutorial video and the first control mentioned above include at least one of the following: In response to a first search operation for action-related instruction, if there is a first action-related instruction video related to the first search operation, the first action-related instruction video will be displayed as the priority search result. In response to opening a preset information flow interface, if it is determined that the duration since the first action-type teaching video was generated or since the last display of the first action-type teaching video meets the preset requirements, the first action-type teaching video will be displayed preferentially in the information flow interface.
4. The method according to claim 1, characterized in that, The first action-based instructional video is the first intelligent instructional video; The first action-type tutorial video and the first control are described, including: In response to the first search operation for action-related teaching, if there is no intelligent teaching video or intelligent teaching exclusive video related to the first search operation, the first search result is displayed, which includes multiple action-related teaching videos. In response to at least one action-related instructional video in the first search result being interacted with, a second control is displayed; In response to the second control being triggered, the first action-type tutorial video and the first control are displayed.
5. The method according to claim 1, characterized in that, Also includes: In response to determining the real-time overall matching degree between the learning video and the first action-type teaching video, the playback speed of the first action-type teaching video is adjusted according to the overall matching degree; The overall matching degree is positively correlated with the playback speed.
6. The method according to claim 5, characterized in that, It also includes at least one of the following: In response to the overall matching degree being lower than the first matching degree threshold, at least one third action class teaching video is recommended. The third action class teaching video is a teaching video with a learning difficulty lower than the first action class teaching video or is used to teach a second target action. The second target action is the guidance action in the first action class teaching video that causes the overall matching degree to be lower than the first matching degree threshold. In response to the overall matching degree being no less than a second matching degree threshold when the playback progress is no less than a progress threshold, a second intelligent teaching exclusive video is generated. The second intelligent teaching exclusive video is a video generated based on the first action-type teaching video and the real-time learning video. Wherein, the first matching degree threshold is lower than the second matching degree threshold.
7. The method according to claim 1 or 6, characterized in that, For any one of the first and second smart teaching exclusive videos, at least one video frame in the smart teaching exclusive video includes a first image and a second image. The first image is determined based on the first video frame in the corresponding action-type instructional video; The second image is the foreground image in the second video frame of the corresponding learning video.
8. The method according to claim 7, characterized in that, The first action-based instructional video is a first intelligent instructional video, and the first image in the second intelligent instructional video is based on the first video frame of the first intelligent instructional video; The first action-type teaching video is a first intelligent teaching exclusive video, and the first image in the second intelligent teaching exclusive video is the first video frame in either the second intelligent teaching video or the first intelligent teaching exclusive video.
9. The method according to claim 1, characterized in that, The simultaneous playback of the first action-related instructional video and the real-time learning video also includes: Real-time display of action guidance information for the learning video; The action guidance information is used to instruct at least one of the following: The learning videos demonstrate the methods for correcting learning actions; Feature information of the learning actions shown in the learning video; The characteristic information of the teaching actions shown in the first action-type teaching video.
10. The method according to claim 6, characterized in that, For any one of the first and second intelligent teaching exclusive videos, the intelligent teaching exclusive video also includes checkpoint indication information; The checkpoint indication information is used to indicate teaching actions in the corresponding learning process where the number of guidance attempts exceeds the second threshold.
11. The method according to claim 1, characterized in that, The segment of interest includes at least one of the following: Video segments in the corresponding second-action instructional videos that are repeated more times than the threshold for the first playback; Based on learners' interaction records, and by understanding learners' interests and preferences, the system identifies segments of interest from the corresponding second-action instructional videos.
12. The method according to claim 7, characterized in that, The frame matching degree between the second video frame and the first video frame is higher than a third preset threshold; The frame matching degree refers to the matching degree between the guidance action shown in the first video frame and the learning action shown in the second video frame. The overall matching degree is positively correlated with the number of video frame pairs whose frame matching degree is higher than the third preset threshold. Each video frame pair includes one video frame from the first action-type teaching video and one video frame from the corresponding learning video.
13. The method according to claim 6, characterized in that, The generation of the second intelligent teaching-specific video also includes: Display a third control; In response to the triggering of the third control, at least one of the second intelligent teaching exclusive video and the first action-type teaching video is displayed in a preset information flow interface; The information flow interface includes at least one of the following: The information feed interface displays the search results; A feed interface displaying multimedia information posted by learner accounts; A feed interface that displays trending information or information recommended based on interests.
14. The method according to claim 1 or 4, characterized in that, The first intelligent teaching video is generated in the following way: Obtain at least one second-action tutorial video based on the interaction log; For each second action category instructional video, obtain the segments of interest in the second action category instructional video based on the pre-trained interest recognition model; Based on the semantic correlation of each segment of interest, the splicing order of each segment of interest is determined, and the segments of interest are spliced according to the splicing order to obtain the first intelligent teaching video; The interest recognition model is trained in a supervised or unsupervised manner on sample segments of interest and is used to identify the learner's level of interest in at least one of the themes, styles, and content of the instructional video.
15. The method according to claim 9, characterized in that, The action guidance information is generated in the following way: For each video frame in the learning video, the video frame is classified by a pre-built action classification model to obtain a classification result. The classification result is used to indicate whether the learning action in the video frame is correct or the learning action is incorrect. The action guidance information is generated based on the classification results; The action classification model is trained using multiple positive and negative samples, where both positive and negative samples are video frames and each has labeled information. The annotation information of the positive samples indicates that the action displayed in the corresponding video frame is correct; the annotation information of the negative samples indicates the type of error in the action displayed in the corresponding video frame.
16. The method according to claim 12, characterized in that, The frame matching degree of the video frame pair is determined in the following way: Motion capture is performed on two video frames in the video frame pair to obtain the motion key points of each of the two video frames. The frame matching degree of the video frame pair is obtained based on the distance between the action key points of the two video frames.
17. A video processing apparatus, characterized in that, include: The video display module is used to display the first action-based tutorial video and the first control. The synchronous playback module, in response to the first control being triggered, records real-time learning actions while playing the first action-type teaching video, so as to play the first action-type teaching video and the real-time learning video simultaneously. The first action-related instructional video is either a first intelligent instructional video or a first intelligent instructional exclusive video; the first intelligent instructional exclusive video is a video generated based on a second intelligent instructional video and historical learning videos. For any one of the first intelligent teaching video and the second intelligent teaching video, the intelligent teaching video is a video segment generated by splicing video segments from at least one second action-type teaching video.
18. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the video processing method according to any one of claims 1-16.
19. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the video processing method according to any one of claims 1-16.
20. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video processing method according to any one of claims 1-16.