Video lecture learning effect improving method and system based on GPT technology

By extracting video and audio through a browser plugin and using GPT technology to generate structured summaries and comprehension questions, combined with a linear regression model to recommend the optimal video segmentation strategy, this approach addresses the problems of difficulty in concentrating, low efficiency, lack of personalization, and insufficient interactivity in traditional online learning, thus achieving an efficient, personalized, and interactive learning experience.

CN120980299APending Publication Date: 2025-11-18BEIJING GUONENG GUOYUAN ENERGY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511256538.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-03
Publication Date
2025-11-18

AI Technical Summary

Technical Problem

In traditional online learning environments, learners struggle to maintain focus, resulting in low learning efficiency, a lack of personalized learning experiences, and issues such as language barriers and insufficient interactivity.

Method used

The system extracts video audio through a browser plugin, uses GPT technology to generate structured summaries and comprehension questions, and combines a linear regression model to recommend the optimal video segmentation strategy. It also provides multilingual support and interactive features to enhance user engagement.

Benefits of technology

It effectively maintains students' attention, improves learning efficiency, provides a personalized learning experience, overcomes language barriers, enhances interactivity, and significantly improves learning outcomes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120980299A_ABST
    Figure CN120980299A_ABST
Patent Text Reader

Abstract

The invention discloses a video lecture learning effect improving method and system based on a GPT technology, and aims to solve the problems that in a traditional online learning environment in the prior art, the attention of students is difficult to continuously concentrate, the learning efficiency is low, personalized learning experience is lacked, language barriers exist, and interactivity is insufficient. The method comprises the following steps: in response to an operation instruction of a user, extracting an audio part of a corresponding video file from an online video platform specified by the user; generating text transcription content corresponding to the audio part; processing the text transcription content by using a natural language processing model, generating a structured abstract, formulating an understanding problem based on the abstract content, and translating the abstract content and the understanding problem into a target language; analyzing user historical interaction data through a linear regression model, and dynamically predicting and recommending an optimal video segmentation strategy; an interaction part independent of the webpage is provided on the webpage content, and the problem is solved through the webpage interaction method and device.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a learning effect improvement method and system. BACKGROUND

[0002] In the traditional online learning environment, learners usually acquire knowledge by watching video lectures. However, this method has some obvious limitations:

[0003] (1) Difficulty in maintaining attention: When watching long videos, learners' attention tends to gradually disperse over time, resulting in low learning efficiency. Studies have shown that long-term passive learning can easily cause fatigue and inattention.

[0004] (2) Low learning efficiency: Traditional video lectures usually require learners to watch the entire process, which not only consumes time but also makes it difficult to quickly obtain key information. Learners often need to spend a lot of time screening and understanding the lecture content, resulting in low learning efficiency.

[0005] (3) Lack of personalized learning experience: Different learners have different learning abilities and learning habits, and traditional learning methods are difficult to meet individual needs. Existing learning tools often lack analysis of user learning behavior and personalized recommendation functions.

[0006] (4) Language barrier: For non-English-speaking learners or Chinese learners, language barriers may affect learning effectiveness. Existing learning tools often lack multilingual support, making it difficult for these learners to access high-quality learning resources.

[0007] (5) Lack of interactivity: Traditional video lectures lack interactivity, making learners feel dull. Existing learning tools often lack interactive learning functions, which cannot effectively stimulate learners' enthusiasm and participation. SUMMARY

[0008] The present application provides a video lecture learning effect improvement method and system based on GPT technology, aiming to solve the problems of students' difficulty in maintaining attention, low learning efficiency, lack of personalized learning experience, language barriers and lack of interactivity in the traditional online learning environment of the prior art.

[0009] In a first aspect, a video lecture learning effect improvement method based on GPT technology is provided, which takes a browser plug-in as a carrier and includes:

[0010] S1, in response to the user's operation instruction to the browser plug-in, the browser plug-in interacts with the web page content, extracts the audio part of the corresponding video file from the currently opened online video platform of the user;

[0011] S2, converting the extracted audio into a specified format as needed, calling a cloud-based speech recognition API or a local speech model to generate text transcription content corresponding to the audio portion; the text transcription content can accurately capture the key content of the lecture, taking into account various accents, dialects and speech nuances, to ensure the generation of high-quality manuscripts;

[0012] S3, processing the text transcription content using a natural language processing model to generate structured summaries, formulating comprehension questions based on the summary content, and translating the summary content and comprehension questions into a target language;

[0013] S4, using a linear regression model to correlate video segment length with user performance, i.e., using video segment length as the independent variable and correct answer rate of multiple-choice questions as the dependent variable, analyzing user historical interaction data through a linear regression model to dynamically predict and recommend the optimal video segmentation strategy;

[0014] S5, providing an interactive part independent of the webpage on the webpage content, which allows users to easily select videos, view summaries and interact with comprehension questions, enhancing user engagement and learning experience.

[0015] In the above scheme, optionally, step S3 specifically includes:

[0016] S31, extracting key information from the text transcription content to generate a structured summary; the key information includes video key points, main points and important findings;

[0017] S32, formulating comprehension questions based on the summary content, and designing the comprehension questions in the form of multiple-choice questions, each of which links back to the relevant excerpts of the original manuscript, showing the questions and their answer basis;

[0018] S33, translating the summary content and comprehension questions into a target language while maintaining the consistency of the translation of academic terminology; the translation process uses advanced language models to achieve accurate and context-aware translation, ensuring the accuracy and readability of the translated content.

[0019] In the above scheme, optionally, in step S4, dynamically predicting and recommending the optimal video segmentation strategy specifically includes:

[0020] Analyzing user historical interaction data through a linear regression model to predict the point at which user performance begins to decline, which corresponds to the optimal video segment length for maintaining user engagement and understanding;

[0021] Updating the linear regression model periodically, and each time updating refers to the new optimal video segment length data to form the optimal video segmentation strategy.

[0022] In the above scheme, optionally, in step S5, the interactive part independent of the webpage is provided in the form of a pop-up window on the webpage content; the interaction with the comprehension question is realized by providing a video jump button on the pop-up window to locate to the original text associated segment.

[0023] In the above scheme, optionally, in step S1, the audio part of the video file is extracted by using digital signal processing technology or multimedia framework FFmpeg.

[0024] In the above scheme, optionally, in step S2, the text transcription content corresponding to the audio part is generated by using Google's speech-to-text technology, which can process multiple accents and dialects.

[0025] In the above scheme, optionally, in step S3, the structured summary and the comprehension question based on the summary content are generated by using OpenAI's GPT technology; the summary content and the comprehension question are translated into the target language by using OpenAI's advanced language model, and accurate and context-aware translation is realized.

[0026] In the above scheme, optionally, the method realizes the automatic processing process from transcription, summary generation to question preparation and translation by integrating AI-generated content technology, i.e., AIGC technology, to ensure efficient, accurate and high-quality output.

[0027] In the above scheme, optionally, the method allows users to easily select videos, summaries and interact with comprehension questions through user-friendly interface design, which combines seamless integration of background AI technology and intuitive front-end design to enhance user engagement and learning experience.

[0028] In a second aspect, a video lecture learning effect improvement system based on GPT technology is provided, comprising:

[0029] A video processing module for responding to user operation instructions and interacting with webpage content to extract the audio part of the corresponding video file from the user-specified online video platform, ensuring that only the audio part is passed to the next step for transcription;

[0030] A speech transcription module for converting the extracted audio into a specified format as needed, calling a cloud-based speech recognition API or a local speech model to generate text transcription content corresponding to the audio part; the text transcription content can accurately capture the key content of the lecture, taking into account various accents, dialects and speech nuances to ensure the generation of high-quality manuscripts;

[0031] a content generation module for utilizing a natural language processing model to perform the following operations: extracting key information from the text transcription content to generate a structured summary; the key information includes video key points, main points, and important findings; formulating comprehension questions based on the summary content and designing them in the form of multiple-choice questions, each linking back to relevant excerpts from the original manuscript, showing the question and its answer basis; translating the summary content and comprehension questions into the target language while maintaining consistency in the translation of academic terminology; the translation process uses advanced language models to achieve accurate and context-aware translation, ensuring the accuracy and readability of the translated content;

[0032] a personalized adaptation module for using a linear regression model to correlate video segment length with user performance, i.e., using video segment length as the independent variable and the accuracy of the multiple-choice answers as the dependent variable, analyzing user historical interaction data through a linear regression model to dynamically predict and recommend the optimal video segmentation strategy;

[0033] a user interface module for providing a web page content-independent interaction section on the web page, which allows users to easily select videos, view summaries, and interact with comprehension questions, enhancing user engagement and learning experience.

[0034] Compared with the prior art, the present application has at least the following beneficial effects:

[0035] Based on further analysis and research of the problems in the prior art, the present application recognizes that in traditional online learning environments, students have difficulty maintaining focus, learning efficiency is low, personalized learning experience is lacking, language barriers exist, and interaction is insufficient. The present application proposes a video lecture learning effect improvement method based on GPT technology. This method extracts audio from videos and transcribes it with high quality, uses natural language processing models to generate structured summaries and comprehension questions, translates the content into the target language, and dynamically predicts and recommends the optimal video segmentation strategy based on user interaction data. Through these technical means, the present application can effectively maintain student attention, improve learning efficiency, provide personalized learning experience, solve language barriers, and enhance interactivity, thereby significantly improving learning effectiveness and user experience. BRIEF DESCRIPTION OF DRAWINGS

[0036] Figure 1 A flowchart of a video lecture learning effect improvement method based on GPT technology is provided for an embodiment of the present application.

[0037] Figure 2 A module architecture block diagram of a video lecture learning effect improvement device based on GPT technology is provided for an embodiment of the present application. DETAILED DESCRIPTION

[0038] In order to make the purpose, technical scheme and advantages of the present application clearer, further detailed description will be given below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and not to limit the present application.

[0039] In the description of the present application: unless otherwise specified, the meaning of "multiple" is two or more. The expressions "include", "contain", "have" and the like also mean "not limited to" (some units, components, materials, steps, etc.).

[0040] The present application aims at a web plug-in application suitable for students and professionals, which can provide summaries, interactive Q&A and multilingual support for video lectures, helping non-English speakers and Chinese learners improve their learning effectiveness. It is particularly useful in disciplines such as history, literature, philosophy, social science, economics, psychology and political science, as OpenAI's GPT technology can help understand complex concepts. Although it is less effective in data-intensive disciplines such as higher mathematics or natural sciences, it has an advantage in processing text-based information, making it a powerful and diverse tool in the humanities and social sciences.

[0041] In one embodiment, a GPT technology-based video lecture learning effectiveness improvement method is provided, which takes a browser plug-in as a carrier and includes steps S1-S5.

[0042] S1, in response to the user's operation instruction to the browser plug-in, the browser plug-in interacts with the web content to extract the audio part of the corresponding video file from the online video platform currently opened by the user.

[0043] In step S1, the browser extension or web script is used to interact with the web content to access and manipulate the video file. The audio part is accurately extracted by using corresponding audio processing technology, avoiding other interference information in the video, such as background music. Since users obtain information from more than one online video platform when using it, the method's compatibility and versatility are further improved by supporting multiple online video platforms.

[0044] In step S1, the browser plug-in is set in the navigation bar of the browser. When the user opens the browser and opens the corresponding online video platform in the browser, the browser plug-in is allowed to interact with the current page content of the browser by clicking the browser plug-in in the navigation bar, and the audio part of the corresponding video file is extracted from the online video platform currently opened by the user.

[0045] In one embodiment, in step S1, the audio part of the video file is extracted by using digital signal processing technology or multimedia framework FFmpeg.

[0046] By using digital signal processing techniques or FFmpeg, the audio extraction of video files can be efficiently implemented, providing high-quality audio input for subsequent audio transcription and processing. Digital signal processing techniques offer high flexibility, allowing for customized processing based on specific requirements. FFmpeg provides a simple and easy-to-use solution, suitable for rapid implementation and automated processing. FFmpeg supports multiple video and audio formats, ensuring good compatibility and allowing for wide application in different video platforms and file formats.

[0047] In step S2, the extracted audio is converted to a specified format as needed, and a cloud-based speech recognition API or a local speech model is called to generate text transcription content corresponding to the audio portion. The text transcription content can accurately capture the key content of the lecture, taking into account various accents, dialects, and subtle differences in speech, ensuring the generation of high-quality manuscripts.

[0048] In one embodiment, in step S2, Google's speech-to-text technology is used to generate text transcription content corresponding to the audio portion. This technology can handle multiple accents and dialects.

[0049] In step S2, before passing the audio to Google's speech-to-text technology, it is necessary to ensure that the audio format meets the requirements of the API. Common audio formats include WAV, FLAC, MP3, etc. If the audio format does not meet the requirements, audio processing tools such as FFmpeg can be used to convert it to a supported format. Speech recognition can be performed by calling a cloud-based speech recognition API or using a local speech model. Cloud-based APIs have higher accuracy and more powerful processing capabilities, but require network connectivity; local models can be used in offline environments, but require more computing resources and storage space. Users can train and optimize the model according to their needs, thereby accurately capturing the key content of the lecture, taking into account various accents, dialects, and subtle differences in speech, to improve the recognition of different speech characteristics.

[0050] In step S2, Google's speech-to-text technology can efficiently generate high-quality text transcription content. This technology has high accuracy, strong processing capability, and good flexibility, meeting the needs of various application scenarios. Through reasonable configuration and optimization, transcription effect and user experience can be further improved.

[0051] In step S3, the text transcription content is processed using a natural language processing model to generate a structured summary, develop comprehension questions based on the summary content, and translate the summary content and comprehension questions into the target language.

[0052] In one embodiment, the step S3 specifically comprises:

[0053] S31, extracting key information from the text transcription content to generate a structured summary; the key information includes video key points, main viewpoints, and important findings;

[0054] S32, formulating comprehension questions based on the summary content, and designing the comprehension questions in the form of multiple-choice questions, each of which is linked back to the relevant excerpts of the original manuscript to show the question and its answer basis;

[0055] S33, translating the summary content and comprehension questions into the target language while maintaining the consistency of the translation of academic terms; the translation process uses advanced language models to achieve accurate and context-aware translation, ensuring the accuracy and readability of the translated content.

[0056] Identify important time nodes or theme transition points in the video and extract the relevant text content. The text analysis capabilities of GPT can be used to guide the model to identify key points through prompts. Extract the core viewpoints and arguments in the text. GPT can extract the main viewpoints and arguments through overall understanding of the text. Identify important conclusions or findings in the text. GPT can identify important findings or conclusions through semantic understanding.

[0057] The generated summary should have a clear structure, such as organizing key information in chronological order, theme classification, or importance ranking. Use the text generation capabilities of GPT to generate summaries according to pre-set templates or frameworks. While extracting key information, ensure the brevity and accuracy of the summary. GPT generates concise and clear summaries through prompt optimization and parameter adjustment.

[0058] Formulate comprehension questions based on the summary content, and design the comprehension questions in the form of multiple-choice questions, each of which is linked back to the relevant excerpts of the original manuscript to show the question and its answer basis. Design different types of comprehension questions, such as factual questions, reasoning questions, and application questions, to comprehensively examine the user's understanding and mastery of the video content. GPT can generate diverse questions based on the summary content.

[0059] According to the difficulty of the video content and the learning level of the user, set the difficulty of the question reasonably. GPT can generate questions of different difficulty levels by adjusting the prompt and parameters. Design multiple choice items for each comprehension question, including the correct answer and the interference item. GPT can generate multiple possible answer options and select appropriate interference items based on the context. Link each multiple-choice question back to the relevant excerpts of the original manuscript to show the question and its answer basis.

[0060] Translate the summary content and comprehension questions into the target language while maintaining consistency in the translation of academic terminology. The translation process uses advanced language models to achieve accurate and contextually aware translations, ensuring the accuracy and readability of the translated content.

[0061] Users can create a library of academic terminology to store commonly used terms and their accurate translations. During the translation process, the library is prioritized to ensure accuracy and consistency. The context of academic terminology is considered to avoid translation errors due to ambiguous meanings. Advanced language models such as GPT-4 from OpenAI are used for translation to ensure accuracy and fluency. GPT can perform contextually aware translation. A translation quality evaluation mechanism is established to automatically evaluate and manually review translation results. Automatic evaluation indicators such as BLEU score and ROUGE score can be used in conjunction with manual translation expert reviews to ensure the quality of the translated content.

[0062] In one embodiment, in step S3, structured summaries and comprehension questions based on summary content are generated using GPT technology from OpenAI; summary content and comprehension questions are translated into target language using advanced language models from OpenAI, and accurate and contextually aware translation is achieved.

[0063] In this embodiment, in step S3, GPT technology from OpenAI is used to process text transcription content, generate structured summaries, formulate comprehension questions, and perform translation. The powerful natural language processing capabilities of GPT technology are fully utilized to ensure that the generated content is of high quality and has good readability and accuracy.

[0064] S4, using a linear regression model to correlate video segment length with user performance, i.e. using video segment length as the independent variable and selecting the correct answer rate as the dependent variable, analyzing user historical interaction data through a linear regression model to dynamically predict and recommend the optimal video segmentation strategy;

[0065] In one embodiment, in step S4, dynamically predicting and recommending the optimal video segmentation strategy specifically includes:

[0066] Through the linear regression model, analyze the user's historical interaction data to predict the point at which the user's performance begins to decline. The length corresponding to this point is the optimal video segment length to maintain user engagement and comprehension;

[0067] Update the linear regression model regularly, and each time update refers to the new optimal video segment length data to form the optimal video segmentation strategy.

[0068] A linear regression model is used to correlate video segment length with user performance. By analyzing user historical interaction data, the optimal video segmentation strategy is dynamically predicted and recommended. The user's historical interaction data during the video learning process, such as viewing time, pause times, and correct answer rate, are collected as the basis for analysis. The video segment length is used as the independent variable, and the correct answer rate of the multiple-choice question is used as the dependent variable to build a linear regression model. Through model analysis, the relationship between the two is found out, and the best video segmentation length is found out to improve the user's learning effect. According to the prediction results of the model, the video segmentation strategy is dynamically adjusted. This requires real-time acquisition of user interaction data and timely updating of model prediction results to provide personalized segmentation strategy recommendations.

[0069] S5, providing a web-independent interactive part on the web content, which allows users to easily select videos, view summaries, and interact with comprehension questions, enhancing user engagement and learning experience.

[0070] In one embodiment, in step S5, a web-independent interactive part is provided in the form of a pop-up window on the web content; by providing a video jump button on the pop-up window to locate to the original text associated segment, interaction with comprehension questions is realized.

[0071] In this embodiment, in addition to the pop-up window, the web-independent interactive part can also have the following forms: (1) a floating toolbar: a floating toolbar is fixedly displayed on the side or bottom of the web page, and the user can expand or retract it at any time. The toolbar can include video selection buttons, summary display areas, and comprehension question interaction modules. (2) a fixed sidebar: a fixed-width sidebar is set on the left or right side of the web page, dedicated to interactive functions. The sidebar can include video lists, summary content, and comprehension questions, and the user can view different content by scrolling the sidebar. (3) a collapsible panel: a collapsible panel is set at the top or bottom of the web page, and the user can see the interactive content, including video selection, summary, and comprehension questions, after clicking the expand button. When collapsed, the panel is hidden and does not occupy much space.

[0072] In addition, it can also be: a group of floating buttons, an independent interactive module area, a dynamic floating label, a full-screen interaction mode, etc. The above content is only an example, and in the actual implementation process, other interactive forms that can achieve similar functions can also be used, and the present application does not make specific limitations.

[0073] Design a simple, intuitive, and easy-to-use interface that allows users to easily select videos, view summaries, and answer comprehension questions. The interface should seamlessly integrate with the webpage content while maintaining independence and not affecting the normal browsing experience. Optimize user experience by optimizing the interaction process, providing immediate feedback, and enhancing user engagement and learning experience. For example, you can set a progress bar, automatically save the answer progress, and other functions to facilitate users to pause and continue learning at any time. Ensure that the interactive part runs normally on different browsers and devices, with good compatibility and stability. Conduct thorough testing and optimization to avoid compatibility or performance issues that may affect user experience.

[0074] In one embodiment, the method integrates AI-generated content technology (AIGC technology) to automate the process from transcription, summary generation, to question preparation and translation, ensuring efficient, accurate, and high-quality output.

[0075] By integrating AIGC technology, the method realizes the automation process from transcription, summary generation, to question preparation and translation. This method not only improves processing efficiency, but also ensures the accuracy and quality of the output content. Through reasonable configuration and optimization, user experience and learning effect can be further improved, suitable for various education and learning scenarios.

[0076] In one embodiment, the method allows users to easily select videos, view summaries, and interact with comprehension questions through a user-friendly interface design that seamlessly integrates backend AI technology with intuitive front-end design to enhance user engagement and learning experience.

[0077] In this embodiment, through user-friendly interface design, combined with seamless integration of backend AI technology and intuitive front-end design, the complete process from video selection, summary viewing to comprehension question interaction is realized. This method not only improves user engagement, but also enhances learning experience.

[0078] In one embodiment, a video lecture learning effect improvement system based on GPT technology is provided, comprising:

[0079] A video processing module for responding to user operation instructions and interacting with webpage content, extracting the audio part of the corresponding video file from the user-specified online video platform, ensuring that only the audio part is passed to the next step for transcription;

[0080] a speech transcription module for converting the extracted audio into a specified format as needed, invoking cloud-based speech recognition APIs or local speech models to generate text transcription content corresponding to the audio segments; the text transcription content can accurately capture the key content of the lecture, taking into account various accents, dialects and speech nuances, ensuring the generation of high-quality transcripts;

[0081] a content generation module for using natural language processing models to perform the following operations: extracting key information from the text transcription content to generate a structured summary; the key information includes video key points, main points and important findings; formulating comprehension questions based on the summary content, and designing the comprehension questions in the form of multiple-choice questions, each linking back to relevant excerpts of the original transcript, showing the question and its answer basis; translating the summary content and comprehension questions into the target language while maintaining consistency in the translation of academic terminology; the translation process uses advanced language models to achieve accurate and context-aware translation, ensuring the accuracy and readability of the translated content;

[0082] a personalized adaptation module for using a linear regression model to correlate video segment length with user performance, i.e., using video segment length as the independent variable and the accuracy of the multiple-choice answers as the dependent variable, analyzing user historical interaction data through the linear regression model to dynamically predict and recommend the optimal video segmentation strategy;

[0083] a user interface module for providing a web page-independent interactive section on the web page content, which allows users to easily select videos, review and interact with comprehension questions, enhancing user engagement and learning experience.

[0084] The specific implementation of each module can be found in the above description of the method for improving the learning effect of video lectures based on GPT technology, and will not be repeated here.

[0085] The technical features of the above embodiments can be combined in any way. To make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present disclosure.

Claims

1. A method for improving the learning effectiveness of video lectures based on GPT technology, characterized in that, The method uses a browser plugin as a carrier and includes: S1, in response to the user's operation command to the browser plugin, the browser plugin interacts with the web page content and extracts the audio part of the corresponding video file from the online video platform currently opened by the user; S2, as needed, convert the extracted audio into a specified format, call the cloud speech recognition API or local speech model, and generate text transcription content corresponding to the audio part; the text transcription content can accurately capture the key content of the lecture, take into account various accents, dialects and subtle differences in speech, and ensure the generation of high-quality transcripts. S3, using a natural language processing model to process the transcribed text content, generate a structured summary, formulate comprehension questions based on the summary content, and translate the summary content and comprehension questions into the target language; S4 uses a linear regression model to link video segment length with user performance, that is, video segment length is used as the independent variable and the correctness of multiple-choice answers is used as the dependent variable. By analyzing historical user interaction data through the linear regression model, the optimal video segmentation strategy is dynamically predicted and recommended. S5 provides an interactive component independent of the webpage content, which allows users to easily select videos, view summaries, and interact with comprehension questions, enhancing user engagement and learning experience.

2. The method for improving the learning effect of video lectures based on GPT technology according to claim 1, characterized in that, Step S3 specifically includes: S31, extract key information from the transcribed text to generate a structured summary; the key information includes key points of the video, main viewpoints, and important findings; S32, Based on the abstract content, formulate comprehension questions and design these comprehension questions as multiple-choice questions. Each multiple-choice question links back to the relevant excerpt in the original manuscript, showing the question and the basis for the answer; S33, the abstract content and comprehension questions are translated into the target language while maintaining consistency in the translation of subject-specific terminology; the translation process uses advanced language models to achieve accurate and context-aware translation, ensuring the accuracy and readability of the translated content.

3. The method for improving the learning effect of video lectures based on GPT technology according to claim 1, characterized in that, In step S4, dynamically predicting and recommending the optimal video segmentation strategy specifically includes: By analyzing historical user interaction data using a linear regression model, the point at which user performance begins to decline can be predicted. The duration corresponding to this point is the optimal video segment length for maintaining user engagement and comprehension. The linear regression model is updated regularly, and each update references the new optimal video segment length data to form the optimal video segmentation strategy.

4. The method for improving the learning effect of video lectures based on GPT technology according to claim 1, characterized in that, In step S5, an interactive element independent of the webpage is provided on the webpage content in the form of a pop-up window; by providing a video jump button on the pop-up window to locate the relevant segment in the original text, interaction with comprehension questions is achieved.

5. The method for improving the learning effect of video lectures based on GPT technology according to claim 1, characterized in that, In step S1, the audio portion of the video file is extracted using digital signal processing technology or the multimedia framework FFmpeg.

6. The method for improving the learning effect of video lectures based on GPT technology according to claim 1, characterized in that, In step S2, Google's speech-to-text technology is used to generate text transcription of the corresponding audio portion. Google's speech-to-text technology can handle various accents and dialects.

7. The method for improving the learning effect of video lectures based on GPT technology according to claim 1, characterized in that, In step S3, OpenAI's GPT technology is used to generate a structured summary and formulate comprehension questions based on the summary content; OpenAI's advanced language model is used to translate the summary content and comprehension questions into the target language, achieving accurate and context-aware translation.

8. The method for improving the learning effect of video lectures based on GPT technology according to claim 1, characterized in that, The method integrates AI-generated content technology, namely AIGC technology, to automate the process from transcription and summary generation to question formulation and translation, ensuring efficient, accurate and high-quality output.

9. The method for improving the learning effect of video lectures based on GPT technology according to claim 1, characterized in that, The method features a user-friendly interface that allows users to easily select videos, view summaries, and interact with comprehension questions. The interface design seamlessly integrates backend AI technology with an intuitive front-end design to enhance user engagement and learning experience.

10. A video lecture learning effectiveness improvement system based on GPT technology, characterized in that, include: The video processing module is used to respond to user operation commands, interact with web page content, extract the audio part of the corresponding video file from the online video platform specified by the user, and ensure that only the audio part is passed to the next step for transcription. The speech transcription module is used to convert the extracted audio into a specified format as needed, call the cloud speech recognition API or local speech model, and generate text transcription content corresponding to the audio part; the text transcription content can accurately capture the key content of the lecture, take into account various accents, dialects and subtle differences in speech, and ensure the generation of high-quality transcripts. The content generation module is used to perform the following operations using a natural language processing model: extracting key information from the transcribed text to generate a structured summary; the key information includes video key points, main viewpoints, and important findings; formulating comprehension questions based on the summary content, and designing the comprehension questions as multiple-choice questions, each of which links back to relevant excerpts from the original manuscript, showing the question and the basis for the answer; translating the summary content and comprehension questions into the target language, while preserving the consistency of the translation of subject-specific terminology; The translation process uses advanced language models to achieve accurate and context-aware translation, ensuring the accuracy and readability of the translated content; The personalized adaptation module is used to link video segment length with user performance using a linear regression model. That is, the video segment length is used as the independent variable and the correctness of multiple-choice answers is used as the dependent variable. By analyzing the user's historical interaction data through the linear regression model, the optimal video segmentation strategy is dynamically predicted and recommended. A user interface module is used to provide interactive elements on web page content that are independent of the web page itself. These interactive elements allow users to easily select videos, view summaries, and interact with comprehension questions, enhancing user engagement and the learning experience.