English video splitting learning method based on large model

Through the English video split learning method based on the big model, combined with user feedback and multi-dimensional data fusion technology, intelligent splitting and personalized adjustment of video content are achieved, solving the problem of insufficient flexibility and accuracy of splitting results in the existing technology, and improving learning efficiency and personalized teaching experience.

CN120429463AActive Publication Date: 2025-08-05读书郎教育科技有限公司
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510511638.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-08-05
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing technology lacks user feedback and dynamic teaching adjustments in video splitting, resulting in low flexibility and accuracy of splitting results and cannot meet personalized teaching needs.

Method used

The English video split learning method based on the big model is adopted. By obtaining the video teaching text and the speech transcription text, dividing it into single sentences and depositing it into the difficulty vocabulary library, integrating videos are generated and integrated videos are generated based on user feedback, and dynamic adjustments of video content are made, including subtitle correction, example sentence supplementation, anchor point playback and slicing interval adjustment.

Benefits of technology

It realizes intelligent splitting and personalized adjustment of video content, accurately identify user learning bottlenecks, eliminates redundant information, dynamically optimizes teaching resources, and significantly improves learning efficiency and personalized teaching experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120429463A_ABST
    Figure CN120429463A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video splitting, in particular to an English video splitting learning method based on a large model. According to the method, a teaching text and a voice transcription text of a teaching video are obtained through ASR and OCR technologies respectively, the text is segmented into single sentences, vocabularies are extracted and stored in a difficulty vocabulary library, and an integrated video meeting the requirements of an examination outline is generated through integration according to the expected learning level and difficulty quantitative indexes of a user. Segmenting the video length according to the learning duration demand of the user; a user playing record, a repetition rate, a complete playing rate and interaction feedback data are collected, a user learning level and video teaching difficulty are comprehensively evaluated through a big data model, and video content is dynamically adjusted. According to the scheme, the limitation of a traditional static splitting method is filled, the effects of effectively reducing knowledge omission and improving the learning efficiency are achieved through dynamic optimization and personalized customization of the teaching content, meanwhile, accurately matched teaching resources are provided for learners of different levels, and the learning efficiency is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video segmentation, and in particular to a method for learning English video segmentation based on a large model. Background Art

[0002] With the widespread use of mobile devices and the continuous improvement of network speeds, short, concise content with high traffic and high reach has become increasingly popular among major platforms, user groups, and the capital market. As an emerging form of internet communication, short videos have not only quickly gained popularity but also helped promote and disseminate related drama video resources.

[0003] The patent document with Chinese patent publication number CN110493637A discloses a video splitting method and device, the technical point of which is to obtain the video content features of the video to be processed and the time information corresponding to the video content features; determine the plot transition points of the video to be processed based on the video content features and the time information corresponding to the video content features; and split the video to be processed based on the plot transition points. By obtaining the video content features of the video to be processed and the time information corresponding to the video content features, and determining the plot transition points of the video to be processed based on the video content features and the time information corresponding to the video content features, the video to be processed can be automatically split according to the plot content based on the determined plot transition points of the video to be processed; at the same time, the invention focuses on the static splitting of video content and the detection of plot transitions, does not introduce user behavior data and interactive feedback, and lacks the ability to adapt to dynamic changes in teaching content and personalized needs, resulting in low flexibility and accuracy of the splitting results in practical applications. Summary of the Invention

[0004] To this end, the present invention provides a method for learning English video segmentation based on a large model, so as to overcome the problem that the existing technology focuses on static video content segmentation, resulting in low levels of user feedback integration and dynamic teaching adjustment, and low video segmentation accuracy.

[0005] To achieve the above objectives, the present invention provides a method for English video segmentation learning based on a large model, comprising:

[0006] Acquire input video data and process the data into standard input data, and integrate the standard input data according to the requirements of the exam difficulty outline to obtain a plurality of integrated videos;

[0007] The integrated video is segmented according to the user's actual learning time requirements to obtain several integrated video segments, including fragmented videos and long videos;

[0008] For any of the integrated video segments, the user's actual video viewing information is obtained, and the learning progress is analyzed based on the actual video viewing information. The video content adjustment direction is determined according to the analysis result, and the video content adjustment method is determined based on the video content adjustment direction, and the adjustment result is fed back to the user end.

[0009] Further, obtaining input video data and processing the data into standard input data includes,

[0010] Obtain video teaching text and voice transcription text;

[0011] The text content is divided into several separate sentences, and the parts of speech of words and phrases in any separate sentence are extracted and stored in the vocabulary database of corresponding difficulty. The corresponding examination syllabus requirements or application ability requirements are determined according to the expected learning level, and several integrated videos are obtained based on the quantitative indicators of difficulty.

[0012] Furthermore, several integrated videos are obtained based on the quantitative indicators of difficulty, including:

[0013] Set a standard matching range and determine the actual matching degree of the corresponding examination syllabus requirements or application ability requirements based on the user's actual learning level expectations;

[0014] The actual matching degree is determined according to the standard matching interval, and the type of the integrated video content is determined according to the determination result, including the first video type, the second video type, and the third video type.

[0015] Furthermore, segmenting the integrated video includes:

[0016] Segment the integrated video based on user feedback:

[0017] If the user feedback is that they need to learn quickly in fragmented time, the integrated video will be split into fragmented videos;

[0018] If the user feedback is to choose continuous deep learning needs, the integrated video will be split into long videos.

[0019] Furthermore, obtaining the user's actual video viewing information and analyzing the learning progress based on the actual video viewing information includes:

[0020] Actual video viewing information includes repetition rate, completion rate, and repetition completion index. Video viewing information is evaluated and compared with preset thresholds.

[0021] If the repetition completion index meets the evaluation criteria, the user's learning progress is determined to match the difficulty level requirements, and feedback measures on the user's learning progress are implemented;

[0022] If the repeated completion index does not meet the evaluation criteria, it is determined that the user's learning progress does not match the current difficulty progress requirements, the type of user's current learning bottleneck is determined, and the direction of video content adjustment is determined.

[0023] Furthermore, determining the user's current learning bottleneck and determining the direction of video content adjustment includes:

[0024] The learning bottleneck types include the first result, the second result, the third result and the fourth result;

[0025] When the learning bottleneck type is determined to be the first result, executing a subtitle error adjustment procedure;

[0026] When the learning bottleneck type is determined to be the second result, executing the example supplement procedure;

[0027] When the learning bottleneck type is determined to be the third result, executing the anchor point playback program;

[0028] When the learning bottleneck type is determined to be the fourth result, executing the segmentation interval adjustment procedure;

[0029] Further, executing the subtitle error adjustment procedure includes,

[0030] Read the real-time delay of subtitles and compare it with the standard delay.

[0031] When the real-time delay is less than or equal to the standard delay, there is no abnormality in the subtitles;

[0032] When the real-time delay is longer than the standard delay, the subtitles are abnormal;

[0033] If the judgment result is that there is no abnormality in the subtitles, the corresponding extension parameters are obtained through the large model analysis learning content, and the video length is extended according to the parameters;

[0034] If the judgment result is that the subtitles are abnormal, the subtitle display parameters are adjusted and the subtitle text content is verified, and the adjusted video viewing information is read and evaluated;

[0035] If the evaluation results do not meet the evaluation criteria and are excessively repetitive, the corresponding extension parameters are obtained through large-scale model analysis of the learning content, and the video duration is extended according to the parameters.

[0036] Further, executing the example supplementary procedure includes,

[0037] Determine the teaching difficulty level based on the type of video content currently being integrated, and extract the language logic template at that difficulty level;

[0038] Select several core words in the current integrated video content that match the teaching difficulty level from the corresponding difficulty vocabulary database;

[0039] The semantics between the core words are associated according to the language logic template to generate candidate sentences, which are stored in the candidate sentence information database and whether to call them is determined according to the actual needs of the user.

[0040] Further, executing the anchor point playback program includes,

[0041] Confirm the starting position of the repeat interval and set an anchor point at the starting position to quickly play back the repeated content from the anchor point;

[0042] Confirm user playback behavior within the standard judgment duration to dynamically adjust the anchor point position.

[0043] Furthermore, executing the segmentation interval adjustment procedure includes:

[0044] Extract the playback information of the target fast-forward interval, evaluate the user's current learning level based on the playback information, and quantify it into an actual learning level value;

[0045] Extracting and integrating playback information of the video, evaluating the current teaching level of the video based on the playback information and quantifying it into an actual teaching level value;

[0046] Among them, the playback information includes the user's historical learning records and interactive feedback;

[0047] Compare the actual teaching level value with the actual learning level value,

[0048] When the actual teaching level value is greater than or equal to the actual learning level value, the target fast-forward interval is the not yet fully mastered interval;

[0049] When the actual teaching level value is less than the actual learning level value, the target fast-forward interval is the complete mastery interval;

[0050] When the result is a fully mastered interval, the target fast-forward interval is removed.

[0051] Compared with the existing technology, the beneficial effect of the present invention is that, by introducing large models and multi-dimensional data fusion technology, it realizes the intelligent splitting and personalized adjustment of video content; compared with the traditional video splitting method that only relies on video content features and fixed time information, this method first uses voice transcription and text analysis to accurately extract key information in the video, and combines the user's actual viewing behavior and learning progress feedback to realize dynamic evaluation of video content and introduce multiple indicators such as completion rate and repetition rate, accurately identifying the bottleneck areas that users encounter in the learning process, and thus executing targeted optimization measures such as subtitle correction, example sentence supplementation, anchor point playback and segmentation interval adjustment, eliminating redundant information while avoiding knowledge omissions, realizing dynamic customization and continuous optimization of teaching resources under different levels of teaching needs, and significantly improving learning efficiency and personalized teaching experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] Figure 1 This is a flowchart of a method for English video segmentation and learning based on a large model according to an embodiment of the present invention;

[0053] Figure 2 A logic decision diagram of a plurality of integrated videos obtained by integrating the difficulty quantification index according to an embodiment of the present invention;

[0054] Figure 3 A logical decision diagram for determining the type of a user's current learning bottleneck and determining the direction of video content adjustment in an embodiment of the present invention;

[0055] Figure 4 This is a logic decision diagram for executing the subtitle error adjustment procedure according to an embodiment of the present invention. DETAILED DESCRIPTION

[0056] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.

[0057] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.

[0058] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.

[0059] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.

[0060] See also Figure 1 As shown, it is a flow chart of a method for English video segmentation learning based on a large model according to an embodiment of the present invention. The invention provides a method for English video segmentation learning based on a large model, comprising:

[0061] Acquire input video data and process the data into standard input data, divide the standard input data into standards according to the requirements of the exam difficulty outline, and integrate the standard input data to obtain a plurality of integrated videos;

[0062] The integrated video is segmented according to the user's actual learning time requirements to obtain several integrated video segments, including fragmented videos and long videos;

[0063] For any of the integrated video segments, obtaining the user's actual video viewing information, analyzing the learning progress based on the actual video viewing information, determining the video content adjustment direction based on the analysis results, and determining the video content adjustment method based on the video content adjustment direction, and feeding back the adjustment results to the user end;

[0064] By introducing large models and multi-dimensional data fusion technology, intelligent splitting and personalized adjustment of video content are achieved; compared with traditional video splitting methods that only rely on video content features and fixed time information, this method first uses voice transcription and text analysis to accurately extract key information from the video, and combines users' actual viewing behavior and learning progress feedback to achieve dynamic evaluation of video content and introduce multiple indicators such as completion rate and repetition rate to accurately identify bottleneck areas that users encounter in the learning process, thereby implementing targeted optimization measures such as subtitle correction, example sentence supplementation, anchor point playback and segmentation interval adjustment, eliminating redundant information while avoiding knowledge omissions, and realizing dynamic customization and continuous optimization of teaching resources under different levels of teaching needs, significantly improving learning efficiency and personalized teaching experience.

[0065] Specifically, obtaining input video data and processing the data into standard input data includes,

[0066] Obtain video teaching text and voice transcription text;

[0067] The text content is divided into several individual sentences, and the parts of speech of words and phrases in any individual sentence are extracted and stored in a vocabulary database of corresponding difficulty. The corresponding test outline requirements or application ability requirements are determined according to the expected learning level, and several integrated videos are obtained based on the quantitative indicators of difficulty.

[0068] In this embodiment, the steps of obtaining the video teaching text and the voice transcription text are:

[0069] Extract teaching text and speech transcription text from teaching videos by integrating automatic speech recognition (ASR) system and optical character recognition (OCR) technology;

[0070] Among them, the integrated automatic speech recognition system recognizes the speech in the video in real time and generates preliminary text;

[0071] Optical character recognition technology recognizes subtitles or blackboard information embedded in the video;

[0072] In this embodiment, the steps of dividing the text content into individual sentences, extracting the parts of speech of words and phrases in the individual sentences and storing them in the vocabulary database of corresponding difficulty are as follows:

[0073] The text content is segmented based on linguistic rules and punctuation boundaries. For the sentence "The students listened attentively and thought about the problems actively," it is identified as two logically continuous but independent semantic units based on punctuation.

[0074] Extract words and phrases and mark their parts of speech. In the sentence "The quick brown fox jumps over the lazy dog.", "quick" is marked as an adjective, "fox" is a noun, and "jumps" is a verb. The system stores the words and their parts of speech information in the corresponding difficulty vocabulary database. Vocabulary that meets the requirements of the College English Test Band 4 is stored in the CET-4 difficulty database, and vocabulary that meets the requirements of the College English Test Band 6 is classified into the CET-6 difficulty database.

[0075] After standardization, the two types of data form unified standard input data, providing a numerical basis for subsequent processing.

[0076] See Figure 2 As shown, it is a logic decision diagram of several integrated videos obtained by integrating the difficulty quantification index according to an embodiment of the present invention;

[0077] Specifically, several integrated videos were obtained based on the quantitative indicators of difficulty, including:

[0078] Set a standard matching range and determine the actual matching degree of the corresponding examination syllabus requirements or application ability requirements based on the user's actual learning level expectations;

[0079] Determining the actual matching degree according to the standard matching interval, and determining the type of the integrated video content according to the determination result, including the first video type, the second video type, and the third video type;

[0080] In this embodiment, based on the learning level expectation set by the user and referring to the requirements of the corresponding examination syllabus, the difficulty of the extracted vocabulary and sentences is quantified and integrated. The calculation formula for the difficulty quantification index is:

[0081] D I =0.4*V ratio +0.35*S complexity +0.25*G difficulty

[0082] in,

[0083] D I : actual matching degree;

[0084] Vratio : The proportion of words beyond the basic vocabulary;

[0085] S complexity : The proportion of complex sentence structures;

[0086] G difficulty : frequency of occurrence of complex grammatical structures;

[0087] 0.4, 0.35 and 0.25 are V ratio 、S complexity and G difficulty The weight of

[0088] If the user sets the expected learning level as the CET-4 exam, the matching intervals are divided.

[0089] Among them, [0.4, 0.6] is the standard matching interval of the CET-4 test;

[0090] When D I In the range of [0.6, 0.8], the integrated video is the third video type, corresponding to the level 4 beyond-syllabus content;

[0091] When D I In the range of [0.4, 0.6], the integrated video is the second video type, corresponding to the level 4 standard difficulty content;

[0092] When D I When the score is lower than 0.4, the integrated video is of the first video type, corresponding to the basic difficulty content of level 4;

[0093] This quantification method not only provides data support for the generation of integrated videos, but also ensures that the generated video content accurately matches user expectations in terms of difficulty.

[0094] Specifically, segmenting the integrated video includes:

[0095] Segment the integrated video based on user feedback:

[0096] If the user feedback is that they need to learn quickly in fragmented time, the integrated video will be split into fragmented videos;

[0097] If the user feedback indicates that they want continuous deep learning, the integrated video will be split into long videos;

[0098] Among them, the fragmented video duration is within 3 to 7 minutes, which meets the needs of fast learning in fragmented time and uses fragmented time for rapid information transmission and review;

[0099] Long videos are between 40 and 60 minutes long, meeting the needs of continuous deep learning and suitable for continuous in-depth explanations and the complete presentation of complex concepts.

[0100] Different segmentations of integrated videos can reduce cognitive load in fragmented learning environments, facilitate repeated viewing, and enhance memory effects; in long-term learning environments, learners are provided with a systematic and coherent knowledge structure, which helps them understand and digest more complex content and meet the needs of continuous learning.

[0101] Specifically, obtaining the user's actual video viewing information and analyzing the learning progress based on the actual video viewing information includes:

[0102] Actual video viewing information includes repetition rate, completion rate, and repetition completion index. Video viewing information is evaluated and compared with preset thresholds.

[0103] If the repetition completion index meets the evaluation criteria, the user's learning progress is determined to match the difficulty level requirements, and feedback measures on the user's learning progress are implemented;

[0104] If the replay completion index does not meet the evaluation criteria, it is determined that the user's learning progress does not match the current difficulty level requirement. The user's current learning bottleneck type is determined and the direction of video content adjustment is determined;

[0105] The specific judgment criteria and calculation formulas for the repetition rate and completion rate data in this embodiment are as follows:

[0106] Repetition rate = (total duration of repeated content / total duration of video) × 100%;

[0107] Completion rate = (number of users who watched the video in full / total number of users who watched it) × 100%;

[0108] In order to more precisely reflect the impact of repetition and completion on learning outcomes, a comprehensive indicator, the repetition and completion index, can be introduced. The calculation formula for this index is:

[0109] Repeat completion index = α × repetition rate + β × (1-completion rate)

[0110] Among them, α and β are weight coefficients, which can be dynamically adjusted according to the learning difficulty;

[0111] Wherein, for the first video type, α=0.6, β=0.8;

[0112] Second video type α=0.5, β=0.5;

[0113] Second video type α=0.4, β=0.2;

[0114] Set the preset threshold = 0.4,

[0115] When the repetition completion index is greater than or equal to the preset threshold, the video has excessive repetition issues and needs feedback and video editing adjustments;

[0116] When the repeated broadcast index is less than the preset threshold, there is no duplication problem in the video content;

[0117] This division method not only quantifies the redundancy of video content, but also combines users' actual viewing data, providing a scientific basis for intelligent feedback and optimization adjustments.

[0118] See Figure 3 As shown, it is a logical decision diagram for determining the type of the user's current learning bottleneck and judging the direction of video content adjustment according to an embodiment of the present invention;

[0119] Specifically, determining the user's current learning bottleneck and determining the direction of video content adjustment include:

[0120] The learning bottleneck types include the first result, the second result, the third result and the fourth result;

[0121] When the learning bottleneck type is determined to be the first result, executing a subtitle error adjustment procedure;

[0122] When the learning bottleneck type is determined to be the second result, executing the example supplement procedure;

[0123] When the learning bottleneck type is determined to be the third result, executing the anchor point playback program;

[0124] When the learning bottleneck type is determined to be the fourth result, the segmentation interval adjustment program is executed

[0125] When the learning bottleneck type is determined to be the first result, the fragmented video focuses on excessive repetition;

[0126] When the learning bottleneck type is determined to be the second result, the key points of the long video are overly repeated;

[0127] When the learning bottleneck type is determined to be the third result, the fragmented video is excessively fast-forwarded to non-key points;

[0128] When the learning bottleneck type is determined to be the fourth result, the long video is excessively fast-forwarded without being the key point.

[0129] Refer to FIG4 , which is a logic decision diagram for executing the subtitle error adjustment program according to an embodiment of the present invention;

[0130] Specifically, executing the subtitle error adjustment procedure includes:

[0131] Read the real-time delay of subtitles and compare it with the standard delay.

[0132] When the real-time delay is less than or equal to the standard delay, there is no abnormality in the subtitles;

[0133] When the real-time delay is longer than the standard delay, the subtitles are abnormal;

[0134] If the judgment result is that there is no abnormality in the subtitles, the corresponding extension parameters are obtained through the large model analysis learning content, and the video length is extended according to the parameters;

[0135] If the judgment result is that the subtitles are abnormal, the subtitle display parameters are adjusted and the subtitle text content is verified, and the adjusted video viewing information is read and evaluated;

[0136] If the evaluation results do not meet the evaluation criteria and are excessively repetitive, the corresponding extension parameters are obtained through large-scale model analysis of the learning content, and the video duration is extended according to the parameters;

[0137] In this example, the standard delay is set to 100 milliseconds, and the open source tool LanguageTool is selected and combined with a Python script to verify the subtitle text content;

[0138] When the actual delay between subtitles and audio does not exceed 100 milliseconds, subtitle synchronization is normal;

[0139] When the actual delay between subtitles and audio exceeds 100 milliseconds, the subtitles will experience delay abnormalities;

[0140] After detecting subtitle anomalies, the system adjusts subtitle display parameters, including subtitle display start time, position, and font, and automatically verifies subtitle text to ensure that the text content is consistent with the voice content.

[0141] Set the standard repetition rate to 50%, and the parameter calculation formula is:

[0142] T = (actual repetition rate - standard repetition rate) × 0.4

[0143] In this embodiment, the actual repetition rate is 65%, T = 15 × 0.4 = 6 seconds;

[0144] After the adjustment is completed, the system will collect real-time viewing data again. If the evaluation result still does not meet the standards and the repetition rate is high, the above calculation process will be repeated, gradually extending the video length until it meets the evaluation standards or reaches the upper limit of the length.

[0145] While ensuring the accuracy of subtitle synchronization, the video length is dynamically adjusted according to user viewing behavior to improve the learning experience of fragmented key learning content.

[0146] Specifically, the example supplementary procedures include:

[0147] Determine the teaching difficulty level based on the type of video content currently being integrated, and extract the language logic template at that difficulty level;

[0148] Select several core words in the current integrated video content that match the teaching difficulty level from the corresponding difficulty vocabulary database;

[0149] According to the language logic template, the semantics between the core words are associated to generate candidate sentences, which are stored in the candidate sentence information database and then determined whether to call them according to the actual needs of the user;

[0150] In the context of learning attributive clauses, this embodiment selects the attributive clause pattern as the language logic template;

[0151] The vocabulary database is the College English Test Band 4 vocabulary database. The system extracts keywords that meet the difficulty requirements of Band 4 by matching the vocabulary that appears in the video content with the vocabulary in the vocabulary database.

[0152] The example sentence generation is based on the basic structure of the attributive clause in the template to form a reasonable sentence pattern.

[0153] Among them, the candidate sentences include video sentences and text sentences;

[0154] Video example sentences are extracted from video teaching text and images;

[0155] Text examples are extracted from speech transcription text;

[0156] After the system verifies the grammar and semantics of the two types of candidate sentences, it marks the newly added candidate sentences that have passed the screening in the repetition interval of the video to facilitate learners' targeted review and reference.

[0157] Specifically, executing the anchor point playback procedure includes:

[0158] Confirm the starting position of the repeat interval and set an anchor point at the starting position to quickly play back the repeated content from the anchor point;

[0159] Confirm user playback behavior within the standard judgment time to dynamically adjust the anchor point position;

[0160] Specifically, the split interval adjustment procedure includes:

[0161] Extract the playback information of the target fast-forward interval, evaluate the user's current learning level based on the playback information, and quantify it into an actual learning level value;

[0162] Extracting and integrating playback information of the video, evaluating the current teaching level of the video based on the playback information and quantifying it into an actual teaching level value;

[0163] Among them, the playback information includes the user's historical learning records and interactive feedback;

[0164] Compare the actual teaching level value with the actual learning level value,

[0165] When the actual teaching level value is greater than or equal to the actual learning level value, the target fast-forward interval is the not yet fully mastered interval;

[0166] When the actual teaching level value is less than the actual learning level value, the target fast-forward interval is the complete mastery interval;

[0167] Among them, when the result is the fully mastered interval, the target fast-forward interval is removed;

[0168] The target fast-forward interval is a video interval in which the corresponding actual operation is high-frequency fast-forward.

[0169] In this embodiment, the standard fast-forward operation frequency value is set to four times per minute, and the user's actual fast-forward operation frequency value is obtained and compared with the standard fast-forward operation frequency value.

[0170] When the user's actual fast-forward operation frequency value is greater than or equal to the standard fast-forward operation frequency value, the actual operation is high-frequency fast-forward;

[0171] When the user's actual fast-forward operation frequency value is less than the standard fast-forward operation frequency value, the actual operation is non-high-frequency fast-forward;

[0172] Read playback information and extract user historical learning records, including repetition rate, completion rate, and historical playback records; collect user interaction feedback data, including pause, mark, and active jump operations, to form an interaction feedback score;

[0173] In this embodiment, the data is processed and evaluated by the large model, and the specific steps are:

[0174] Read and normalize all data, and construct feature vectors containing dimensions such as repetition rate, completion rate, interactive feedback, and viewing time. Noise reduction and missing value filling techniques are used to ensure data quality.

[0175] The preprocessed feature vector is input into the neural network model pre-trained with English teaching input videos. The model comprehensively analyzes the historical playback records, repetition rate, completion rate and completion of after-class exercises, and outputs an actual learning level value L in the range of 0 to 100. user ,This value reflects the user’s understanding of the overall content of the current integrated video through the target fast-forward interval;

[0176] Extract text and video frame features from the integrated video content, construct a video teaching difficulty feature vector, including vocabulary difficulty, average sentence length, proportion of long and difficult sentences, and grammatical structure complexity, input the big data model, and output a video actual teaching level value L in the range of 0 to 100 video ,This value reflects the difficulty of integrating the overall content of the video into teaching;

[0177] The actual teaching level value L video and the actual learning level value L user Make a comparison,

[0178] When the actual teaching level value L video Greater than or equal to the actual learning level value L user When the target fast-forward interval is the interval that has not yet been fully mastered;

[0179] When the actual teaching level value L video Less than the actual learning level value L user When , the target fast-forward interval is the complete mastery interval;

[0180] Perform content removal on the target fast-forward interval that has been fully mastered, removing it from the integrated video and retaining the "not yet fully mastered" area as the key content for the user's subsequent review and learning;

[0181] The comprehensive judgment method based on big data models visualizes the user's learning level and the difficulty of video teaching by refining the steps of data collection, feature engineering, model reasoning and interval comparison, and realizes dynamic adjustment and personalized optimization of the target fast-forward interval, so that the teaching method can accurately meet the user's personalized needs.

[0182] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.

[0183] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.

Claims

1. A method for English video segmentation learning based on a large model, characterized by: include, Acquire input video data and process the data into standard input data, and integrate the standard input data according to the requirements of the exam difficulty outline to obtain a plurality of integrated videos; The integrated video is segmented according to the user's actual learning time requirements to obtain several integrated video segments, including fragmented videos and long videos; For any of the integrated video segments, the user's actual video viewing information is obtained, and the learning progress is analyzed based on the actual video viewing information. The video content adjustment direction is determined according to the analysis result, and the video content adjustment method is determined based on the video content adjustment direction, and the adjustment result is fed back to the user end.

2. The method for English video segmentation learning based on a large model according to claim 1 is characterized in that: Acquiring input video data and processing said data into standard input data comprises, Obtain video teaching text and voice transcription text; The text content is divided into several separate sentences, and the parts of speech of words and phrases in any separate sentence are extracted and stored in the vocabulary database of corresponding difficulty. The corresponding examination syllabus requirements or application ability requirements are determined according to the expected learning level, and several integrated videos are obtained based on the quantitative indicators of difficulty.

3. The method for English video segmentation learning based on a large model according to claim 2 is characterized in that: According to the quantitative indicators of difficulty, several integrated videos are obtained, including: Set a standard matching range and determine the actual matching degree of the corresponding examination syllabus requirements or application ability requirements based on the user's actual learning level expectations; The actual matching degree is determined according to the standard matching interval, and the type of the integrated video content is determined according to the determination result, including the first video type, the second video type, and the third video type.

4. The method for English video segmentation learning based on a large model according to claim 1 is characterized in that: Segmentation of integrated video includes: Segment the integrated video based on user feedback: If the user feedback is that they need to learn quickly in fragmented time, the integrated video will be split into fragmented videos; If the user feedback is to choose continuous deep learning needs, the integrated video will be split into long videos.

5. The method for English video segmentation learning based on a large model according to claim 1 is characterized in that: Obtain the user's actual video viewing information and analyze the learning progress based on the actual video viewing information, including: Actual video viewing information includes repetition rate, completion rate, and repetition completion index. Video viewing information is evaluated and compared with preset thresholds. If the repetition completion index meets the evaluation criteria, the user's learning progress is determined to match the difficulty level requirements, and feedback measures on the user's learning progress are implemented; If the repeated completion index does not meet the evaluation criteria, it is determined that the user's learning progress does not match the current difficulty progress requirements, the type of user's current learning bottleneck is determined, and the direction of video content adjustment is determined.

6. The method for English video segmentation learning based on a large model according to claim 5 is characterized in that: Determine the user's current learning bottleneck and determine the direction of video content adjustment, including: The learning bottleneck types include the first result, the second result, the third result and the fourth result; When the learning bottleneck type is determined to be the first result, executing a subtitle error adjustment procedure; When the learning bottleneck type is determined to be the second result, executing the example supplement procedure; When the learning bottleneck type is determined to be the third result, executing the anchor point playback program; When the learning bottleneck type is determined to be the fourth result, a segmentation interval adjustment procedure is executed.

7. The method for learning English video segmentation based on a large model according to claim 6 is characterized in that: Executing the subtitle error adjustment procedure includes, Read the real-time delay of subtitles and compare it with the standard delay. When the real-time delay is less than or equal to the standard delay, there is no abnormality in the subtitles; When the real-time delay is longer than the standard delay, the subtitles are abnormal; If the judgment result is that there is no abnormality in the subtitles, the corresponding extension parameters are obtained through the large model analysis learning content, and the video length is extended according to the parameters; If the judgment result is that the subtitles are abnormal, the subtitle display parameters are adjusted and the subtitle text content is verified, and the adjusted video viewing information is read and evaluated; If the evaluation results do not meet the evaluation criteria and are excessively repetitive, the corresponding extension parameters are obtained through large-scale model analysis of the learning content, and the video duration is extended according to the parameters.

8. The method for learning English video segmentation based on a large model according to claim 6 is characterized in that: Execution of example supplementary procedures includes, Determine the teaching difficulty level based on the type of video content currently being integrated, and extract the language logic template at that difficulty level; Select several core words in the current integrated video content that match the teaching difficulty level from the corresponding difficulty vocabulary database; The semantics between the core words are associated according to the language logic template to generate candidate sentences, which are stored in the candidate sentence information database and whether to call them is determined according to the actual needs of the user.

9. The method for learning English video segmentation based on a large model according to claim 6, characterized in that: Executing the anchor playback procedure includes, Confirm the starting position of the repeat interval and set an anchor point at the starting position to quickly play back the repeated content from the anchor point; Confirm user playback behavior within the standard judgment duration to dynamically adjust the anchor point position.

10. The method for English video segmentation learning based on a large model according to claim 6 is characterized in that: The execution of the split interval adjustment procedure includes: Extract the playback information of the target fast-forward interval, evaluate the user's current learning level based on the playback information, and quantify it into an actual learning level value; Extracting and integrating playback information of the video, evaluating the current teaching level of the video based on the playback information and quantifying it into an actual teaching level value; Among them, the playback information includes user historical learning records and interactive feedback; Compare the actual teaching level value with the actual learning level value, When the actual teaching level value is greater than or equal to the actual learning level value, the target fast-forward interval is the not yet fully mastered interval; When the actual teaching level value is less than the actual learning level value, the target fast-forward interval is the complete mastery interval; When the result is a fully mastered interval, the target fast-forward interval is removed.

Citation Information

Patent Citations

  • Video segmentation method and device

    CN111046839A

  • Diversified college piano course teaching system and construction method

    CN118071554A

  • Learning condition management method for course interaction video, electronic equipment and storage medium

    CN118313969A

  • Video processing method and device based on large model and ASR

    CN119110128A

  • Teaching video reconstruction method and system based on intelligent agent

    CN119271844A