An online training video recording method and system

The method uses a neural network to improve video clarity and highlight key content in online training videos, enhancing the learning experience by clearly identifying and emphasizing important moments.

CN119325013BActive Publication Date: 2025-07-15GUANGZHOU LANFAN INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411258892.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-10
Publication Date
2025-07-15
Estimated Expiration
2044-09-10

AI Technical Summary

Technical Problem

The existing video recording methods cannot improve the clarity of the video recording screen, nor can they highlight the key contents of the training process.

Method used

By obtaining recorded video images and adding timestamps, identifying the target subject, using the preset neural network model to extract features and performing denoising processing, combining high-definition recording videos, and inserting split-screen nodes and their sub-screens at the timestamp location, including auxiliary information for key content.

Benefits of technology

It significantly improves the clarity and information volume of videos, facilitates post-editing and positioning, enhances the viewing and interactivity of videos, and improves educational effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119325013B_ABST
    Figure CN119325013B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of communication technologies, and discloses a method and a system for recording an online training video. The method includes obtaining a recorded video image of a target subject, and adding time stamps to key content in the recorded video image; identifying the target subject in each video frame in the recorded video image to obtain a first video picture; extracting features of the first video picture through a preset neural network model, and performing denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image picture; combining the high-definition recorded video image pictures in chronological order to obtain a high-definition recorded video; inserting a split-screen node and its corresponding secondary picture into the high-definition recorded video according to the time stamps to obtain a video recording result, where the secondary picture includes auxiliary content corresponding to the key content. The recorded video obtained by this method has high clarity, prominent key content, and is convenient for trainees to learn.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of communication technologies, and in particular, to a method and system for recording online training videos. Background Art

[0002] With the continuous development and popularization of Internet technologies, online training, as a new learning method, has gradually been recognized and accepted by a large number of users. Usually, the training content is recorded during the training process, and the clarity of the recorded training video is directly related to the later review. Currently, screen recording software is usually used to record the online training process. Although this method can record the training content completely, it cannot improve the clarity of the recorded video image and cannot highlight the key content of the training process. Summary of the Invention

[0003] The present invention provides a method and system for recording online training videos, which solves the problems that the existing video recording methods cannot improve the clarity of the recorded video image and cannot highlight the key content of the training process.

[0004] To solve the above technical problems, a first aspect of the present invention provides a method for recording an online training video, including:

[0005] Obtaining a recorded video image of a target subject and adding time stamps to the key content in the recorded video image;

[0006] Identifying the target subject in each video frame in the recorded video image to obtain a first video picture;

[0007] Extracting features of the first video picture through a preset neural network model and performing denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image;

[0008] Combining the high-definition recorded video images in chronological order to obtain a high-definition recorded video;

[0009] Inserting split-screen nodes and their corresponding sub-pictures into the high-definition recorded video according to the time stamps to obtain a video recording result, where the sub-picture includes auxiliary content corresponding to the key content.

[0010] As one preferred solution, the preset neural network model includes a deep residual network layer, a first convolutional neural network layer, a recurrent neural network layer, a second convolutional neural network layer, a third convolutional neural network layer, and a fourth convolutional neural network layer.

[0011] As one preferred solution, extracting features of the first video picture through a preset neural network model and performing denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image, including:

[0012] Extract the image features of the first video frame through the depth residual network layer, and extract the spatial features of the first video frame through the first convolutional neural network layer;

[0013] Process the spatial features of the first video frame in the current video frame and the spatial features of the first video frame in the previous video frame through the recurrent neural network layer to obtain the temporal features of the first video frame in the current video frame;

[0014] Extract the state features of the first video frame through the second convolutional neural network layer, and calculate the contribution value of the first video frame according to the state features and the image features;

[0015] Fuse the temporal features and the spatial features of the corresponding video frame to obtain spatio-temporal features, and based on the spatio-temporal features, perform convolutional processing on the first video frame through the third convolutional neural network layer to obtain a convolutional matrix;

[0016] Perform residual processing on the convolutional matrix through the fourth convolutional neural network layer to obtain the first video frame after denoising processing, and combine the first video frame after denoising processing according to the contribution value to obtain a high-definition recorded video image frame.

[0017] As one of the preferred solutions, the time stamp includes a start time and an end time; wherein,

[0018] Inserting a split screen node and its corresponding sub-picture into the high-definition recorded video according to the time stamp to obtain a video recording result includes:

[0019] Obtain a sub-picture and its corresponding sub-video frame, and set the time length of the sub-picture through the time stamp;

[0020] Extract video frames from the high-definition recorded video based on the time stamp and use them as the main video frames in the main picture;

[0021] Create a new video frame on the frame corresponding to the time stamp, and add the main video frame and the sub-video frame to obtain a new fused video frame;

[0022] Set a split screen node at the start time, and insert the new fused video frame according to the split screen node according to the time stamp to replace the original video frame to obtain a video recording result.

[0023] As one of the preferred solutions, creating a new video frame on the frame corresponding to the time stamp and adding the main video frame and the sub-video frame to obtain a new fused video frame includes:

[0024] Create a new video frame on the frame corresponding to the timestamp, and divide the new video frame into regions to obtain a first region and a second region;

[0025] Add the main video frame to the first region and add the secondary video frame to the second region to obtain a new fused video frame.

[0026] As one preferred solution, after inserting a split-screen node and its corresponding secondary picture into the high-definition recorded video according to the timestamp and obtaining the video recording result, it further includes:

[0027] Perform a condensation process on the video recording result to obtain a condensed video.

[0028] As one preferred solution, performing a condensation process on the video recording result to obtain a condensed video includes:

[0029] Identify each high-definition recorded video frame in the video recording result to obtain the second video picture corresponding to each high-definition recorded video frame;

[0030] Select target video pictures that meet the preset conditions from the second video pictures to obtain a target picture list;

[0031] Perform a cropping process on the video recording result according to the target picture list, obtain several video segments and combine them to obtain a condensed video.

[0032] As one preferred solution, selecting target video pictures that meet the preset conditions from the second video pictures to obtain a target picture list includes:

[0033] Calculate the feature similarity of the second video pictures, traverse each high-definition recorded video frame, and compare the feature similarity of the second video picture in the current high-definition recorded video frame with the feature similarity of the second video picture in the adjacent high-definition recorded video frame;

[0034] Take two second video pictures whose comparison results meet the preset comparison conditions and whose appearance duration reaches the preset duration threshold as the same target video picture, and combine the target video pictures whose pixels are not less than the preset pixel threshold to obtain a target picture list.

[0035] As one preferred solution, selecting target video pictures that meet the preset conditions from the second video pictures to obtain a target picture list includes:

[0036] Extract the color data of the second video pictures and remove the background color data from the color data to obtain content color data;

[0037] If the content color data belongs to a pre-constructed color database and the appearance duration of the second video frame corresponding to the content color data reaches a preset duration threshold, then this video frame is taken as the target video frame;

[0038] Combine the target video frames to obtain a target video frame list.

[0039] The second aspect of the present invention provides an online training video recording system, including:

[0040] A video image processing module, configured to obtain a recorded video image of a target subject and add time stamps to key content in the recorded video image;

[0041] A video frame recognition module, configured to recognize the target subject in each video frame within the recorded video image to obtain a first video frame;

[0042] A video image denoising module, configured to extract features of the first video frame through a preset neural network model and perform denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image frame;

[0043] A first video generation module, configured to combine the high-definition recorded video image frames in chronological order to obtain a high-definition recorded video;

[0044] A second video generation module, configured to insert split-screen nodes and their corresponding sub-frames in the high-definition recorded video according to the time stamps to obtain a video recording result, where the sub-frames include auxiliary content corresponding to the key content.

[0045] Compared with the prior art, the beneficial effects of the embodiments of the present invention are at least one of the following:

[0046] (1) In this application, a recorded video image of a target subject is captured, and time stamps are added at important content for subsequent processing; a preset neural network model is used to extract features of the target subject's frame in each frame of the video, and denoising processing is performed based on these features to improve the clarity of the video frame; the processed high-definition frames are recombined to generate a high-definition recorded video; according to the previously added time stamps, split-screen nodes and their sub-frames are inserted at the corresponding positions in the high-definition video, and these sub-frames contain auxiliary information of the key content, thereby enhancing the information volume and viewing pleasure of the video and improving the quality of the recorded video, and generating a high-definition and information-rich video recording result;

[0047] (2) Timestamps are added at key points in this application, which not only facilitates later editing and quick positioning but also provides an intuitive reference point for learners, making it convenient for reviewing and sharing specific content. Through the feature extraction and denoising processing technology of the neural network model, noise and interference in the video are effectively removed, significantly enhancing the video clarity and enabling learners to obtain a more delicate and realistic visual experience. The design of the split-screen nodes and their sub-images makes the video content presentation more flexible and diverse. It not only increases the layering of the video but also greatly enriches the video content, allowing for customized display according to the needs and interests of learners, thus enhancing the viewing interest and interactivity. It can effectively convey knowledge points and details in the fields of education, training, etc., helping learners better understand and master the content they have learned, thereby improving the educational effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] To more clearly illustrate the technical solutions of the present invention, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0049] Figure 1 is a flowchart of a method for recording an online training video provided by an embodiment of the present invention;

[0050] Figure 2 is a flowchart of step S3 provided by an embodiment of the present invention;

[0051] Figure 3 is a flowchart of step S5 provided by an embodiment of the present invention;

[0052] Figure 4 is a flowchart of step S53 provided by an embodiment of the present invention;

[0053] Figure 5 is a flowchart of step S6 provided by an embodiment of the present invention;

[0054] Figure 6 is a flowchart of step S62 provided by an embodiment of the present invention;

[0055] Figure 7 is a flowchart of step S62 provided by another embodiment of the present invention;

[0056] Figure 8 is a structural diagram of an online training video recording system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings and embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The purpose of providing these embodiments is to make the disclosure of the present invention more thorough and comprehensive. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0058] In the description of the present application, the terms "first", "second", "third", etc. are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", "third", etc. may explicitly or implicitly include one or more of such features. In the description of the present application, unless otherwise stated, the meaning of "a plurality" is two or more.

[0059] In the description of the present application, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the communication inside two components. The terms "vertical", "horizontal", "left", "right", "up", "down" and similar expressions used herein are only for the purpose of illustration, rather than indicating or implying that the system or component referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present invention. The term "and / or" used herein includes any and all combinations of one or more of the related listed items. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0060] In the description of the present application, it should be noted that unless otherwise defined, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs. The terms used in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances.

[0061] In one embodiment, as Figure 1 shown, the first aspect of the present invention provides an online training video recording method, including:

[0062] S1. Obtain the recorded video image of the target subject, and add timestamps to the key content in the recorded video image; wherein, the timestamp includes a start time and an end time; the target subject is the training files commonly used by trainers for assistance during the online training process, such as PPTs, documents, pictures, etc.; the recorded video image is usually recorded using tools with recording functions such as cameras; the key content is the content that needs to be learned and mastered by the learners during the training process. Usually, the key content will be emphasized or demonstrated by the trainer during the training process. Therefore, timestamps including the start time and the end time are added to identify the key content during recording for subsequent processing.

[0063] S2. Identify the target subject in each video frame of the recorded video image to obtain a first video frame.

[0064] Specifically, since the target subject can be text or a picture, but during the training display, for the convenience of the learners to view, the training files are all arranged in a certain way, and the recorded video image cannot be filled with the target subject. Therefore, it is necessary to identify and process the target subject in each video frame of the recorded video image to obtain the first video frame, that is, the image frame containing the key content.

[0065] In this application, first, the recorded video image is processed to obtain its corresponding grayscale image and pixel image respectively; then, the particle swarm algorithm is used to search and calculate the optimal grayscale threshold (that is, the optimal parameter obtained from this search process) for the grayscale image with the first preset parameters (the parameters include resolution, weight, learning factor, and speed), and based on the optimal grayscale threshold obtained from the previous search, the cooperative particle swarm algorithm is used to search for the grayscale image with the second preset parameters and update the optimal grayscale threshold. Repeat this update step until the preset search times are reached to obtain the final optimal grayscale threshold and use it to segment the original image to obtain several image blocks; then determine the boundary range of the pixel image and set a central seed point at the center position of the pixel image, and randomly select multiple initial seed points in the pixel image, and use the growth criterion of merging adjacent pixels with pixel value differences less than the pixel threshold between the central seed point and the initial seed points and pixels with color distances (such as Euclidean distance, Manhattan distance, etc.) less than the color threshold between pixels to grow until there are no new pixels, obtaining several segmented regions; finally, compare the several image blocks and the several segmented regions, take the overlapping part as the initial video frame, and use the boundary of the segmented region to optimize the boundary of the initial video frame to obtain the first video frame.

[0066] This application uses a multi-particle algorithm and a region growing algorithm to identify and segment images, combines global search and local optimization, integrates different modality information from the image through the multi-particle algorithm, and constructs a more accurate similarity criterion using this information through the growing algorithm, so as to divide the image region more precisely, improve the segmentation accuracy and robustness, and also improve the efficiency and convergence speed of the algorithm and enhance the self-adaptability of the algorithm.

[0067] S3. Extract the features of the first video frame through a preset neural network model, and perform denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image frame;

[0068] In one embodiment, the preset neural network model includes a deep residual network layer, a first convolutional neural network layer, a recurrent neural network layer, a second convolutional neural network layer, a third convolutional neural network layer, and a fourth convolutional neural network layer; specifically, the preset neural network model can be trained through the original and denoised initial video image samples. Input the original video image samples into the preset neural network model, output the target video image samples after denoising, calculate the error loss between it and the denoised initial video image samples, and train the preset neural network model through this error loss to obtain the trained preset neural network model, and use it for denoising in the solution of this application to improve the clarity of the video image.

[0069] In one embodiment, step S3 is as Figure 2 shown and includes:

[0070] S31. Extract the image features of the first video frame through the deep residual network layer, and extract the spatial features of the first video frame through the first convolutional neural network layer;

[0071] S32. Process the spatial features of the first video frame in the current video frame and the spatial features of the first video frame in the previous video frame through the recurrent neural network layer to obtain the temporal features of the first video frame in the current video frame;

[0072] S33. Extract the state features of the first video frame through the second convolutional neural network layer, and calculate the contribution value of the first video frame according to the state features and the image features;

[0073] S34. Fuse the temporal features and the spatial features of the corresponding video frame to obtain spatio-temporal features, and based on the spatio-temporal features, perform convolutional processing on the first video frame through the third convolutional neural network layer to obtain a convolutional matrix;

[0074] S35. Perform residual processing on the convolutional matrix through the fourth convolutional neural network layer to obtain a first video frame after denoising processing, and combine the first video frame after denoising processing according to the contribution value to obtain a high-definition recorded video image frame.

[0075] Specifically, the present application uses a deep residual network layer to extract the image features of the video frame corresponding to the video frame, which can quickly extract the important information in the frame image; extracts the spatial features of the first video frame through the first convolutional neural network layer, and the spatial features are used to represent the positions of the noise elements in the video frame, and uses the characteristics of the recurrent neural network layer to summarize and extract the temporal features of the video frame corresponding to the current video frame based on the spatial features of the video frames corresponding to the current and previous video frames, and so on, to obtain the temporal features of the video frames corresponding to the entire video frame sequence according to the temporal features of the video frame corresponding to the current video frame; then by fusing the temporal features and spatial features of the video images corresponding to the video frames, the temporal and spatial correlations between adjacent frame images are established to more help identify and extract the corresponding target subject and remove noise; at the same time, considering that the proportions of different video frames are different, and the proportions of the various video components in the same video frame are also different, the present application calculates the contribution value of the corresponding video frame through the image features and hidden state features corresponding to the video frame in the video frame, and combines the denoised video frames in chronological order according to the proportion corresponding to the contribution value to obtain a high-definition recorded video image frame.

[0076] The present application removes noise through a trained preset neural network model, significantly improving the quality of the video image. Viewers and learners can more clearly see the details in the image, thus enhancing the overall visual experience. And the noise reduction process helps to accurately capture and display the important information in the video image, reducing the interference and confusion of noise to the image details, so as to better summarize and highlight the key content of the training process.

[0077] S4. Combine the high-definition recorded video image frames in chronological order to obtain a high-definition recorded video; specifically, the high-definition recorded video can be used as video material, and through video editing software such as Adobe Premiere Pro, FinalCut Pro, Wondershare Filmora, etc., after operations such as adding the video material to the timeline, adjusting the order of the video material, editing and adjusting the video, and setting the export parameters, a high-definition recorded video is generated.

[0078] S5. Insert split-screen nodes and their corresponding sub-images into the high-definition recorded video according to the time stamps to obtain a video recording result, and the sub-images include the auxiliary content corresponding to the key content;

[0079] In one embodiment, step S5 is as Figure 3 shown and includes:

[0080] S51, obtaining a secondary picture and its corresponding secondary video frame, and setting the time length of the secondary picture by the timestamp; specifically, the secondary picture contains the auxiliary file content corresponding to the training file to further supplement the training file, which may be annotations, small window videos or other auxiliary content, etc. The start time in the timestamp is used as the start time of the secondary picture, and the end time in the timestamp is used as the end time of the secondary picture, so that the secondary picture can be played synchronously with the main picture.

[0081] S52, extracting a video frame from the high-definition recorded video based on the timestamp and using it as a main video frame in the main picture;

[0082] S53, creating a new video frame on the frame corresponding to the timestamp and adding the main video frame and the auxiliary video frame to obtain a new fused video frame;

[0083] In one embodiment, step S53 is as follows: Figure 4 As shown, including:

[0084] S531, creating a new video frame on the frame corresponding to the timestamp, and dividing the new video frame into regions to obtain a first region and a second region;

[0085] S532: Add the main video frame to the first area, and add the auxiliary video frame to the second area to obtain a new fused video frame.

[0086] Specifically, the present application merges the main and sub-video frames by creating a new fused video frame, and adds the corresponding main and sub-video frames to the corresponding areas according to the timestamp, and displays the contents of two or more video sources in the same frame, which can enrich the information content of the video and enable learners to obtain more visual information at the same time, thereby enhancing the expressiveness and attractiveness of the video; when dividing the area, a unique video layout can be designed as needed to achieve personalized visual effects; through area division and content fusion to optimize the presentation of video content, important content is more prominent, and secondary information is reasonably supplemented, thereby improving the overall quality and viewing experience of the video, so that learners can effectively learn the key content of the training process.

[0087] S54, setting a split screen node at the start time, and inserting the new fused video frame according to the split screen node according to the timestamp to replace the original video frame, to obtain a video recording result;

[0088] Specifically, in the present application, a split-screen node is set at a specified time point to mark the starting position where the split-screen operation is to be performed, and the exact positions of one or more video frames are correspondingly determined according to the split-screen node, and the original video frames are replaced with new fused video frames at these positions to generate a video recording result containing new content. By inserting new fused video frames, the present application can introduce additional information or perspectives into the high-definition recorded video, thereby greatly enhancing the richness and diversity of the video content; it can quickly add or modify content in the high-definition recorded video without reprocessing the entire video file, thus improving the efficiency of video production.

[0089] In one embodiment, after step S5, the method further includes:

[0090] S6. Perform condensation processing on the video recording result to obtain a condensed video;

[0091] In one embodiment, step S6 is as Figure 5 shown and includes:

[0092] S61. Identify each high-definition recorded video frame in the video recording result to obtain the second video picture corresponding to each high-definition recorded video frame;

[0093] S62. Select target video pictures that meet preset conditions from the second video pictures to obtain a target picture list;

[0094] S63. Perform cropping processing on the video recording result according to the target picture list, obtain several video segments and combine them to obtain a condensed video.

[0095] Specifically, the recording of the online training process is for the convenience of re-learning by learners in the later stage. However, not all parts of the entire training process are necessarily key learning content. Learners view the playback video to obtain a lot of useful information. Therefore, after obtaining the video recording result, the present application also performs condensation processing on it to extract the key content, which is convenient for improving learning efficiency. Among them, the method for identifying the second video picture can refer to the identification process of the first video picture, which will not be elaborated here; and the preset conditions include that the appearance duration reaches the duration threshold. Generally, key content often requires trainers to focus on explaining, so more time will be spent on this part. By statistically analyzing the duration of key content in previous training videos, the shortest duration threshold can be evaluated. Video pictures exceeding this value are target video pictures to form a target video picture list. The video segments with target video pictures can be cropped and combined in chronological order to obtain a condensed video.

[0096] By selecting target video frames that meet the preset conditions, this application can remove redundant or unnecessary parts in the video, making the final condensed video content more compact and refined, directly focusing on the information that learners need to focus on for learning, so as to improve the viewing efficiency. Moreover, for the condensed video obtained through cropping, its file size will be greatly reduced compared with the original video recording result, which can not only save storage space and reduce storage costs, but also reduce the bandwidth required during video transmission and improve the transmission efficiency.

[0097] In one embodiment, step S62 is as Figure 6 shown and includes:

[0098] S6211. Calculate the feature similarity of the second video frame, traverse each high-definition recorded video frame, and compare the feature similarity of the second video frame in the current high-definition recorded video frame with the feature similarity of the second video frame in the adjacent high-definition recorded video frame;

[0099] S6212. Take two second video frames whose comparison results meet the preset comparison conditions and whose appearance duration reaches the preset duration threshold as the same target video frame, and combine the target video frames with pixel values not less than the preset pixel threshold to obtain a target frame list.

[0100] Specifically, the feature similarity of video frames can be calculated through SSIM based on structural similarity. By calculating the brightness, contrast, and structural information of video frames, the similarity between video frames can be evaluated more comprehensively; the preset comparison condition is the similarity between video frames. It should be noted that the preset values and threshold values in this text are all set according to the actual situation and actual needs, and no specific limitations are made here. This application determines the target video frame through the feature similarity of video frames, and eliminates video frames with insufficient appearance duration and too small pixel values, which can effectively avoid taking useless video frames as target video frames for video condensation, thereby improving the video condensation efficiency and providing better and useful video image data for the later learning of learners.

[0101] In another embodiment, step S62 is as Figure 7 shown and includes:

[0102] S6221. Extract the color data of the second video frame, and remove the background color data from the color data to obtain content color data;

[0103] S6222. If the content color data belongs to a pre-constructed color database and the appearance duration of the second video frame corresponding to the content color data reaches the preset duration threshold, then take this video frame as the target video frame;

[0104] S6223. Combine the target video frames to obtain a list of target video frames.

[0105] Specifically, during online training, trainers usually highlight or interpret key points in the training documents being presented. After removing the color of the video frames corresponding to the training documents themselves, the remaining color data is the color data corresponding to the text added later during direct presentation of the training documents. Since the text added later is used to highlight key points and for differentiation, it has a significant color difference from the color of the training documents themselves, and the two cannot be converted into each other. Moreover, the more important the content, the more text is added, thus making the accuracy of the obtained target video frames extremely high. Among them, the pre-constructed color database includes the annotation colors commonly used during training, and usually combines the target video frames in chronological order. By setting a threshold for the appearance duration of the second video frame corresponding to the content color data in this application, it is possible to further filter out the frames that continuously appear in the video and are important, which helps to remove short and unimportant frames and only retain the key learning content valuable to the learners.

[0106] In the embodiment of this application, in view of the problem that the existing video recording method cannot improve the clarity of the video recording frames and cannot highlight the key content during the training process, a method for recording online training videos is designed. It realizes obtaining a recorded video image of the target subject and adding timestamps to the key content in the recorded video image; identifying the target subject in each video frame of the recorded video image to obtain the first video frame; extracting the features of the first video frame through a preset neural network model and performing denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image frame; combining the high-definition recorded video image frames in chronological order to obtain a high-definition recorded video; inserting split-screen nodes and their corresponding sub-frames into the high-definition recorded video according to the timestamps to obtain the video recording result, and the sub-frame includes the auxiliary content corresponding to the key content. This technical solution enhances the information content and viewing experience of the video and improves the quality of the recorded video, generating a high-definition and information-rich video recording result.

[0107] It should be noted that although the steps in the above flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders.

[0108] In another embodiment, as Figure 8 shown, the second aspect of the present invention provides an online training video recording system, including:

[0109] The video image processing module 10 is used to obtain the recorded video image of the target subject and add time stamps to the key content in the recorded video image;

[0110] The video frame recognition module 20 is used to recognize the target subject in each video frame of the recorded video image to obtain a first video frame;

[0111] The video image denoising module 30 is used to extract the features of the first video frame through a preset neural network model and perform denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image frame;

[0112] The first video generation module 40 is used to combine the high-definition recorded video image frames in chronological order to obtain a high-definition recorded video;

[0113] The second video generation module 50 is used to insert split-screen nodes and their corresponding secondary images into the high-definition recorded video according to the time stamps to obtain a video recording result, and the secondary images include auxiliary content corresponding to the key content.

[0114] It should be noted that each module in the above-mentioned online training video recording system can be implemented in whole or in part by software, hardware, and their combinations. The above-mentioned modules can be embedded in the processor of the computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above-mentioned modules. For the specific limitations of an online training video recording system, refer to the limitations of an online training video recording method in the above text. The two have the same functions and effects and will not be elaborated here.

[0115] In summary, the present invention relates to the field of communication technology, and discloses an online training video recording method and system. The method includes obtaining a recorded video image of a target subject and adding time stamps to the key content in the recorded video image; recognizing the target subject in each video frame of the recorded video image to obtain a first video frame; extracting the features of the first video frame through a preset neural network model and performing denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image frame; combining the high-definition recorded video image frames in chronological order to obtain a high-definition recorded video; inserting split-screen nodes and their corresponding secondary images into the high-definition recorded video according to the time stamps to obtain a video recording result, and the secondary images include auxiliary content corresponding to the key content; solving the problem that the existing video recording method cannot improve the clarity of the video recording screen and cannot highlight the key content of the training process, and facilitating the later learning of the trainees.

[0116] Each embodiment in this specification is described in a progressive manner. For the parts that are the same or similar in each embodiment, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the relevant content. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as within the scope described in this specification.

[0117] The above embodiments only represent several preferred embodiments of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the technical principle of the present invention, several improvements and substitutions can be made, and these improvements and substitutions should also be regarded as within the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the protection scope of the claims.

Claims

1. A method for recording online training videos, characterized in that, Including: Obtain the recorded video image of the target subject, and add timestamps to the key content in the recorded video image; Identify the target subject in each video frame of the recorded video image to obtain a first video picture; Extract the features of the first video picture through a preset neural network model, and perform denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image picture; Combine the high-definition recorded video image pictures in chronological order to obtain a high-definition recorded video; Insert split-screen nodes and their corresponding sub-pictures into the high-definition recorded video according to the timestamps to obtain a video recording result, where the sub-pictures include auxiliary content corresponding to the key content; Perform condensation processing on the video recording result to obtain a condensed video; The identifying the target subject in each video frame of the recorded video image to obtain a first video picture includes: Process the recorded video image to obtain a corresponding grayscale image and pixel image; use the particle swarm algorithm to search and calculate the optimal grayscale threshold for the grayscale image with the first preset parameter, and based on the optimal grayscale threshold obtained from the previous search, use the cooperative particle swarm algorithm to search the grayscale image with the second preset parameter and update the optimal grayscale threshold. Repeat the update step until the preset search times are reached to obtain the final optimal grayscale threshold and use the final optimal grayscale threshold to segment the original image to obtain several image blocks; determine the boundary range of the pixel image and set a central seed point at the center position of the pixel image, randomly select multiple initial seed points in the pixel image, and use the growth criterion of merging adjacent pixels with a pixel value difference less than the pixel threshold and pixels with a color distance less than the color threshold between pixels with the central seed point until there are no new pixels to stop, obtaining several segmentation regions; finally, compare the several image blocks and several segmentation regions, take the overlapping part as the initial video picture, and optimize the boundary of the initial video picture with the boundary of the segmentation region to obtain the first video picture; The performing condensation processing on the video recording result to obtain a condensed video includes: Identify each high-definition recorded video frame in the video recording result to obtain a second video picture corresponding to each high-definition recorded video frame; Select target video pictures that meet the preset conditions from the second video pictures to obtain a target picture list; Perform cropping processing on the video recording result according to the target picture list to obtain several video segments and combine them to obtain a condensed video; The selecting target video pictures that meet the preset conditions from the second video pictures to obtain a target picture list includes: Calculate the feature similarity of the second video picture, traverse each high-definition recorded video frame, and compare the feature similarity of the second video picture located in the current high-definition recorded video frame with the feature similarity of the second video picture located in the adjacent high-definition recorded video frame; Take two second video frames whose comparison results meet the preset comparison conditions and the appearance duration reaches the preset duration threshold as the same target video frame, and combine the target video frames with pixels not less than the preset pixel threshold to obtain a target frame list; Or extract the color data of the second video frame, and remove the background color data from the color data to obtain content color data; If the content color data belongs to a pre-constructed color database and the appearance duration of the second video frame corresponding to the content color data reaches the preset duration threshold, then regard this video frame as a target video frame; Combine the target video frames to obtain a target video frame list; The time stamp includes a start time and an end time; Among them, the inserting split-screen nodes and their corresponding sub-pictures into the high-definition recorded video according to the time stamp to obtain a video recording result includes: Obtain a sub-picture and its corresponding sub-video frame, and set the time length of the sub-picture through the time stamp; Based on the time stamp, extract video frames from the high-definition recorded video and use them as the main video frames in the main picture; Create a new video frame on the frame corresponding to the time stamp, and add the main video frame and the sub-video frame to obtain a new fused video frame; Set a split-screen node at the start time, and insert the new fused video frame according to the time stamp according to the split-screen node to replace the original video frame to obtain a video recording result.

2. The online training video recording method according to claim 1, wherein The preset neural network model includes a deep residual network layer, a first convolutional neural network layer, a recurrent neural network layer, a second convolutional neural network layer, a third convolutional neural network layer, and a fourth convolutional neural network layer.

3. The online training video recording method according to claim 2, wherein The extracting the features of the first video frame through a preset neural network model and denoising the recorded video image based on the features to obtain a high-definition recorded video image frame includes: Extract the image features of the first video frame through the deep residual network layer, and extract the spatial features of the first video frame through the first convolutional neural network layer; Process the spatial features of the first video frame in the current video frame and the spatial features of the first video frame in the previous video frame through the recurrent neural network layer to obtain the temporal features of the first video frame in the current video frame; Extract the state features of the first video frame through the second convolutional neural network layer, and calculate the contribution value of the first video frame according to the state features and the image features; Fuse the temporal features and the spatial features of the corresponding video frames to obtain spatio-temporal features, and based on the spatio-temporal features, perform convolutional processing on the first video frame through the third convolutional neural network layer to obtain a convolutional matrix; Perform residual processing on the convolutional matrix through the fourth convolutional neural network layer to obtain the first video frame after denoising processing, and combine the first video frame after denoising processing according to the contribution value to obtain a high-definition recorded video image frame.

4. A method for recording an online training video according to claim 1, characterized in that, The creating a new video frame on the frame corresponding to the time stamp, and adding the main video frame and the sub-video frame to obtain a new fused video frame includes: On the frame corresponding to the timestamp, create a new video frame and divide the new video frame into regions to obtain a first region and a second region; Add the main video frame to the first region and add the secondary video frame to the second region to obtain a new fused video frame.

5. An online training video recording system, which is used to implement the online training video recording method as described in claim 1, and is characterized in that, It includes: A video image processing module, configured to obtain a recorded video image of a target subject and add a timestamp to key content in the recorded video image; A video frame recognition module, configured to recognize the target subject in each video frame in the recorded video image to obtain a first video frame; A video image denoising module, configured to extract features of the first video frame through a preset neural network model and perform denoising processing on the recorded video image based on the features to obtain a high-definition recorded video image; A first video generation module, configured to combine the high-definition recorded video images in chronological order to obtain a high-definition recorded video; A second video generation module, configured to insert a split-screen node and its corresponding secondary screen in the high-definition recorded video according to the timestamp to obtain a video recording result, where the secondary screen includes auxiliary content corresponding to the key content.

Citation Information

Patent Citations

  • Video denoising method, device and equipment and storage medium

    CN112686828A

  • Multimedia course online editing and making method

    CN114173201A