Video processing method and device, medium, electronic equipment and product
By de-redundant processing of teaching videos and adding target annotation data, the problem of insufficient video processing in the existing technology is solved, the refined processing of video content and the support of diversified annotation needs is achieved, and the visual effect and practicality of teaching videos are improved.
Patent Information
- Application Number
- CN202510043602.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-05-23
AI Technical Summary
Existing video processing software cannot accurately identify the redundant parts when processing teaching videos, resulting in the processed videos still containing a large amount of invalid information, unable to effectively highlight key operational steps or key information of teaching needs, and cannot flexibly support diversified teaching video annotation needs.
By obtaining multiple initial video frames in the pending video, it is deredundantly processed to generate key video frames, which contain teaching requirements information. Then obtain the user's target teaching requirements information, identify the target location in the key video frame, and add labeled data of the target teaching requirements information at that location to generate the target video.
Effectively remove redundant information in the pending video, highlight key information of teaching needs, improve the visual effect of teaching videos, and flexibly support diversified teaching video labeling needs, and improve the practicality and user experience of teaching videos.
Smart Images

Figure CN120034704A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of multimedia technology, and in particular to a video processing method, device, medium, electronic equipment and product. Background Art
[0002] Existing video processing software cannot accurately identify redundant parts when processing teaching videos, resulting in the processed videos still containing a large amount of invalid information and being unable to effectively highlight key operating steps or key information of teaching needs, resulting in insufficient visual effects. In addition, different teaching videos may require different annotation methods, and existing video processing software cannot flexibly support diverse needs, limiting the customization and personalization of teaching videos. Therefore, teaching videos processed by existing video processing software still have the problem of insufficient practicality, which not only affects the learners' experience, but also hinders the effective dissemination of educational content. Summary of the invention
[0003] In order to overcome the problems existing in the related art, the present disclosure provides a video processing method, device, medium, electronic device and product.
[0004] According to a first aspect of an embodiment of the present disclosure, a video processing method is provided, the method comprising: Obtain multiple initial video frames in the video to be processed; Performing redundancy removal processing on the plurality of the initial video frames to generate key video frames, wherein the key video frames include a plurality of teaching demand information, wherein the teaching demand information includes at least one of teaching content information, teaching target information and teaching method information; Acquire target teaching demand information of the user, and identify the target position corresponding to the target teaching demand information in the key video frame; The target annotation data of the target teaching requirement information is added to the target position in the key video frame to generate a target video.
[0005] Optionally, the obtaining a plurality of initial video frames in the video to be processed includes: The video to be processed is divided to obtain at least one divided segment, and a preset number of video frames are respectively extracted from the at least one divided segment to obtain the multiple initial video frames.
[0006] Optionally, performing redundancy removal processing on the plurality of initial video frames to generate key video frames includes: In the case where it is determined that there are two adjacent initial video frames that match, determining the matching initial video frames as redundant video frames; In the case of determining that there is one redundant video frame, taking the redundant video frame as a key video frame; When it is determined that there are a plurality of redundant video frames, the redundant video frame with the highest definition among the plurality of redundant video frames is used as a key video frame.
[0007] Optionally, determining that there are two adjacent initial video frames that match each other includes: Obtaining a pixel value of each pixel in a plurality of the initial video frames; In the case where it is determined that the pixel difference between the pixel values corresponding to each pixel in two consecutive initial video frames is less than or equal to a preset pixel threshold, it is determined that the two consecutive initial video frames match.
[0008] Optionally, adding target annotation data of the target teaching requirement information at the target position in the key video frame to generate a target video includes: Performing image enhancement processing on the key video frame to generate an enhanced video frame; The target annotation data of the target teaching requirement information is added to the target position corresponding to the target teaching requirement information in the enhanced video frame to generate a target video.
[0009] Optionally, the method further comprises: Obtaining the video format requirement of the user; The video format requirement is used as the output video format of the target video.
[0010] According to a second aspect of an embodiment of the present disclosure, a video processing device is provided, the device comprising: A first acquisition module is used to acquire a plurality of initial video frames in the video to be processed; A first generating module is used to perform redundancy removal processing on the plurality of the initial video frames to generate key video frames, wherein the key video frames include a plurality of teaching demand information, and the teaching demand information includes at least one of teaching content information, teaching target information and teaching method information; A second acquisition module is used to acquire the user's target teaching demand information and identify the target position corresponding to the target teaching demand information in the key video frame; The second generating module is used to add the target annotation data of the target teaching requirement information at the target position in the key video frame to generate the target video.
[0011] Optionally, the first acquisition module is further used to divide the video to be processed to obtain at least one divided segment, and extract a preset number of video frames from the at least one divided segment to obtain the multiple initial video frames.
[0012] Optionally, the first generating module is further used to, when it is determined that there are two adjacent initial video frames that match, determine that the matching initial video frame is a redundant video frame; when it is determined that there is one redundant video frame, use the redundant video frame as a key video frame; when it is determined that there are multiple redundant video frames, use the redundant video frame with the highest clarity among the multiple redundant video frames as the key video frame.
[0013] Optionally, the first generating module is also used to obtain the pixel value of each pixel in the multiple initial video frames; when it is determined that the pixel difference between the pixel values corresponding to each pixel in two consecutive initial video frames is less than or equal to a preset pixel threshold, it is determined that the two consecutive initial video frames match.
[0014] Optionally, the second generating module is also used to perform image enhancement processing on the key video frame to generate an enhanced video frame; add target annotation data of the target teaching requirement information at the target position corresponding to the target teaching requirement information in the enhanced video frame to generate a target video.
[0015] Optionally, the device further comprises: The output module is used to obtain the video format requirement of the user; and use the video format requirement as the output video format of the target video.
[0016] According to a third aspect of an embodiment of the present disclosure, there is provided a non-temporary computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods described in the first aspect are implemented.
[0017] According to a fourth aspect of an embodiment of the present disclosure, there is provided an electronic device, including: a memory having a computer program stored thereon; A processor is used to execute the computer program in the memory to implement the steps of any one of the methods in the first aspect.
[0018] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which implements the steps of any one of the methods in the first aspect when executed by a processor.
[0019] The above technical scheme obtains multiple initial video frames in the video to be processed; performs redundancy removal processing on the multiple initial video frames to generate key video frames, wherein the key video frames include multiple teaching demand information, wherein the teaching demand information includes at least one of teaching content information, teaching target information and teaching method information; obtains the user's target teaching demand information, and identifies the target position corresponding to the target teaching demand information in the key video frame; adds the target annotation data of the target teaching demand information to the target position in the key video frame to generate the target video. In this way, by performing redundancy removal processing on the initial video frames to obtain key video frames, redundant information in the video to be processed can be removed, thereby effectively highlighting the key information of the teaching demand and improving the visual effect of the teaching video; combining the user's target teaching demand information, adding the target annotation data of the target teaching demand information to the target position of the key video frame to generate the target video, different target annotation data can be added for different target teaching demand information of the user, thereby meeting the annotation requirements in different teaching scenarios, and thus effectively improving the practicality of the teaching video and the user's experience.
[0020] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings are used to provide a further understanding of the present disclosure and constitute a part of the specification. Together with the following specific embodiments, they are used to explain the present disclosure but do not constitute a limitation of the present disclosure. In the accompanying drawings: Figure 1 is a flow chart of a video processing method according to an exemplary embodiment; Figure 2 is based on Figure 1 The embodiment shows a flow chart of a video processing method; Figure 3 is based on Figure 2 A flowchart of a video processing method shown in the illustrated embodiment; Figure 4 is based on Figure 1 A flowchart of another video processing method shown in the illustrated embodiment; Figure 5 is based on Figure 1 A flowchart of another video processing method shown in the illustrated embodiment; Figure 6 is a block diagram of a video processing device according to an exemplary embodiment; Figure 7 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0022] The specific implementation of the present disclosure is described in detail below in conjunction with the accompanying drawings. It should be understood that the specific implementation described herein is only used to illustrate and explain the present disclosure, and is not used to limit the present disclosure.
[0023] Before introducing the specific implementation methods of the present disclosure in detail, the application scenarios of the present disclosure are first described as follows. The present disclosure can be applied to the application scenario of video processing of teaching videos. When processing teaching videos, the existing video processing software cannot accurately identify redundant parts, resulting in the processed video still containing a large amount of invalid information, and cannot effectively highlight the key operating steps or key information of teaching needs, resulting in insufficient performance in terms of visual effects. In addition, different teaching videos may require different annotation methods, and the existing video processing software cannot flexibly support diverse needs, which limits the customization and personalization of teaching videos. Therefore, the teaching videos processed by the existing video processing software still have the problem of insufficient practicality, which not only affects the learners' experience, but also hinders the effective dissemination of educational content.
[0024] In order to solve the above technical problems, the present disclosure provides a video processing method, by obtaining multiple initial video frames in a video to be processed; performing redundancy removal processing on the multiple initial video frames to generate a key video frame, the key video frame includes multiple teaching demand information, the teaching demand information includes at least one of teaching content information, teaching target information and teaching method information; obtaining the user's target teaching demand information, and identifying the target position corresponding to the target teaching demand information in the key video frame; adding the target annotation data of the target teaching demand information at the target position in the key video frame to generate a target video. In this way, by performing redundancy removal processing on the initial video frame to obtain the key video frame, the redundant information in the video to be processed can be removed, so that the key information of the teaching demand can be effectively highlighted and the visual effect of the teaching video can be improved; combining the user's target teaching demand information, adding the target annotation data of the target teaching demand information at the target position of the key video frame to generate the target video, and different target annotation data can be added according to the user's different target teaching demand information, and the diversified needs can be flexibly supported, so that the annotation needs in different scenarios can be better met, and the practicality of the teaching video and the user's experience can be effectively improved.
[0025] The specific implementation of the present disclosure will be described in detail below with reference to the specific drawings.
[0026] Figure 1 is a flowchart of a video processing method according to an exemplary embodiment. Figure 1 As shown, the method comprises the following steps: Step 101: Acquire multiple initial video frames in the video to be processed.
[0027] The video to be processed may be a teaching video or a training video.
[0028] In this step, the video to be processed is divided to obtain at least one divided segment, and a preset number of video frames are respectively extracted from the at least one divided segment to obtain the multiple initial video frames.
[0029] Step 102, performing redundancy removal processing on the plurality of initial video frames to generate key video frames, wherein the key video frames include a variety of teaching demand information.
[0030] The teaching demand information includes at least one of teaching content information, teaching target information and teaching method information. The teaching content information may include teaching knowledge points or skill points, the teaching target information is used to represent the ability level or learning outcomes that students should achieve after completing their studies, and the teaching method information may include the teaching strategies or means used by teachers.
[0031] In this step, when it is determined that there are two adjacent initial video frames that match, the matching initial video frames are determined to be redundant video frames; when it is determined that there is one redundant video frame, the redundant video frame is used as a key video frame; when it is determined that there are multiple redundant video frames, the redundant video frame with the highest definition among the multiple redundant video frames is used as a key video frame.
[0032] It should be noted that if the current teaching scenario is how to tighten screws correctly, the teaching content information may include teaching knowledge points or skill points, such as the types of screws (such as cross slots, slotted slots, hex sockets, etc.), the correct use of screwdrivers or wrenches, and the skills of tightening screws. The teaching goal information is used to characterize the ability level or learning outcomes that students should achieve after completing their studies, such as students being able to identify different types of screws and corresponding tools, students being able to use screwdrivers or wrenches correctly, and students being able to understand and implement the correct tightening sequence and strength. The teaching method information may include the teaching strategies or means used by teachers, such as theoretical lectures and explanations of different types of screws and their applications, practical demonstrations of how to use screwdrivers or wrenches correctly, and allowing students to try to tighten screws under guidance.
[0033] Step 103, obtaining the user's target teaching demand information, and identifying the target position corresponding to the target teaching demand information in the key video frame.
[0034] The target teaching requirement information may include at least one of teaching content information, teaching target information and teaching method information.
[0035] In this step, if the target teaching requirement information is to learn how to use a torque wrench to correctly tighten a cross-slot screw, the features of the cross-slot screw and the torque wrench in the key video frame are extracted, and the positions of the cross-slot screw and the torque wrench are detected, and the detected positions of the cross-slot screw and the torque wrench are used as the target positions.
[0036] Step 104, adding target annotation data of the target teaching requirement information to the target position in the key video frame to generate a target video.
[0037] The target annotation data may be text annotation, graphic annotation, numerical annotation or animation annotation.
[0038] For example, text annotation can be a brief text description to explain the target or ongoing action in the picture; text annotation can also be a short label to identify objects or concepts, such as "torque wrench", "screw", etc. Graphic annotation can be to frame the target object or area with a rectangular box. Graphic annotation can also be to point to a specific part with an arrow to emphasize or guide the line of sight. Graphic annotation can also highlight important areas by changing color or brightness. Numerical annotation can be to show the angle or direction of a specific action through a numerical value. Animated annotation can be to show the movement path or direction of an object through dynamic arrows, and animated annotation can also be to change the highlighted area over time to show the change of action.
[0039] In this step, the key video frame is subjected to image enhancement processing to generate an enhanced video frame; the target annotation data of the target teaching demand information is added to the target position corresponding to the target teaching demand information in the enhanced video frame to generate a target video.
[0040] The above technical scheme, by performing redundancy removal processing on the initial video frame to obtain the key video frame, can remove redundant information in the video to be processed, thereby effectively highlighting the key information of the teaching needs and improving the visual effect of the teaching video; combining the user's target teaching demand information, adding the target annotation data of the target teaching demand information at the target position of the key video frame to generate the target video, can add different target annotation data according to the user's different target teaching demand information, flexibly support diversified needs, so as to better meet the annotation needs in different scenarios, and then can effectively improve the practicality of the teaching video and the user experience.
[0041] Optionally, in Figure 1 The step 101 of obtaining a plurality of initial video frames in the video to be processed may include: The video to be processed is divided to obtain at least one divided segment, and a preset number of video frames are respectively extracted from the at least one divided segment to obtain the multiple initial video frames. In this step, the entire video to be processed can be divided into several smaller segments, each of which is called a segment. The length of the segment can be determined according to actual needs, for example, the length of each segment can be a few seconds or a few minutes. A certain number of video frames are extracted from each segment, and these frames are called initial video frames. The number of extracted initial video frames can be preset as needed, for example, 5 frames or 10 frames can be extracted from each segment.
[0042] The above technical solution obtains the initial video frame by extracting the video to be processed. The initial video frame contains all the visual information in the video to be processed, which helps to quickly determine the content of the video to be processed.
[0043] Figure 2 is based on Figure 1 The embodiment shows a flow chart of a video processing method, such as Figure 2 As shown, in Figure 1 The step 102 of performing redundancy removal processing on the plurality of initial video frames to generate key video frames may include: Step 1021: When it is determined that there are two adjacent initial video frames that match, determine that the matching initial video frames are redundant video frames.
[0044] The redundant video frames may be those video frames that appear in the plurality of initial video frames and have little or no change compared with the preceding and following adjacent frames.
[0045] In this step, the pixel value of each pixel in a plurality of the initial video frames is obtained; when it is determined that the pixel difference between the pixel values corresponding to each pixel in two consecutive initial video frames is less than or equal to a preset pixel threshold, it is determined that the two consecutive initial video frames match.
[0046] Step 1022: When it is determined that there is one redundant video frame, use the redundant video frame as a key video frame.
[0047] Step 1023: When it is determined that there are multiple redundant video frames, use the redundant video frame with the highest definition among the multiple redundant video frames as a key video frame.
[0048] In this step, when it is determined that there are multiple redundant video frames, clarity evaluation indicators of the multiple redundant video frames are determined, and the clarity evaluation indicators include image sharpness, contrast, noise level and image quality index, which are used to evaluate the clarity of the redundant video frames. The redundant video frame with the highest clarity evaluation index among the multiple redundant video frames is used as the key video frame.
[0049] The above technical solution determines the key video frame by comparing the clarity of the redundant video frame, which can ensure that the key video frame is both representative and can provide high-quality visual information, thereby improving the overall efficiency and accuracy of video processing.
[0050] Figure 3 is based on Figure 2 The embodiment shown is a flowchart of a video processing method, such as Figure 3 As shown, in Figure 2 The determination in step 1021 that there are two adjacent initial video frames that match each other may include: S11, obtaining a pixel value of each pixel in a plurality of the initial video frames.
[0051] In this step, the pixel value of each pixel in each of the initial video frames is obtained. For a color image, the pixel value of each pixel can be represented by a (R, G, B) triplet, where the value range of each color channel is usually between 0 and 255; for a grayscale image, the pixel value of each pixel can be represented by a grayscale value, which represents an intensity level from black (0) to white (255).
[0052] S12: When it is determined that the pixel difference between the pixel values corresponding to each pixel in two consecutive initial video frames is less than or equal to a preset pixel threshold, determine that the two consecutive initial video frames match.
[0053] In this step, any two consecutive initial video frames among the multiple initial video frames are taken as the first initial video frame and the second initial video frame, and the pixel value corresponding to each pixel in the first initial video frame and the second initial video frame is obtained by traversing each pixel in the first initial video frame and the second initial video frame, and the pixel difference between the pixel values corresponding to each pixel in the first initial video frame and the second initial video frame is determined. When it is determined that the pixel difference between the pixel values corresponding to each pixel in the first initial video frame and the second initial video frame is less than or equal to a preset pixel threshold, it is determined that the two consecutive initial video frames match.
[0054] The above technical solution can accurately detect subtle changes between two initial video frames through pixel-by-pixel comparison, retain more detail information, and judge whether the two initial video frames match by comparing the pixel difference between the pixel values of each pixel in two consecutive initial video frames and determining whether these differences are less than or equal to a preset pixel threshold. The pixel threshold can be adjusted according to the actual application scenario to adapt to different types of video content, thereby effectively improving the flexibility of video processing.
[0055] Figure 4 is based on Figure 1 The embodiment shown is a flowchart of another video processing method, such as Figure 4 As shown, in Figure 1 Adding the target annotation of the target teaching requirement information to the target position in the key video frame in step 104 to generate the target video may include: Step 1041, performing image enhancement processing on the key video frame to generate an enhanced video frame.
[0056] In this step, the key video frame can be enhanced by increasing or decreasing the overall brightness of the key video frame and adjusting the contrast; the key video frame can also be enhanced by adjusting the saturation and brightness of the key video frame; the key video frame can also be enhanced by a variety of methods such as image sharpening, histogram equalization, noise reduction, edge detection and image fusion to obtain the enhanced video.
[0057] Step 1042, adding target annotation data of the target teaching requirement information to the target position corresponding to the target teaching requirement information in the enhanced video frame to generate a target video.
[0058] The target annotation data may be text annotation, graphic annotation, numerical annotation or animation annotation.
[0059] In this step, the target teaching demand information may include at least one of teaching content information, teaching target information and teaching method information. If the target teaching demand information is to learn how to use a torque wrench to correctly tighten a cross-slot screw. Extract the features of the cross-slot screw and the torque wrench in the key video frame, detect the positions of the cross-slot screw and the torque wrench, and use the detected positions of the cross-slot screw and the torque wrench as the target position corresponding to the target teaching demand information. Add the target annotation of the target teaching demand information at the target position to generate a target video.
[0060] It should be noted that the target annotation data can be text annotation, graphic annotation, numerical annotation or animation annotation. Text annotation can be a brief text description used to explain the target or ongoing action in the picture; text annotation can also be a short label to identify objects or concepts, such as "torque wrench", "screw", etc. Graphic annotation can be to frame the target object or area with a rectangular box. Graphic annotation can also be to point to a specific part with an arrow to emphasize or guide the line of sight. Graphic annotation can also highlight important areas by changing color or brightness. Numerical annotation can be to display the angle or direction of a specific action through a numerical value. Animation annotation can be to display the moving path or direction of an object through dynamic arrows, and animation annotation can also be to change the highlighted area over time to show the change of action.
[0061] The above technical scheme generates a target video by adding target annotation data of the target teaching demand information at the target position of the enhanced video frame in combination with the user's target teaching demand information. It can add different target annotation data according to the user's different target teaching demand information, flexibly support diversified needs, so as to better meet the annotation needs in different scenarios, and thus effectively improve the practicality of the teaching video and the user experience.
[0062] Figure 5 is based on Figure 1 The embodiment shown is a flowchart of another video processing method, as shown in FIG. Figure 5 As shown, the method also includes: Step 105: Obtain the video format requirement of the user.
[0063] The video format requirements may include MP4 (MPEG-4 Part 14), AVI (Audio Video Interleave), MOV (QuickTime File Format), WMV (Windows Media Video), and other video formats.
[0064] Step 106: Using the video format requirement as the output video format of the target video.
[0065] In this step, after the video processing is completed, a video encoder and packaging format corresponding to the video format requirement are used to generate a final video file, and the video encoding method and packaging format of the video file meet the user's video format requirement.
[0066] The above technical solution determines the output video format of the target video according to the video format requirements of the user, which can ensure that the output video format of the target video is compatible with the media playback software currently being used by the user, avoiding the problem of format unsupport, and can also select a format that provides higher image quality or smaller file size according to the different video format requirements of the user, thereby effectively improving the user experience and satisfaction.
[0067] Figure 6 is a block diagram of a video processing device according to an exemplary embodiment. Figure 6 , the video processing device 600 may include: A first acquisition module 601 is used to acquire a plurality of initial video frames in a video to be processed; A first generating module 602 is used to perform redundancy removal processing on the plurality of the initial video frames to generate key video frames, wherein the key video frames include a plurality of teaching demand information, and the teaching demand information includes at least one of teaching content information, teaching target information and teaching method information; The second acquisition module 603 is used to acquire the user's target teaching demand information and identify the target position corresponding to the target teaching demand information in the key video frame; The second generating module 604 is used to add the target annotation data of the target teaching requirement information at the target position in the key video frame to generate the target video.
[0068] The video to be processed may be a teaching video or a training video. The teaching demand information includes at least one of teaching content information, teaching target information and teaching method information. The teaching content information may include teaching knowledge points or skill points, the teaching target information is used to characterize the ability level or learning outcomes that students should achieve after completing their studies, and the teaching method information may include the teaching strategies or means used by teachers. The target annotation data may be text annotation, graphic annotation, numerical annotation or animation annotation.
[0069] It should be noted that if the current teaching scenario is how to tighten screws correctly, the teaching content information may include teaching knowledge points or skill points, such as the types of screws (such as cross slot, slot, hexagon socket, etc.), the correct use of screwdrivers or wrenches and the skills of tightening screws. The teaching goal information is used to characterize the ability level or learning outcomes that students should achieve after completing their studies, such as students being able to identify different types of screws and corresponding tools, students being able to use screwdrivers or wrenches correctly, and students being able to understand and implement the correct tightening sequence and strength. The teaching method information may include the teaching strategies or means used by teachers, such as theoretical teaching and explanation of different types of screws and their application scenarios, practical demonstration of how to use screwdrivers or wrenches correctly, and allowing students to try to tighten screws under guidance. If the target teaching demand information is to learn how to use a torque wrench to correctly tighten cross-slot screws. Extract the features of the cross-slot screws and torque wrenches in the key video frames, detect the positions of the cross-slot screws and torque wrenches, and use the detected positions of the cross-slot screws and torque wrenches as the target positions.
[0070] The above technical scheme obtains multiple initial video frames in the video to be processed; performs redundancy removal processing on the multiple initial video frames to generate key video frames, wherein the key video frames include multiple teaching demand information, wherein the teaching demand information includes at least one of teaching content information, teaching target information and teaching method information; obtains the user's target teaching demand information, and identifies the target position corresponding to the target teaching demand information in the key video frame; adds the target annotation data of the target teaching demand information to the target position in the key video frame to generate the target video. In this way, by performing redundancy removal processing on the initial video frames to obtain key video frames, redundant information in the video to be processed can be removed, thereby effectively highlighting the key information of the teaching demand and improving the visual effect of the teaching video; combining the user's target teaching demand information, adding the target annotation data of the target teaching demand information to the target position of the key video frame to generate the target video, different target annotation data can be added for different target teaching demand information of the user, thereby meeting the annotation requirements in different teaching scenarios, and thus effectively improving the practicality of the teaching video and the user's experience.
[0071] Optionally, the first acquisition module 601 is further used to divide the video to be processed to obtain at least one divided segment, and extract a preset number of video frames from the at least one divided segment to obtain the multiple initial video frames.
[0072] The entire video to be processed can be divided into several smaller segments, each of which is called a segment. The length of the segment can be determined according to actual needs. For example, the length of each segment can be a few seconds or a few minutes. A certain number of video frames are extracted from each segment, and these frames are called initial video frames. The number of extracted initial video frames can be preset as needed, for example, 5 frames or 10 frames can be extracted from each segment.
[0073] The above technical solution obtains the initial video frame by extracting the video to be processed. The initial video frame contains all the visual information in the video to be processed, which helps to quickly determine the content of the video to be processed.
[0074] Optionally, the first generating module 602 is further used to, when it is determined that there are two adjacent initial video frames that match, determine that the matching initial video frame is a redundant video frame; when it is determined that there is one redundant video frame, use the redundant video frame as a key video frame; when it is determined that there are multiple redundant video frames, use the redundant video frame with the highest clarity among the multiple redundant video frames as the key video frame.
[0075] The redundant video frames may be those video frames that appear in the multiple initial video frames and have little or no change compared with the adjacent frames before and after. By obtaining the pixel value of each pixel in the multiple initial video frames; when it is determined that the pixel difference between the pixel values corresponding to each pixel in two consecutive initial video frames is less than or equal to the preset pixel threshold, it is determined that the two consecutive initial video frames match. When it is determined that there are multiple redundant video frames, the clarity evaluation index of the multiple redundant video frames is determined, and the clarity evaluation index includes image sharpness, contrast, noise level and image quality index, which are used to evaluate the clarity of the redundant video frames. The redundant video frame with the highest clarity evaluation index among the multiple redundant video frames is used as the key video frame.
[0076] The above technical solution determines the key video frame by comparing the clarity of the redundant video frame, which can ensure that the key video frame is both representative and can provide high-quality visual information, thereby improving the overall efficiency and accuracy of video processing.
[0077] Optionally, the first generating module 602 is further used to obtain the pixel value of each pixel in the multiple initial video frames; when it is determined that the pixel difference between the pixel values corresponding to each pixel in two consecutive initial video frames is less than or equal to a preset pixel threshold, it is determined that the two consecutive initial video frames match.
[0078] The above technical solution can accurately detect subtle changes between two initial video frames through pixel-by-pixel comparison, retain more detail information, and judge whether the two initial video frames match by comparing the pixel difference between the pixel values of each pixel in two consecutive initial video frames and determining whether these differences are less than or equal to a preset pixel threshold. The pixel threshold can be adjusted according to the actual application scenario to adapt to different types of video content, thereby effectively improving the flexibility of video processing.
[0079] Optionally, the second generating module 604 is further used to perform image enhancement processing on the key video frame to generate an enhanced video frame; add target annotation data of the target teaching requirement information at the target position corresponding to the target teaching requirement information in the enhanced video frame to generate a target video.
[0080] Among them, the key video frame can be subjected to image enhancement processing by increasing or decreasing the overall brightness of the key video frame and adjusting the contrast; the key video frame can also be subjected to image enhancement processing by adjusting the saturation and brightness of the key video frame; the key video frame can also be subjected to image enhancement processing by a variety of methods such as image sharpening, histogram equalization, noise reduction, edge detection and image fusion to obtain the enhanced video.
[0081] It should be noted that the target annotation data can be text annotation, graphic annotation, numerical annotation or animation annotation. Text annotation can be a brief text description used to explain the target or ongoing action in the picture; text annotation can also be a short label to identify objects or concepts, such as "torque wrench", "screw", etc. Graphic annotation can be to frame the target object or area with a rectangular box. Graphic annotation can also be to point to a specific part with an arrow to emphasize or guide the line of sight. Graphic annotation can also highlight important areas by changing color or brightness. Numerical annotation can be to display the angle or direction of a specific action through a numerical value. Animation annotation can be to display the moving path or direction of an object through dynamic arrows, and animation annotation can also be to change the highlighted area over time to show the change of action.
[0082] The above technical scheme generates a target video by adding target annotation data of the target teaching demand information at the target position of the enhanced video frame in combination with the user's target teaching demand information. It can add different target annotation data according to the user's different target teaching demand information, flexibly support diversified needs, so as to better meet the annotation needs in different scenarios, and thus effectively improve the practicality of the teaching video and the user experience.
[0083] Optionally, the device 600 further includes: The output module 605 is used to obtain the video format requirement of the user; and use the video format requirement as the output video format of the target video.
[0084] The video format requirement may include MP4 (MPEG-4 Part 14), AVI (Audio Video Interleave), MOV (QuickTime File Format), WMV (Windows Media Video), etc. After the video processing is completed, a video encoder and packaging format corresponding to the video format requirement are used to generate a final video file, and the video encoding method and packaging format of the video file meet the user's video format requirement.
[0085] The above technical solution determines the output video format of the target video according to the video format requirements of the user, which can ensure that the output video format of the target video is compatible with the media playback software currently being used by the user, avoiding the problem of format unsupport, and can also select a format that provides higher image quality or smaller file size according to the different video format requirements of the user, thereby effectively improving the user experience and satisfaction.
[0086] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.
[0087] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, and when the program instructions are executed by a processor, the steps of the method for determining a job plan provided by the present disclosure are implemented.
[0088] Figure 7 FIG. 7 is a block diagram of an electronic device 700 according to an exemplary embodiment. Figure 7 As shown, the electronic device 700 may include: a processor 701 , a memory 702 . The electronic device 700 may also include one or more of a multimedia component 703 , an input / output (I / O) interface 704 , and a communication component 705 .
[0089] Among them, the processor 701 is used to control the overall operation of the electronic device 700 to complete all or part of the steps in the above video processing method. The memory 702 is used to store various types of data to support the operation of the electronic device 700. Such data may include, for example, instructions for any application or method operating on the electronic device 700, as well as application-related data, such as contact data, messages sent and received, pictures, audio, video, and so on. The memory 702 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disc. The multimedia component 703 may include a screen and an audio component. Among them, the screen can be, for example, a touch screen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone, and the microphone is used to receive external audio signals. The received audio signal can be further stored in the memory 702 or sent through the communication component 705. The audio component also includes at least one speaker for outputting audio signals. The I / O interface 704 provides an interface between the processor 701 and other interface modules, and the above other interface modules can be a keyboard, a mouse, buttons, etc. These buttons can be virtual buttons or physical buttons. The communication component 705 is used for wired or wireless communication between the electronic device 700 and other devices. Wireless communication, such as Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, 4G, NB-IOT, eMTC, or other 5G, etc., or a combination of one or more of them, is not limited herein. Therefore, correspondingly, the communication component 705 may include: a Wi-Fi module, a Bluetooth module, an NFC module, and so on.
[0090] In an exemplary embodiment, the electronic device 700 can be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to execute the above-mentioned video processing method.
[0091] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, and when the program instructions are executed by the processor, the steps of the above-mentioned video processing method are implemented. For example, the computer-readable storage medium can be the above-mentioned memory 702 including program instructions, and the above-mentioned program instructions can be executed by the processor 701 of the electronic device 700 to complete the above-mentioned video processing method.
[0092] In another exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program executable by a programmable device. The computer program has a code portion for executing the above-mentioned video processing method when executed by the programmable device.
[0093] The preferred embodiments of the present disclosure are described in detail above in conjunction with the accompanying drawings; however, the present disclosure is not limited to the specific details in the above embodiments. Within the technical concept of the present disclosure, a variety of simple modifications can be made to the technical solution of the present disclosure, and these simple modifications all fall within the protection scope of the present disclosure.
[0094] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any modification, use or adaptation of the present disclosure, which follows the general principles of the present disclosure and includes common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.
[0095] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present disclosure will not further describe various possible combinations.
[0096] In addition, various embodiments of the present disclosure may be arbitrarily combined, and as long as they do not violate the concept of the present disclosure, they should also be regarded as the contents disclosed by the present disclosure.
Claims
1. A video processing method, characterized in that: The method comprises: Obtain multiple initial video frames in the video to be processed; Performing redundancy removal processing on the plurality of the initial video frames to generate key video frames, wherein the key video frames include a plurality of teaching demand information, wherein the teaching demand information includes at least one of teaching content information, teaching target information and teaching method information; Acquire target teaching demand information of the user, and identify the target position corresponding to the target teaching demand information in the key video frame; The target annotation data of the target teaching requirement information is added to the target position in the key video frame to generate a target video.
2. The video processing method according to claim 1, characterized in that: The step of obtaining a plurality of initial video frames in the video to be processed includes: The video to be processed is divided to obtain at least one divided segment, and a preset number of video frames are respectively extracted from the at least one divided segment to obtain the multiple initial video frames.
3. The video processing method according to claim 1, characterized in that: The performing redundancy removal processing on the plurality of initial video frames to generate key video frames comprises: In the case where it is determined that there are two adjacent initial video frames that match, determining the matching initial video frames as redundant video frames; In the case of determining that there is one redundant video frame, taking the redundant video frame as a key video frame; When it is determined that there are a plurality of redundant video frames, the redundant video frame with the highest definition among the plurality of redundant video frames is used as a key video frame.
4. The video processing method according to claim 3, characterized in that: Determining whether two adjacent initial video frames match each other includes: Obtaining a pixel value of each pixel in a plurality of the initial video frames; In the case where it is determined that the pixel difference between the pixel values corresponding to each pixel in two consecutive initial video frames is less than or equal to a preset pixel threshold, it is determined that the two consecutive initial video frames match.
5. The video processing method according to claim 1, characterized in that: The step of adding target annotation data of the target teaching requirement information to the target position in the key video frame to generate the target video includes: Performing image enhancement processing on the key video frame to generate an enhanced video frame; The target annotation data of the target teaching requirement information is added to the target position corresponding to the target teaching requirement information in the enhanced video frame to generate a target video.
6. The video processing method according to claim 1, characterized in that: The method further comprises: Obtaining the video format requirement of the user; The video format requirement is used as the output video format of the target video.
7. A video processing device, characterized in that: The device comprises: A first acquisition module is used to acquire a plurality of initial video frames in the video to be processed; A first generating module is used to perform redundancy removal processing on the plurality of the initial video frames to generate key video frames, wherein the key video frames include a plurality of teaching demand information, and the teaching demand information includes at least one of teaching content information, teaching target information and teaching method information; A second acquisition module is used to acquire the user's target teaching demand information and identify the target position corresponding to the target teaching demand information in the key video frame; The second generating module is used to add the target annotation data of the target teaching requirement information at the target position in the key video frame to generate the target video.
8. The video processing device according to claim 7, characterized in that: The first acquisition module is further used to divide the video to be processed to obtain at least one divided segment, and extract a preset number of video frames from the at least one divided segment to obtain the multiple initial video frames.
9. The video processing device according to claim 7, characterized in that: The first generating module is further used to, when it is determined that there are two adjacent initial video frames that match, determine that the matching initial video frame is a redundant video frame; when it is determined that there is one redundant video frame, use the redundant video frame as a key video frame; when it is determined that there are multiple redundant video frames, use the redundant video frame with the highest clarity among the multiple redundant video frames as the key video frame.
10. The video processing device according to claim 9, characterized in that: The first generating module is also used to obtain the pixel value of each pixel in the multiple initial video frames; when it is determined that the pixel difference between the pixel values corresponding to each pixel in two consecutive initial video frames is less than or equal to a preset pixel threshold, it is determined that the two consecutive initial video frames match.
11. The video processing device according to claim 7, characterized in that: The second generating module is also used to perform image enhancement processing on the key video frame to generate an enhanced video frame; and add target annotation data of the target teaching requirement information at the target position corresponding to the target teaching requirement information in the enhanced video frame to generate a target video.
12. The video processing device according to claim 7, characterized in that: The device also includes: The output module is used to obtain the video format requirement of the user; and use the video format requirement as the output video format of the target video.
13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
14. An electronic device, characterized in that: include: a memory having a computer program stored thereon; A processor, configured to execute the computer program in the memory to implement the steps of the method according to any one of claims 1 to 6.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.
Citation Information
Cited By
Video enhancement method and device
CN120431001A