Video processing method and apparatus, device, and medium

By extracting frames and features from the original video, multiple sets of cut frames are determined, solving the problem of inaccurate cut results in existing technologies and achieving more efficient video processing and a better user experience.

WO2025261435A1PCT designated stage Publication Date: 2025-12-26BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/102041
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-20
Filing Date
2025-06-19
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

Existing video editing technologies struggle to provide precise cuts at different granularities, resulting in inefficient video processing and a poor user experience.

Method used

By extracting frames from the original video, feature groups with multiple feature granularities are extracted. Based on these feature groups, multiple cut frames are determined, and cut results are output according to different feature granularities. A network model is used for feature extraction and fusion to improve the accuracy of the cut results.

Benefits of technology

It achieves more accurate cut results at different granularities, improves video processing efficiency, meets more user needs, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025102041_26122025_PF_FP_ABST
    Figure CN2025102041_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides a video processing method and apparatus, a device, and a medium. The video processing method comprises: performing frame extraction on an original video to be processed, to obtain a reference image frame sequence; according to multiple different feature granularities, performing feature extraction on the reference image frame sequence, to obtain multiple feature groups, each feature group comprising multiple frames of feature images; on the basis of the multiple feature groups, determining multiple groups of cut frames; on the basis of the cut frames, outputting a cut result for the original video to a user, according to the different feature granularities. The present video processing method provides a more comprehensive cut result for a user from different granularities, so that the cut result is more accurate under the limitations of different granularities, meeting more needs of the user, improving video processing efficiency, and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Video processing method, apparatus, device, and medium

[0001] Cross Reference to Related Applications

[0002] This application claims priority to Chinese Patent Application No. 202410800080.2, filed on June 20, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD

[0003] The present disclosure relates to a video processing method, apparatus, device, and medium. BACKGROUND

[0004] With the continuous development and improvement of digital media technology, the application of video has become more and more extensive. For example, video is widely used in advertising, film, and art fields. Video not only brings people a lot of convenience, but also adds more interest to people's life. In order to make the video more vivid, people usually edit the existing video through technical means to edit and adjust the video or create a new one. For example, the video can be edited by combining background music, video materials, special effects, subtitles, transitions, etc., so as to obtain a new video. SUMMARY

[0005] The present disclosure provides a video processing method, apparatus, device, and medium.

[0006] According to a first aspect, a video processing method is provided, the method comprising:

[0007] frame extraction is performed on a to-be-processed original video to obtain a reference image frame sequence;

[0008] feature extraction is performed on the reference image frame sequence according to a plurality of different feature granularities to obtain a plurality of feature groups; the feature groups include a plurality of feature image frames;

[0009] a plurality of groups of cut mirror frames are determined based on the plurality of feature groups;

[0010] a cut mirror result for the original video is output to a user according to different feature granularities based on the cut mirror frames.

[0011] According to a second aspect, a video processing apparatus is provided, the apparatus comprising:

[0012] a frame extraction module configured to perform frame extraction on a to-be-processed original video to obtain a reference image frame sequence;

[0013] an extraction module configured to perform feature extraction on the reference image frame sequence according to a plurality of different feature granularities to obtain a plurality of feature groups; the feature groups include a plurality of feature image frames;

[0014] The determining module is used to determine multiple sets of cut-off frames based on the multiple feature groups;

[0015] The output module is used to output the cut-out results for the original video to the user based on the cut-out frames and according to different feature granularities.

[0016] According to a third aspect, a computer-readable storage medium is provided, the storage medium storing a computer program that, when executed by a processor, implements the method described in any one of the first aspects above.

[0017] According to a fourth aspect, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any one of the first aspects. Attached Figure Description

[0018] To more clearly illustrate the technical solutions of the embodiments in this specification, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 is a schematic diagram of a video processing scenario according to an exemplary embodiment of the present disclosure;

[0020] Figure 2 is a flowchart illustrating a video processing method according to an exemplary embodiment of the present disclosure;

[0021] Figure 3 is a flowchart illustrating another video processing method according to an exemplary embodiment of the present disclosure;

[0022] Figure 4 is a schematic diagram of another video processing scenario according to an exemplary embodiment of the present disclosure;

[0023] Figure 5 is a block diagram of a video processing apparatus according to an exemplary embodiment of the present disclosure; and

[0024] Figure 6 is a schematic block diagram of an electronic device provided in some embodiments of this disclosure. Detailed Implementation

[0025] To enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.

[0026] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0027] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure refers to and includes any and all possible combinations of one or more of the associated listed items.

[0028] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0029] With the continuous development and improvement of digital media technology, the application of video has become increasingly widespread. For example, video is widely used in advertising, film, and art. Video not only brings people many conveniences but also adds more fun to their lives. To make videos more vivid, people often use technical means to edit existing videos, adjusting or creating new content. For example, videos can be edited by combining background music, video footage, special effects, subtitles, transitions, and other elements to create new videos.

[0030] During editing, it may be necessary to delete redundant segments, rearrange the order of different segments, or insert additional clips. Therefore, in actual editing, video editing involves cutting the video, which means dividing it into segments by comparing different scenes. These segments may differ in scene or semantic information. For example, a high jump video can be divided into several stages: the approach run, the takeoff, the flight, and the landing. The transition points between these stages can be used as cut points.

[0031] This disclosure provides a video processing method that extracts frames from the original video to obtain a reference image frame sequence. Features are then extracted from the reference image frame sequence according to different feature granularities to obtain multiple feature groups. Based on these feature groups, multiple sets of cut-off frames are determined. Finally, based on these multiple sets of cut-off frames, cut-off results for the original video are output to the user according to different feature granularities. This provides users with more comprehensive cut-off results at different granularities, making the cut-off results more accurate under the constraints of different granularities, meeting more user needs, improving video processing efficiency, and enhancing the user experience.

[0032] Referring to Figure 1, it is a schematic diagram of a video processing scenario according to an exemplary embodiment.

[0033] As shown in Figure 1, firstly, the original video A can be acquired, and frame extraction is performed on the original video A to obtain n reference images. Each reference image corresponds to a playback moment in the original video A. The n reference images can be arranged according to the playback moment corresponding to the reference image, forming a reference image frame sequence LA. Then, the reference image frame sequence LA can be input into network model 1. Network model 1 performs dimensionality reduction on the reference images in the reference image frame sequence LA and extracts the image features of each reference image to obtain a set of initial image features. Here, network model 1 can be, for example, a two-dimensional convolutional neural network. It is understood that this embodiment does not limit the specific type of network model 1.

[0034] Next, a set of initial image features extracted by network model 1 can be input into network model 2. Network model 2 processes this set of initial image features according to various different feature granularities, thereby obtaining feature group C1, feature group C2, and feature group C3. Different feature groups correspond to different feature granularities, and each feature group includes multiple frames of feature images. For example, feature group C1 corresponds to coarse feature granularity, feature group C2 corresponds to medium feature granularity, and feature group C3 corresponds to fine feature granularity. Network model 2 can be, for example, an FPN (Feature Pyramid Network). It is understood that this embodiment does not limit the specific type of network model 1.

[0035] Then, feature groups C1, C2, and C3 can be input into network model 3 for context fusion processing to obtain fused feature group R1 corresponding to feature group C1, fused feature group R2 corresponding to feature group C2, and fused feature group R3 corresponding to feature group C3. Network model 3 can be, for example, a Transformer Encoder model; however, this embodiment does not limit the specific type of network model 3. A preset algorithm is then used to calculate the temporal similarity of fused feature groups R1, R2, and R3 respectively, resulting in temporal similarity maps M1, M2, and M3 corresponding to fused feature groups R1, R2, and R3.

[0036] Finally, the temporal similarity maps M1, M2, and M3 can be input into network model 4. Network model 4, based on these maps, determines the confidence level of each feature image in feature groups C1, C2, and C3. Network model 4 can be, for example, a CLS (Constrained Linear System) model; however, this embodiment does not limit the specific type of network model 4. The cut-off frames for feature groups C1, C2, and C3 can be determined based on the confidence level of each feature image. For any feature group, the reference image corresponding to the feature image with a confidence level greater than a preset threshold can be used as the cut-off frame. The playback time corresponding to each cut-off frame can be used as the cut-off result and output to the user. For example, if the confidence scores of feature images X11, X15, and X19 in feature group C1 are all greater than a preset threshold, then reference images K11 (corresponding to feature image X11), K15 (corresponding to feature image X15), and K19 (corresponding to feature image X19) can be used as cut frames. Playback time t1 (corresponding to reference image K11), playback time t5 (corresponding to reference image K15), and playback time t9 (corresponding to reference image K19) can be used as cut result 1 and output to the user. Similarly, cut result 2 (determined based on feature group C2) and cut result 3 (determined based on feature group C3) can be output to the user.

[0037] The present disclosure will now be described in detail with reference to specific embodiments.

[0038] Figure 2 is a flowchart illustrating a video processing method according to an exemplary embodiment. This method can be applied to a terminal device. In this embodiment, for ease of understanding, an example is given using a terminal device capable of video processing applications. Those skilled in the art will understand that the terminal device may include, but is not limited to, mobile terminal devices such as smartphones, smart wearable devices, tablets, laptops, and desktop computers. The method may include the following steps:

[0039] As shown in Figure 2, in step 201, the original video to be processed is frame-by-frame extracted to obtain multiple reference image frames, and in step 202, features are extracted from the reference image frames according to multiple different feature granularities to obtain multiple feature groups.

[0040] In this embodiment, the original video to be processed can be a video that requires camera cuts. The original video can be processed by frame extraction to obtain multiple reference image frames. Each reference image frame corresponds to a playback moment, and these multiple reference image frames can form a set of reference image frame sequences according to the order of the playback moments.

[0041] After obtaining the reference image frame sequence, features can be extracted from the sequence at various feature granularities to obtain multiple feature groups. Each feature group includes multiple feature images, and different feature groups correspond to different feature granularities. For example, the various feature granularities include feature granularity a, feature granularity b, and feature granularity c. Features can be extracted from each reference image frame in the sequence according to feature granularity a, resulting in a feature group Ta corresponding to feature granularity a. Feature group Ta includes the feature image corresponding to each reference image frame after feature extraction. Similarly, feature extraction from each reference image frame in the sequence according to feature granularity b yields a feature group Tb corresponding to feature granularity b, and feature extraction from each reference image frame in the sequence according to feature granularity c yields a feature group Tc corresponding to feature granularity c.

[0042] In this embodiment, feature granularity represents the degree of feature refinement. A higher degree of feature granularity means the features at that granularity are closer to the texture details in the image. Conversely, a lower degree of feature granularity means the features at that granularity are closer to the semantic contours in the image. Different downsampling ratios can be used to downsample the reference image frame sequence, resulting in multiple feature groups at various feature granularities. Different downsampling ratios correspond to different feature granularities. For example, a larger downsampling ratio indicates more prominent global semantic features extracted from the image; therefore, a larger downsampling ratio corresponds to a coarser feature granularity (i.e., a lower degree of feature granularity refinement). A smaller downsampling ratio indicates more prominent local texture features extracted from the image; therefore, a smaller downsampling ratio corresponds to a finer feature granularity (i.e., a higher degree of feature granularity refinement).

[0043] It is understood that in this embodiment, the feature granularity can be divided into at least two types: coarse feature granularity and fine feature granularity, or it can be divided into more than two types of feature granularity. This embodiment does not limit the specific division of feature granularity.

[0044] In step 203, multiple sets of cut frames are determined based on multiple feature groups.

[0045] In this embodiment, multiple cut frames can be determined based on multiple feature groups, wherein one feature group corresponds to a set of cut frames, that is, a set of cut frames can be determined at each feature granularity. Specifically, for any feature group, firstly, the temporal similarity information corresponding to the feature group can be obtained; then, based on the temporal similarity information, the confidence level corresponding to the feature image in the feature group can be determined; finally, based on the confidence level, a set of cut frames corresponding to the feature group can be determined.

[0046] Each set of cut frames can include multiple cut images, each representing a moment in the original video where a scene transition occurs. For example, in a sports video, a cut image might be a frame where the camera switches from the playing field to the audience. The multiple cut images in each set can divide the original video into several different stages at their respective feature granularities. Different feature granularities can determine different cut images, and the number of cut images included in different sets of cut frames can be the same or different at different feature granularities. Generally, a set of cut frames at a coarser granularity usually includes fewer cut images than a set of cut frames at a finer granularity.

[0047] For example, based on the feature set T1 at the coarse-grained level, a set of cut frames Q1 can be determined. The n cut frames included in this set of cut frames Q1 can divide the original video into n+1 stages according to the coarse-grained level. Based on the feature set T2 at the fine-grained level, a set of cut frames Q2 can be determined. The m cut frames included in this set of cut frames Q2 can divide the original video into m+1 stages according to the fine-grained level, and generally n is less than m.

[0048] In step 204, based on the multiple sets of cut frames, the cut results for the original video are output to the user.

[0049] In this embodiment, multiple cut times corresponding to the original video timeline can be determined based on these multiple sets of cut frames, serving as the cut results for the original video. Each set of cut times corresponds to a feature granularity. Specifically, each frame in the original video corresponds to a playback time on the original video timeline. The time corresponding to each cut frame in each set of cut frames on that timeline can be determined based on the original video timeline, serving as the cut time. Since each set of cut frames corresponds to a feature granularity, the cut times corresponding to a set of cut frames also correspond to a feature granularity. Finally, multiple sets of cut times can be output to the user according to different feature granularities.

[0050] For example, at the coarse-grained level, a set of cut frames Q1 is determined, which includes cut images S11, S12, and S13. Cut image S11 corresponds to time t11 on the original image time axis, cut image S12 corresponds to time t12 on the original image time axis, and cut image S13 corresponds to time t13 on the original image time axis. At the fine-grained level, a set of cut frames Q2 is defined, including cut images S21, S22, S23, S24, and S25. Cut image S21 corresponds to time t21 on the original image timeline, cut image S22 corresponds to time t22, cut image S23 corresponds to time t23, cut image S24 corresponds to time t24, and cut image S25 corresponds to time t25. Times t11, t12, and t13 can be used as a set of cut times at the coarse-grained level, and times t21, t22, t23, t24, and t25 can be used as a set of cut times at the fine-grained level. These two sets of cut times are then output to the user according to the coarse-grained and fine-grained levels, respectively. For example, markers can be added to cut moments at different granularities on different timelines to show users the cut moments. Alternatively, markers can be added to cut moments at different granularities on the same timeline using different marker formats (e.g., different colors or different shapes) to show users the cut moments.

[0051] In this embodiment, a specific application scenario is as follows: for a video themed around cooking, the processing method of this embodiment can produce a set of cut frames at a coarse-grained level and a set of cut frames at a fine-grained level. The set of cut frames at the coarse-grained level can divide the video into stages such as "washing vegetables," "chopping vegetables," "stir-frying vegetables," and "dish display." However, for example, in the "chopping vegetables" stage, it cannot distinguish what type of vegetables are being chopped.

[0052] Furthermore, a set of cut frames at a finer granular level can further divide the "washing vegetables" stage in the video into the "washing cucumbers" stage and the "washing tomatoes" stage; the "cutting vegetables" stage in the video into the "cutting cucumbers" stage and the "cutting tomatoes" stage; and the "stir-frying" stage in the video into the "adding vegetables" stage, the "adding seasonings" stage, and the "stir-frying" stage, etc.

[0053] After obtaining the cut-off results, users can edit or trim the video accordingly, such as adding background music or subtitles for a specific segment.

[0054] This embodiment is not limited to the application scenarios described above and can also be applied to other scenarios. The video processing method provided in this disclosure involves extracting frames from the original video to obtain a reference image frame sequence. Multiple feature groups are then extracted from the reference image frame sequence according to various feature granularities. Based on these feature groups, multiple sets of cut frames are determined. Based on these multiple sets of cut frames, cut results for the original video are output to the user according to different feature granularities. This provides users with more comprehensive cut results from different granularities, making the cut results more accurate under the constraints of different granularities, meeting more user needs, improving video processing efficiency, and enhancing the user experience.

[0055] Figure 3 is a flowchart illustrating another video processing method according to an exemplary embodiment, which describes the process of determining a set of cut frames, including the following steps:

[0056] As shown in Figure 3, in step 301, the temporal similarity information corresponding to the feature group is determined.

[0057] In this embodiment, firstly, a set of feature groups can be obtained. This set of feature groups may include multiple frames of feature images obtained according to a feature granularity. The temporal similarity information corresponding to the multiple frames of feature images included in the feature group can be determined. The temporal similarity information corresponding to the feature group can be information that represents the similarity between feature images corresponding to different times within the feature group.

[0058] In one implementation, the similarity between each pair of feature images in the feature group can be directly calculated, and then the temporal similarity information corresponding to the feature group can be determined based on the pairwise similarity between the feature images in the feature group.

[0059] In another implementation, context fusion processing can be performed on multiple frames of feature images within the feature group to obtain a multi-frame fused feature image. Specifically, the context feature images before and after the current frame feature image can be obtained. For example, the n frames before and the n frames after the current frame feature image can be used as the context feature images of the current frame feature image. Then, the current frame feature image and its context feature images can be fused temporally to obtain a single frame fused feature image. This temporal feature fusion can be performed using a Transformer Encoder model. It is understood that temporal feature fusion can also be performed in other reasonable ways, and this embodiment is not limited in this regard.

[0060] Finally, the temporal similarity information corresponding to the feature group can be determined based on the multi-frame fused feature images. Specifically, the similarity between each pair of multi-frame fused feature images can be calculated. Then, based on the pairwise similarity between the multi-frame fused feature images, a temporal similarity map can be generated as the temporal similarity information corresponding to the feature group. This temporal similarity map includes multiple regions, each region corresponding to two fused feature images, and the pixel value of the pixels in this region represents the similarity between the two fused feature images.

[0061] For example, the number of frames in the fused feature image can be 10, which are W0 to W9. The similarity between each pair of W0 to W9 can be calculated, resulting in 100 similarity values: D00, D01, ..., D99. Here, D00 represents the similarity between W0 and W0, D01 represents the similarity between W0 and W1, ..., D99 represents the similarity between W9 and W9, and so on. Then, these 100 similarity values ​​are converted into 100 color values ​​and displayed in 100 regions of the image, thus obtaining a temporal similarity map. Figure 4 shows the temporal similarity map obtained based on these 100 similarities. Region 401 displays the color value converted from D00, region 402 displays the color value converted from D40, region 403 displays the color value converted from D90, and region 404 displays the color value converted from D99.

[0062] In step 302, based on the aforementioned temporal similarity information, the confidence level corresponding to each feature image in the feature group is determined, and in step 303, based on the confidence level, a set of cut frames corresponding to the feature group is determined.

[0063] In this embodiment, based on the aforementioned temporal similarity information, the confidence level corresponding to each feature image in the feature group is determined, and based on the confidence level, a set of cut-off frames corresponding to the feature group is determined. Specifically, a pre-trained model can be used to process the temporal similarity information corresponding to the feature group to predict the confidence level corresponding to each feature image in the feature group. The pre-trained model can be, for example, a CLS model (Constrained Linear System). Reference image frames corresponding to feature images with confidence levels greater than or equal to a preset threshold can be used as cut-off frames. For a set of feature groups, a set of cut-off frames can be obtained, which may include one or more cut-off frames.

[0064] Since this embodiment determines the confidence level of the feature image in the feature group based on the temporal similarity information corresponding to the feature group, and determines a set of cut frames corresponding to the feature group based on the confidence level of the feature image in the feature group, the accuracy of the cut results is improved.

[0065] It should be noted that although the operations of the methods of this disclosure embodiment are described in a specific order in the above embodiments, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowcharts may be executed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0066] Corresponding to the aforementioned video processing method embodiments, this disclosure also provides embodiments of a video processing apparatus.

[0067] As shown in FIG5, FIG5 is a block diagram of a video processing apparatus according to an exemplary embodiment of the present disclosure. The apparatus may include: a frame extraction module 501, an extraction module 502, a determination module 503, and an output module 504.

[0068] The frame extraction module 501 is used to extract frames from the original video to be processed, so as to obtain a reference image frame sequence.

[0069] The extraction module 502 is used to extract features from the reference image frame sequence according to various different feature granularities to obtain multiple feature groups, each of which includes multiple feature images.

[0070] The determination module 503 is used to determine multiple sets of cut frames based on multiple feature groups.

[0071] The output module 504 is used to output the cutting results of the original video to the user based on multiple sets of cutting frames and according to different feature granularities.

[0072] In some implementations, the extraction module 502 is configured to: perform downsampling processing on the reference image frame sequence according to a variety of different downsampling ratios to obtain multiple feature groups at a variety of different feature granularities.

[0073] In other embodiments, the determining module 503 may include: a similarity submodule, a confidence submodule, and a mirror-cutting module (not shown in the figure).

[0074] The similarity submodule is used to determine the temporal similarity information corresponding to any feature group.

[0075] The confidence submodule is used to determine the confidence level of the feature images in any feature group based on the temporal similarity information corresponding to that feature group.

[0076] The slicing module is used to determine a set of slicing frames corresponding to any feature group based on the confidence level of the feature image in that feature group.

[0077] In other implementations, the similarity submodule is configured to: perform context fusion processing on the multi-frame feature images in the feature group respectively to obtain multi-frame fused feature images, and determine the temporal similarity information corresponding to the feature group based on the multi-frame fused feature images.

[0078] In other implementations, the similarity submodule can perform context fusion processing on multiple frames of feature images in the feature group to obtain multi-frame fused feature images in the following manner: obtain the context feature images before and after the frame feature image, and fuse the frame feature image and the context feature images with temporal features to obtain a single frame fused feature image.

[0079] In other implementations, the similarity submodule can determine the temporal similarity information corresponding to the feature group based on the multi-frame fused feature images in the following manner: calculating the pairwise similarity between the multi-frame fused feature images, and generating a temporal similarity map based on the pairwise similarity between the multi-frame fused feature images, which serves as the temporal similarity information corresponding to the feature group. The temporal similarity map includes multiple regions, each region corresponding to two frames of fused feature images, and the pixel values ​​of the pixels in the region represent the similarity between the two frames of fused feature images.

[0080] In other embodiments, the output module 504 is configured to: determine multiple cut times corresponding to the original video timeline based on multiple cut frames, as cut results for the original video. Each cut time corresponds to a feature granularity, and multiple cut times are output to the user according to different feature granularities.

[0081] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the embodiments of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0082] This disclosure provides an electronic device according to some embodiments. The electronic device includes a processor and a memory, and can be used to implement a client or server. The memory stores computer-executable instructions (e.g., one or more computer program modules) non-transitoryly. The processor executes the computer-executable instructions, which, when executed by the processor, can perform one or more steps of the video processing method described above, thereby implementing the video processing method described above. The memory and processor can be interconnected via a bus system and / or other forms of connection mechanisms (not shown).

[0083] For example, a processor can be a central processing unit (CPU), a graphics processing unit (GPU), or other form of processing unit with data processing and / or program execution capabilities. For instance, a CPU can be based on x86 or ARM architectures. A processor can be a general-purpose processor or a special-purpose processor, capable of controlling other components in an electronic device to perform desired functions.

[0084] For example, the memory may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), USB storage, flash memory, etc. One or more computer program modules may be stored on the computer-readable storage medium, and the processor may run one or more computer program modules to implement various functions of the electronic device. Various application programs and various data, as well as various data used and / or generated by the application programs, may also be stored in the computer-readable storage medium.

[0085] It should be noted that the specific functions and technical effects of the electronic devices in the embodiments of this disclosure can be referred to the description of the video processing method above, and will not be repeated here.

[0086] Figure 6 is a schematic block diagram of an electronic device provided in some embodiments of this disclosure. This electronic device 920 is, for example, suitable for implementing the video processing method provided in the embodiments of this disclosure. The electronic device 920 can be a terminal device, etc., and can be used to implement a client or server. The electronic device 920 can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), wearable electronic devices, etc., as well as fixed terminals such as digital TVs, desktop computers, smart home devices, etc. It should be noted that the electronic device 920 shown in Figure 6 is merely an example and does not impose any limitations on the functionality and scope of use of the embodiments of this disclosure.

[0087] As shown in Figure 6, the electronic device 920 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 921, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 922 or a program loaded from a storage device 928 into a random access memory (RAM) 923. The RAM 923 also stores various programs and data required for the operation of the electronic device 920. The processing unit 921, ROM 922, and RAM 923 are interconnected via a bus 924. An input / output (I / O) interface 925 is also connected to the bus 924.

[0088] Typically, the following devices can be connected to I / O interface 925: input devices 926 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 927 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 928 including, for example, magnetic tapes, hard disks, etc.; and communication devices 929. Communication device 929 allows electronic device 920 to communicate wirelessly or wiredly with other electronic devices to exchange data. Although Figure 6 shows electronic device 920 with various devices, it should be understood that it is not required to implement or have all the devices shown, and electronic device 920 may alternatively implement or have more or fewer devices.

[0089] For example, according to embodiments of this disclosure, the video processing method described above can be implemented as a computer software program. For instance, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program including program code for performing the video processing method described above. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 929, or installed from a storage device 928, or installed from a ROM 922. When the computer program is executed by the processing device 921, the functions defined in the video processing method provided by embodiments of this disclosure can be implemented.

[0090] This disclosure provides a storage medium according to some embodiments. For example, the storage medium may be a non-transitory computer-readable storage medium for storing non-transitory computer-executable instructions. When the non-transitory computer-executable instructions are executed by a processor, the video processing method described in the embodiments of this disclosure can be implemented. For example, when the non-transitory computer-executable instructions are executed by a processor, one or more steps in the video processing method described above can be performed.

[0091] For example, the storage medium can be used in the aforementioned electronic device; for instance, the storage medium may include the memory in the electronic device.

[0092] For example, the storage medium may include a memory card for a smartphone, a storage component for a tablet computer, a hard disk for a personal computer, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), portable compact disc read-only memory (CD-ROM), flash memory, or any combination of the above storage media, or other suitable storage media.

[0093] For example, the description of the storage medium can be found in the description of the memory in the embodiments of the electronic device, and will not be repeated here. The specific functions and technical effects of the storage medium can be found in the description of the video processing method above, and will not be repeated here.

[0094] It should be noted that, in the context of this disclosure, a computer-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, capable of transmitting, propagating, or transmitting a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0095] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0096] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.

Claims

1. A video processing method, comprising: Frames are extracted from the original video to be processed to obtain a reference image frame sequence; The reference image frame sequence is subjected to feature extraction according to various different feature granularities to obtain multiple feature groups; each feature group includes multiple feature images. Based on the aforementioned multiple feature groups, multiple sets of cut-off frames are determined; Based on the cut frames, the cut results for the original video are output to the user according to different feature granularities.

2. The method according to claim 1, wherein, The reference image frame is subjected to feature extraction according to multiple different feature granularities to obtain multiple feature groups, including: The reference image frame sequence is downsampled at various downsampling rates to obtain multiple feature groups at various feature granularities; wherein different feature groups correspond to different feature granularities.

3. The method according to claim 1 or 2, wherein, For any given feature group, a set of cut frames is determined based on the feature group in the following manner: Determine the temporal similarity information corresponding to the feature group; Based on the temporal similarity information, the confidence level corresponding to the feature image in the feature group is determined; Based on the confidence level, a set of cut frames corresponding to the feature group is determined.

4. The method according to claim 3, wherein, Determining the temporal similarity information corresponding to the feature group includes: Context fusion processing is performed on the feature images of the feature group to obtain multi-frame fused feature images; Based on the multi-frame fused feature image, the temporal similarity information corresponding to the feature group is determined.

5. The method according to claim 4, wherein, A fused feature image is obtained by performing context fusion processing on any frame of feature image in the feature group in the following manner: Obtain the context feature images before and after the frame feature image; The frame feature image and the context feature image are fused using temporal features to obtain a fused feature image.

6. The method according to claim 4 or 5, wherein, The step of determining the temporal similarity information corresponding to the feature group based on the multi-frame fused feature image includes: Calculate the pairwise similarity between the multi-frame fused feature images; Based on the pairwise similarity between the multi-frame fused feature images, a temporal similarity map is generated as the temporal similarity information corresponding to the feature group; wherein, the temporal similarity map includes multiple regions, each region corresponding to two frames of fused feature images, and the pixel value of the pixel in the region represents the similarity between the two frames of fused feature images.

7. The method according to any one of claims 1-6, wherein, The step of outputting the cut results for the original video to the user based on the multiple sets of cut frames and according to different feature granularities includes: Based on the multiple sets of cut frames, determine the corresponding multiple sets of cut times on the original video timeline as the cut results for the original video; wherein, each set of cut times corresponds to a feature granularity. According to the different feature granularities, the multiple sets of cut-off times are output to the user.

8. A video processing apparatus, comprising: The frame extraction module is configured to extract frames from the original video to be processed, thereby obtaining a reference image frame sequence. The extraction module is configured to extract features from the reference image frame sequence according to multiple different feature granularities to obtain multiple feature groups; The feature group includes multiple frames of feature images; The determination module is configured to determine multiple sets of cut-off frames based on the multiple feature groups; The output module is configured to output the cut results of the original video to the user based on the cut frame and according to different feature granularities.

9. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed in a computer, the computer performs the video processing method according to any one of claims 1-7.

10. An electronic device comprising a memory and a processor, wherein, The memory stores executable code, and when the processor executes the executable code, it implements the video processing method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Video splitting method and device, equipment and computer readable storage medium

    CN113408332A

  • Video scene segmentation method, device and equipment and computer readable storage medium

    CN114283351A

  • Video scene segmentation method and related product

    CN117376603A

  • Fine-grain object segmentation in video with deep features and multi-level graphical models

    US20200074185A1