Video shot boundary localization method, device and electronic device
Through the neural network model, the problem of inaccurate positioning of adjacent lenses in the prior art is solved, and higher positioning accuracy is achieved.
Patent Information
- Application Number
- CN202110923476.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-12
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-08-12
AI Technical Summary
The existing video lens boundary positioning methods are not accurate enough when dealing with adjacent lenses with small overall motion changes but severe local motion changes.
The neural network model is used to predict the initial boundary frame and the initial predicted boundary frame is subdivided by the changes of the chunked average gradient matrix to improve positioning accuracy.
Without significantly losing positioning speed, the accuracy of lens boundary positioning is improved, and the actual test effect is significant.
Smart Images

Figure CN113610821B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and in particular, to a method, device and electronic device for video shot boundary localization. Background Art
[0002] Video shot boundary localization is one of the important steps in video content understanding. A video shot is a relatively independent video unit semantically. In terms of time sequence, a shot is a set of continuous actions of objects within a frame in the time domain. Shot boundary localization refers to the process of detecting and locating the boundary frames of a shot and segmenting the video into independent shots.
[0003] The basic idea of a shot boundary localization algorithm is to determine the boundary frames of a shot based on the physical feature differences between adjacent shots. In order to make the video shot transition smoother, several buffer frames are often inserted between two adjacent shots, and the overall visual change is not significant, making boundary detection a difficult problem. In addition, during the video shooting process, there are situations such as camera jitter, noise, and light intensity change, which also greatly affect the effect of shot boundary positioning.
[0004] Current shot boundary localization methods include methods based on histogram difference and methods based on deep learning. Among them, a histogram is a method for describing the distribution of image color features. By the similarity of histograms, the similarity between images can be judged, and based on this, whether there is a critical change in the image scene can be judged to achieve shot boundary localization. However, due to the very complex transformation of video shots, especially for adjacent shots with small overall motion changes but large local motion changes, the existing methods based on histogram difference and methods based on deep learning are not accurate enough in localization. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, device and electronic device for video shot boundary localization to improve the localization accuracy.
[0006] An embodiment of the present invention provides a method for video shot boundary localization, including:
[0007] Determine at least one initial boundary frame of a target video to be detected according to a trained neural network model;
[0008] For each of the initial boundary frames, calculate the block average gradient of each target video frame in the video frame sequence corresponding to the initial boundary frame; wherein, each target video frame in the video frame sequence corresponding to the initial boundary frame is each video frame within a preset range centered on the initial boundary frame in the target video; the block average gradient of the target video frame includes the gradient of the average pixel gray value of each sub-region when the target video frame is divided into multiple sub-regions;
[0009] Determine the target boundary frame corresponding to the initial boundary frame from each of the target video frames according to the block average gradient of each of the target video frames in the video frame sequence corresponding to the initial boundary frame;
[0010] Determine the target boundary frames corresponding to each of the initial boundary frames as the shot boundary localization result of the target video.
[0011] Further, the calculating the block average gradient of each target video frame in the video frame sequence corresponding to the initial boundary frame includes:
[0012] For each target video frame in the video frame sequence corresponding to the initial boundary frame, divide the target video frame into a plurality of sub-regions according to a preset partitioning method;
[0013] Calculate the average pixel gray value of each of the sub-regions;
[0014] According to the average pixel gray values of the sub-regions, calculate the gradient amplitude corresponding to each of the sub-regions;
[0015] Construct a block average gradient matrix of the target video frame from the gradient amplitudes corresponding to the sub-regions.
[0016] Further, the block average gradient of the target video frame is a block average gradient matrix constructed from the gradient amplitudes of the average pixel gray values of the sub-regions of the target video frame; the determining the target boundary frame corresponding to the initial boundary frame from each of the target video frames according to the block average gradients of each of the target video frames in the video frame sequence corresponding to the initial boundary frame includes:
[0017] Calculate the boundary feature matrix of each target video frame in the video frame sequence corresponding to the initial boundary frame, where the boundary feature matrix of the target video frame is the difference between the block average gradient matrix of the target video frame and the block average gradient matrix of the previous video frame of the target video frame;
[0018] Count the number of target elements corresponding to each boundary feature matrix, where the number of target elements corresponding to the boundary feature matrix is the number of elements in the boundary feature matrix whose element values are greater than a preset gradient threshold;
[0019] Determine the target boundary frame corresponding to the initial boundary frame according to the number of target elements corresponding to each boundary feature matrix.
[0020] Further, the determining the target boundary frame corresponding to the initial boundary frame according to the number of target elements corresponding to each boundary feature matrix includes:
[0021] Determine the target video frame corresponding to the boundary feature matrix with the number of target elements greater than the preset number as the candidate boundary frame;
[0022] When there is one such candidate boundary frame, determine the candidate boundary frame as the target boundary frame corresponding to the initial boundary frame;
[0023] When there are multiple such candidate boundary frames, determine the last candidate boundary frame in the video frame sequence as the target boundary frame corresponding to the initial boundary frame.
[0024] Further, the method further includes:
[0025] Obtain a plurality of sample videos and the boundary frame annotation data of each sample video, and each sample video includes a preset number of video frames;
[0026] Train the neural network model to be trained according to the plurality of sample videos and their boundary frame annotation data to obtain the trained neural network model.
[0027] Further, the obtaining a plurality of sample videos and the boundary frame annotation data of each sample video includes:
[0028] Obtain the original video and its boundary frame annotation data from the ClipShots dataset;
[0029] Perform splitting processing on the original video and its boundary frame annotation data to obtain a plurality of sample videos and the boundary frame annotation data of each sample video.
[0030] An embodiment of the present invention further provides a video shot boundary positioning device, including:
[0031] The first determination module is used to determine at least one initial boundary frame of the target video to be detected according to the trained neural network model;
[0032] The gradient calculation module is used to calculate the block average gradient of each target video frame in the video frame sequence corresponding to each initial boundary frame; wherein, each target video frame in the video frame sequence corresponding to the initial boundary frame is each video frame within a preset range centered on the initial boundary frame in the target video; the block average gradient of the target video frame includes the gradient of the average pixel gray value of each sub-region when the target video frame is divided into a plurality of sub-regions;
[0033] The second determination module is used to determine the target boundary frame corresponding to the initial boundary frame from each target video frame according to the block average gradients of each target video frame in the video frame sequence corresponding to the initial boundary frame;
[0034] A third determination module, configured to determine the target boundary frames corresponding to each of the initial boundary frames as the shot boundary localization result of the target video.
[0035] Further, the gradient calculation module is specifically configured to:
[0036] For each target video frame in the video frame sequence corresponding to the initial boundary frame, divide the target video frame into a plurality of sub-regions according to a preset division method;
[0037] Calculate the average pixel gray value of each of the sub-regions;
[0038] According to the average pixel gray values of the sub-regions, calculate the gradient amplitude corresponding to each of the sub-regions;
[0039] Construct a block-average gradient matrix of the target video frame from the gradient amplitudes corresponding to the sub-regions.
[0040] An embodiment of the present invention further provides an electronic device, including a memory and a processor. A computer program that can run on the processor is stored in the memory. When the processor executes the computer program, the above video shot boundary localization method is implemented.
[0041] An embodiment of the present invention further provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is run by a processor, the above video shot boundary localization method is executed.
[0042] In the video shot boundary positioning method, device and electronic device provided by the embodiments of the present invention, the method comprises: determining at least one initial boundary frame of a target video to be detected according to a trained neural network model; for each initial boundary frame, determining the block average gradient of each target video frame in a video frame sequence corresponding to the initial boundary frame; wherein each target video frame in the video frame sequence corresponding to the initial boundary frame is each video frame in a preset range centered on the initial boundary frame in the target video; the block average gradient of the target video frame comprises the gradient of the average pixel grayscale value of each sub-region when the target video frame is divided into a plurality of sub-regions; according to the block average gradient of each target video frame in the video frame sequence corresponding to the initial boundary frame, determining the target boundary frame corresponding to the initial boundary frame from each target video frame; and determining the target boundary frame corresponding to each initial boundary frame as the shot boundary positioning result of the target video. The embodiment of the present invention uses a neural network model to predict the initial boundary frame, and adds a post-processing based on block average gradient to solve the problem of inaccurate positioning of adjacent shot boundaries with small overall motion changes but drastic local motion changes. Compared with the existing histogram difference-based method and deep learning-based method, the positioning accuracy is improved without significantly losing the positioning speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the specific implementation methods of the present invention or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0044] Figure 1 A schematic diagram of a flow chart of a method for locating a video shot boundary provided by an embodiment of the present invention;
[0045] Figure 2 A schematic flow chart of another method for locating a video shot boundary provided by an embodiment of the present invention;
[0046] Figure 3 A schematic diagram of the structure of a neural network provided by an embodiment of the present invention;
[0047] Figure 4 A schematic diagram of the structure of a video shot boundary positioning device provided by an embodiment of the present invention;
[0048] Figure 5 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0049] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative work belong to the protection scope of the present invention.
[0050] In an advertising video, in order not to disrupt the continuity of the shots, it is necessary to quickly and accurately locate the shot boundaries so as to insert the content provided by the advertiser after a complete shot. Currently, methods based on deep learning are commonly used to predict shot boundaries. However, due to the very complex transformation of video shots, the positioning of this method is not accurate enough. Based on this, a robust video shot boundary positioning method, device, and electronic device provided by the embodiments of the present invention can effectively improve the boundary positioning accuracy without significantly sacrificing the positioning speed, and have achieved good results in practical applications.
[0051] For the convenience of understanding this embodiment, a video shot boundary positioning method disclosed in the embodiments of the present invention will be introduced in detail first.
[0052] The embodiments of the present invention provide a video shot boundary positioning method, which can be executed by an electronic device with data processing capabilities. The electronic device can be a mobile phone, a laptop, a desktop computer, etc. Refer to Figure 1 The flowchart of a video shot boundary positioning method shown, the method mainly includes the following steps S102 to step S108:
[0053] Step S102, determine at least one initial boundary frame of the target video to be detected according to the trained neural network model.
[0054] Input the target video to be detected into the trained neural network model, and obtain at least one initial boundary frame output by the neural network model. The neural network model can be, but is not limited to, a model based on 3D CNN (Convolutional Neural Networks). By reasoning the target video to be detected through a model such as a 3D CNN-based model, the probability (range 0-1) that each video frame in the target video is a boundary frame can be obtained. The model discretizes the output probability with a threshold of 0.5 to 0 or 1, where 0 represents a non-boundary frame and 1 represents a boundary frame.
[0055] Step S104: For each initial boundary frame, determine the block average gradient of each target video frame in the video frame sequence corresponding to the initial boundary frame; wherein, each target video frame in the video frame sequence corresponding to the initial boundary frame is each video frame within a preset range centered on the initial boundary frame in the target video; the block average gradient of a target video frame includes the gradient of the average pixel gray value of each sub-region when the target video frame is divided into multiple sub-regions.
[0056] The video frame sequence corresponding to the initial boundary frame can be determined by the windowing method, and the above preset range is equal to the length of the window. The length of the window can be set according to actual needs. For example, if the length of the window is 32, then a window with a length of 32 can be added centered at the position where the model output is 1 in step S102 (i.e., the position where the initial boundary frame is located), with 16 frames on each side of the position where the model output is 1. These 33 video frames within the window constitute the video frame sequence corresponding to the initial boundary frame.
[0057] For the sake of easy understanding, in some possible embodiments, the above step S104 can be implemented through the following sub-steps 1 to sub-step 4:
[0058] Sub-step 1: For each target video frame in the video frame sequence corresponding to the initial boundary frame, divide the target video frame into multiple sub-regions according to a preset division method.
[0059] The above preset division method can be set according to actual needs. For example, each target video frame is evenly divided into 16×16 sub-regions.
[0060] Sub-step 2: Calculate the average pixel gray value of each sub-region.
[0061] Calculate the average value of the gray values of each pixel point in each sub-region to obtain the average pixel gray value of the sub-region.
[0062] Sub-step 3: According to the average pixel gray values of each sub-region, calculate the gradient amplitude corresponding to each sub-region.
[0063] The gradient (i.e., the first-order differential) of the average pixel gray value function H(u, v) in the sub-region (u, v) is a vector with magnitude and direction. Let G u and G v represent the gradients along the u direction and the v direction respectively. Then this gradient vector can be expressed by the following formula 1:
[0064]
[0065] The magnitude of this vector (i.e., the gradient amplitude) can be expressed by the following formula 2:
[0066]
[0067] Sub-step 4: Construct the block-average gradient matrix of the target video frame from the gradient magnitudes corresponding to each sub-region.
[0068] Taking the example that the target video frame is divided into 16×16 sub-regions, the block-average gradient matrix M of the target video frame i can be expressed by the following formula 3:
[0069]
[0070] where M i represents the block-average gradient matrix of the target video frame f i , and G x,y (1≤x≤16, 1≤y≤16) represents the gradient magnitude corresponding to the sub-region (x, y).
[0071] Step S106: Determine the target boundary frame corresponding to the initial boundary frame from each target video frame in the video frame sequence corresponding to the initial boundary frame according to the block-average gradients of each target video frame.
[0072] For ease of understanding, in some possible embodiments, the above step S106 can be implemented through the following sub-steps a to c:
[0073] Sub-step a: Calculate the boundary feature matrix of each target video frame in the video frame sequence corresponding to the initial boundary frame. The boundary feature matrix of the target video frame is the difference between the block-average gradient matrix of the target video frame and the block-average gradient matrix of the previous video frame of the target video frame.
[0074] The boundary feature matrices of each target video frame in the window can be calculated in sequence according to the arrangement order of the target video frames in the video frame sequence (such as from left to right). Since the block-average gradient matrix of the previous video frame of the first target video frame is not calculated, the first target video frame can be not considered when calculating the boundary feature matrix. The boundary feature matrix is defined as the difference between the block-average gradient matrices of adjacent frames and can be expressed by the following formula 4:
[0075] F i = M i - M i-1 (Formula 4)
[0076] where F i represents the boundary feature matrix of the target video frame f i , M i represents the block-average gradient matrix of the target video frame f i , and M i-1 represents the block-average gradient matrix of the target video frame f i-1The block-averaged gradient matrix.
[0077] Sub-step b: Statistically obtain the number of target elements corresponding to each boundary feature matrix. The number of target elements corresponding to the boundary feature matrix is the number of elements in the boundary feature matrix whose element values are greater than a preset gradient threshold.
[0078] The preset gradient threshold can be set according to actual needs and is not limited here. For example, if the preset gradient threshold is 0.3, then count the number of elements in each boundary feature matrix whose element values are greater than 0.3 to obtain the number of target elements corresponding to the boundary feature matrix.
[0079] Sub-step c: Determine the target boundary frame corresponding to the initial boundary frame according to the number of target elements corresponding to each boundary feature matrix.
[0080] For the above sub-step c, in one possible implementation, the target video frame corresponding to the boundary feature matrix with the number of target elements greater than a preset quantity can be determined as a candidate boundary frame; when there is one candidate boundary frame, determine this candidate boundary frame as the target boundary frame corresponding to the initial boundary frame; when there are multiple candidate boundary frames, determine the last candidate boundary frame in the video frame sequence as the target boundary frame corresponding to the initial boundary frame.
[0081] The setting of the above preset quantity is related to the number of sub-regions divided. For example, the preset quantity is set to 0.6 times the number of sub-regions, rounded up or down. If the number of sub-regions is 256, the preset quantity can be 153 or 154. Taking the preset quantity of 154 as an example, if the number of target elements corresponding to a boundary feature matrix is 188, then determine the target video frame corresponding to this boundary feature matrix as a candidate boundary frame.
[0082] For the above sub-step c, in another possible implementation, the number of target elements corresponding to each boundary feature matrix can be divided by the total number of elements in the boundary feature matrix to obtain the target ratio corresponding to each boundary feature matrix; the target video frame corresponding to the boundary feature matrix with the target ratio greater than a preset ratio can be determined as a candidate boundary frame; when there is one candidate boundary frame, determine this candidate boundary frame as the target boundary frame corresponding to the initial boundary frame; when there are multiple candidate boundary frames, determine the last candidate boundary frame in the video frame sequence as the target boundary frame corresponding to the initial boundary frame.
[0083] Similarly, the above preset ratio can be set according to actual needs and is not limited here. For example, the preset ratio is 0.6.
[0084] Step S108: Determine the target boundary frame corresponding to each initial boundary frame as the shot boundary localization result of the target video.
[0085] The video shot boundary localization method provided by the embodiments of the present invention ensures relatively accurate initial prediction by predicting the initial boundary frame through a neural network model, and uses the change of the block-average gradient matrix to subdivide the initially predicted boundary frame, solving the problem that the adjacent shot boundary localization is inaccurate when the overall motion change is small but the local motion change is drastic. Compared with the existing methods based on histogram difference and deep learning methods, the localization accuracy is improved without significantly sacrificing the localization speed, and the actual test shows a significant improvement effect.
[0086] For ease of understanding, the embodiments of the present invention also provide another video shot boundary localization method, which adopts a two-stage scheme: First, use a model architecture based on 3D CNN to infer the input target video to obtain the initial shot boundary localization result (i.e., the initial boundary frame); then apply windowing and block-gradient-based subdivision techniques on the initial shot boundary localization result to improve the localization accuracy, as Figure 2 shown. The specific steps are as follows: Input the target video to be detected into the model based on 3D CNN to obtain at least one initial boundary frame; centered on the position of the initial boundary frame, add a window with a length of 32 (i.e., Figure 2 the windowing process in); Calculate the block-average gradient matrix of all target video frames within the window, calculate the boundary feature matrix of each target video frame in the window from left to right, and count the number of elements N i in each boundary feature matrix that are greater than the preset gradient threshold. N i represents the number of target elements corresponding to the target video frame f i . If N i / total number of elements > 0.6, then update the current target video frame f i to the target boundary frame. If there are multiple target boundary frames in the window, select the last one as the target boundary frame (i.e., Figure 2 the subdivision process in). In this way, an accurate boundary localization result can be generated for the target video to be detected. For the parts not described in detail here, reference can be made to the corresponding content of the foregoing embodiments, and details will not be repeated here.
[0087] The embodiments of the present invention also provide a training method for the above convolutional neural network, including the following process: Obtain a plurality of sample videos and the boundary frame annotation data of each sample video, and each sample video includes a preset number of video frames; According to the plurality of sample videos and their boundary frame annotation data, train the neural network model to be trained to obtain the trained neural network model.
[0088] The above preset number of frames is related to the input data setting of the input layer in the neural network model, and the preset number of frames can be set according to actual needs and is not limited here. For example, the preset number of frames is 128.
[0089] In a possible implementation, the original video and its boundary frame annotation data can be obtained from the ClipShots dataset; the original video and its boundary frame annotation data are split to obtain multiple sample videos and the boundary frame annotation data for each sample video. For example, using the ClipShots dataset for training, the original videos in the ClipShots dataset are split into multiple sample videos of 128 frames, and those with less than 128 frames are discarded or filled with the last frame repeatedly. In this way, using the existing ClipShots dataset to obtain training data can reduce the workload of users.
[0090] See Figure 3 The structural schematic diagram of a neural network shown in the figure. This neural network is a 3D CNN, which mainly includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a fourth convolutional layer, a fifth convolutional layer, a global average pooling layer, a first linear layer, and a second linear layer arranged in sequence. The vector dimension 128 output by the second linear layer is the same as the above-mentioned preset number of frames. When training the neural network model, each iteration loads a batch of training data and inputs it into the 3D CNN shown in Figure 3 the figure for training. The 3D CNN outputs a vector sequence of 128 dimensions, and this vector sequence represents the probability that each frame is a boundary frame.
[0091] In another possible implementation, for fragmented short-shot videos, manual annotation can also be performed to obtain boundary frame annotation data for training the neural network model.
[0092] In summary, the embodiments of the present invention propose a fast and accurate video shot boundary localization technical process, which can be effectively applied to various types of videos to solve the problems of insufficient robustness and accuracy in the existing technologies, and can make up for the deficiencies of the existing technologies.
[0093] Corresponding to the above video shot boundary localization method, the embodiments of the present invention also provide a video shot boundary localization device. See Figure 4 the structural schematic diagram of a video shot boundary localization device shown in the figure. This device includes:
[0094] A first determination module 42, configured to determine at least one initial boundary frame of a target video to be detected according to the trained neural network model;
[0095] The gradient calculation module 44 is configured to calculate the block-average gradient of each target video frame in the video frame sequence corresponding to the initial boundary frame for each initial boundary frame. Among them, each target video frame in the video frame sequence corresponding to the initial boundary frame is each video frame within a preset range centered on the initial boundary frame in the target video. The block-average gradient of a target video frame includes the gradient of the average pixel gray value of each sub-region when the target video frame is divided into multiple sub-regions.
[0096] The second determination module 46 is configured to determine the target boundary frame corresponding to the initial boundary frame from each target video frame according to the block-average gradients of each target video frame in the video frame sequence corresponding to the initial boundary frame.
[0097] The third determination module 48 is configured to determine the target boundary frames corresponding to each initial boundary frame as the shot boundary localization result of the target video.
[0098] The video shot boundary localization device provided by the embodiments of the present invention ensures relatively accurate initial prediction by predicting the initial boundary frame through a neural network model, and uses the change of the block-average gradient matrix to subdivide the initially predicted boundary frame, solving the problem of inaccurate localization of adjacent shot boundaries with small overall motion changes but large local motion changes. Compared with the existing methods based on histogram difference and deep learning methods, the localization accuracy is improved without significantly sacrificing the localization speed, and the actual test shows a significant improvement effect.
[0099] Further, the gradient calculation module 44 is specifically configured to: for each target video frame in the video frame sequence corresponding to the initial boundary frame, divide the target video frame into multiple sub-regions according to a preset division method; calculate the average pixel gray value of each sub-region; calculate the gradient amplitude corresponding to each sub-region according to the average pixel gray values of each sub-region; and construct a block-average gradient matrix of the target video frame from the gradient amplitudes corresponding to each sub-region.
[0100] Further, the block-average gradient of the target video frame is a block-average gradient matrix constructed from the gradient amplitudes of the average pixel gray values of each sub-region of the target video frame. The second determination module 46 is specifically configured to: calculate the boundary feature matrix of each target video frame in the video frame sequence corresponding to the initial boundary frame, where the boundary feature matrix of the target video frame is the difference between the block-average gradient matrix of the target video frame and the block-average gradient matrix of the previous video frame of the target video frame; count the number of target elements corresponding to each boundary feature matrix, where the number of target elements corresponding to the boundary feature matrix is the number of elements with element values greater than a preset gradient threshold in the boundary feature matrix; and determine the target boundary frame corresponding to the initial boundary frame according to the number of target elements corresponding to each boundary feature matrix.
[0101] Further, the second determination module 46 is further configured to: determine a target video frame corresponding to a boundary feature matrix with a target element number greater than a preset number as a candidate boundary frame; when there is one candidate boundary frame, determine the candidate boundary frame as the target boundary frame corresponding to the initial boundary frame; when there are multiple candidate boundary frames, determine the last candidate boundary frame in the video frame sequence as the target boundary frame corresponding to the initial boundary frame.
[0102] Further, the apparatus further includes a model training module connected to the first determination module 42. The model training module is configured to: obtain a plurality of sample videos and boundary frame annotation data of each sample video, where each sample video includes a preset number of video frames; train a neural network model to be trained according to the plurality of sample videos and their boundary frame annotation data to obtain a trained neural network model.
[0103] Further, the model training module is specifically configured to: obtain an original video and its boundary frame annotation data from the ClipShots dataset; perform splitting processing on the original video and its boundary frame annotation data to obtain a plurality of sample videos and boundary frame annotation data of each sample video.
[0104] The apparatus provided in this embodiment has the same implementation principle and the same technical effects as those in the foregoing method embodiment. For a brief description, for the parts not mentioned in the apparatus embodiment, reference may be made to the corresponding content in the foregoing method embodiment.
[0105] See Figure 5 , the embodiment of the present invention further provides an electronic device 100, including: a processor 50, a memory 51, a bus 52, and a communication interface 53. The processor 50, the communication interface 53, and the memory 51 are connected through the bus 52; the processor 50 is configured to execute an executable module stored in the memory 51, such as a computer program.
[0106] Among them, the memory 51 may include a random access memory (Random Access Memory, abbreviated as RAM), and may also include a non-volatile memory (non-volatile memory, abbreviated as NVM), such as at least one disk memory. Through at least one communication interface 53 (which may be wired or wireless), a communication connection between the system network element and at least one other network element is realized, and the Internet, a wide area network, a local area network, a metropolitan area network, etc. can be used.
[0107] The bus 52 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 5 only a bidirectional arrow is used in Figure 5 , but it does not mean that there is only one bus or one type of bus.
[0108] Among them, the memory 51 is used to store programs. After receiving an execution instruction, the processor 50 executes the program. The method executed by the device defined by the process disclosed in any one of the foregoing embodiments of the present invention can be applied to the processor 50 or implemented by the processor 50.
[0109] The processor 50 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 50 or the instructions in the form of software. The above-mentioned processor 50 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 51, and the processor 50 reads the information in the memory 51 and combines its hardware to complete the steps of the above method.
[0110] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, it executes the video shot boundary localization method described in the foregoing method embodiments. The computer-readable storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM for short), RAMs, magnetic disks, or optical discs.
[0111] In all the examples shown and described herein, any specific value should be construed as merely exemplary, rather than as a limitation. Therefore, other examples of the exemplary embodiments may have different values.
[0112] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of methods and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0113] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings, direct couplings, or communication connections to each other can be through some communication interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0114] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0115] In addition, in each embodiment of the present invention, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit.
[0116] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for video shot boundary localization, characterized in that, it includes: Determine at least one initial boundary frame of the target video to be detected according to the trained neural network model; For each of the initial boundary frames, calculate the block-average gradient of each target video frame in the video frame sequence corresponding to the initial boundary frame; wherein, each target video frame in the video frame sequence corresponding to the initial boundary frame is each video frame within a preset range centered on the initial boundary frame in the target video; the block-average gradient of the target video frame includes the gradient of the average pixel gray value of each sub-region when the target video frame is divided into multiple sub-regions; According to the block-average gradients of each target video frame in the video frame sequence corresponding to the initial boundary frame, determine the target boundary frame corresponding to the initial boundary frame from each target video frame, including: calculate the boundary feature matrix of each target video frame in the video frame sequence corresponding to the initial boundary frame, the boundary feature matrix of the target video frame is the difference between the block-average gradient matrix of the target video frame and the block-average gradient matrix of the previous video frame of the target video frame; count the number of target elements corresponding to each boundary feature matrix, the number of target elements corresponding to the boundary feature matrix is the number of elements whose element values are greater than the preset gradient threshold in the boundary feature matrix; according to the number of target elements corresponding to each boundary feature matrix, determine the target boundary frame corresponding to the initial boundary frame, the block-average gradient of the target video frame is a block-average gradient matrix constructed by the gradient amplitudes of the average pixel gray values of each sub-region of the target video frame; Determine the target boundary frames corresponding to each of the initial boundary frames as the shot boundary localization result of the target video.
2. The video shot boundary localization method according to claim 1, characterized in that, the calculating the block-average gradient of each target video frame in the video frame sequence corresponding to the initial boundary frame includes: For each target video frame in the video frame sequence corresponding to the initial boundary frame, divide the target video frame into multiple sub-regions according to a preset division method; Calculate the average pixel gray value of each sub-region; According to the average pixel gray values of each sub-region, calculate the gradient amplitude corresponding to each sub-region; Construct a block-average gradient matrix of the target video frame from the gradient amplitudes corresponding to each sub-region.
3. The video shot boundary localization method according to claim 1, characterized in that, the determining the target boundary frame corresponding to the initial boundary frame according to the number of target elements corresponding to each boundary feature matrix includes: Determine the target video frames corresponding to the boundary feature matrices with the number of target elements greater than the preset number as candidate boundary frames; When there is one candidate boundary frame, determine the candidate boundary frame as the target boundary frame corresponding to the initial boundary frame; When there are multiple candidate boundary frames, determine the last candidate boundary frame in the video frame sequence as the target boundary frame corresponding to the initial boundary frame.
4. The video shot boundary localization method according to claim 1, characterized in that, the method further includes: obtaining a plurality of sample videos and boundary frame annotation data of each of the sample videos, each of the sample videos including a preset number of video frames; training a neural network model to be trained according to the plurality of sample videos and their boundary frame annotation data to obtain a trained neural network model.
5. The video shot boundary localization method according to claim 4, characterized in that, the obtaining of a plurality of sample videos and boundary frame annotation data of each of the sample videos includes: obtaining an original video and its boundary frame annotation data from the ClipShots dataset; performing a splitting process on the original video and its boundary frame annotation data to obtain a plurality of sample videos and boundary frame annotation data of each of the sample videos.
6. A video shot boundary localization device, characterized in that, it includes: a first determination module, configured to determine at least one initial boundary frame of a target video to be detected according to a trained neural network model; a gradient calculation module, configured to calculate the block average gradient of each target video frame in the video frame sequence corresponding to the initial boundary frame for each of the initial boundary frames; wherein, each target video frame in the video frame sequence corresponding to the initial boundary frame is each video frame within a preset range centered on the initial boundary frame in the target video; the block average gradient of the target video frame includes the gradient of the average pixel gray value of each sub-region when the target video frame is divided into a plurality of sub-regions; a second determination module, configured to determine the target boundary frame corresponding to the initial boundary frame from each of the target video frames according to the block average gradients of each of the target video frames in the video frame sequence corresponding to the initial boundary frame: calculating the boundary feature matrix of each of the target video frames in the video frame sequence corresponding to the initial boundary frame, the boundary feature matrix of the target video frame being the difference between the block average gradient matrix of the target video frame and the block average gradient matrix of the previous video frame of the target video frame; counting the number of target elements corresponding to each of the boundary feature matrices, the number of target elements corresponding to the boundary feature matrix being the number of elements in the boundary feature matrix whose element values are greater than a preset gradient threshold; determining the target boundary frame corresponding to the initial boundary frame according to the number of target elements corresponding to each of the boundary feature matrices, the block average gradient of the target video frame being a block average gradient matrix constructed by the gradient amplitudes of the average pixel gray values of each of the sub-regions of the target video frame; a third determination module, configured to determine the target boundary frames corresponding to each of the initial boundary frames as the shot boundary localization result of the target video.
7. The video shot boundary localization device according to claim 6, characterized in that, the gradient calculation module is specifically configured to: for each target video frame in the video frame sequence corresponding to the initial boundary frame, divide the target video frame into a plurality of sub-regions according to a preset division method; calculate the average pixel gray value of each of the sub-regions; According to the average pixel gray value of each of the sub-regions, the gradient amplitude corresponding to each of the sub-regions is calculated; From the gradient amplitudes corresponding to each of the sub-regions, a block-average gradient matrix of the target video frame is constructed.
8. An electronic device, comprising a memory and a processor, wherein a computer program capable of running on the processor is stored in the memory, characterized in that, when the processor executes the computer program, the method according to any one of claims 1-5 is implemented.
9. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is run by a processor, the method according to any one of claims 1-5 is executed.
Citation Information
Patent Citations
Movement identification method and system based on gradient boundary map and multi-mode convolutional fusion
CN108288016A
Time sequence action fragment segmentation method based on boundary search intelligent agent
CN111950393A