A method, system and storage medium for generating video descriptions based on grid graphs
By adopting a grid diagram-based method in video description generation, the problem of incoherence of computing complexity and inter-frame information in the prior art is solved, and efficient video description generation and quality improvement are achieved.
Patent Information
- Application Number
- CN202510300684.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-03-14
AI Technical Summary
The existing video description generation scheme has problems with computing complexity, computing resource consumption, and generation time, and inter-frame information is incoherent.
Using a video description generation method based on grid diagrams, the first image of k-frames is extracted by equally spaced, the optical flow fraction is calculated, the empty graph is constructed, and the image blocks are spliced in grid mode, and adjusted to a single grid-like image to input into the LVLM model.
The calculation complexity, computing resource consumption and generation time of video description generation are reduced, while ensuring the consistency of inter-frame information and improving the generation quality of video description.
Smart Images

Figure CN119815139B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video description, and particularly to a method, a system and a storage medium for generating video description based on a grid graph. Background Art
[0002] Video description generation refers to the process in which a computer automatically generates a text description for a video. Existing video description generation schemes are all based on the LVLM model (Large Vision Language Model), that is, video frames are sent into the LVLM model, and a generation-type text prompt P is attached, and the LVLM model will output the corresponding video description.
[0003] The general steps of existing video description generation schemes are as follows:
[0004] Extract k frame images, each frame image has a size of W × H × C , and adjust each frame image to an image of T × T × C ( T usually 336 / 448), and splice them into a video frame sequence to obtain a video image with a size of k × T × T × C as the input of the LVLM model, and then k × T × T × C of the video image and the generation-type text prompt P are input into the LVLM model together to obtain the generated video description.
[0005] Existing video description generation schemes have the following defects:
[0006] Extract k frame images and splice them into a k × T × T × C video image as the input of the LVLM model. As the number of extracted frames k increases, it will cause a substantial increase in the input data of the LVLM model, thus greatly increasing the problems of computational complexity, computational resource consumption, and too long generation time; if fewer frames k of images are extracted, it will cause too long a time span between frames, resulting in the problem of discontinuous inter-frame information. Summary of the Invention
[0007] The present invention proposes a video description generation method, system and storage medium based on a grid graph to solve the technical problems of existing video description generation computational complexity, computational resource consumption, long generation time and incoherent information between frames.
[0008] One aspect of the present invention is to provide a method for generating video description based on a grid graph, the method comprising the following steps:
[0009] S101. Obtaining original video V ;
[0010] S102, from the original video obtained V Medium pitch extraction k Frame first image; wherein, k It should satisfy the square root;
[0011] The size of the first image in each frame is W × H × C ;in, W Indicates the length of the first image in each frame, H Indicates the width of the first image in each frame, C Represents the original video V The number of channels;
[0012] S103, using the extracted k Frame first image, calculate the original video V If the optical flow score is lower than the preset threshold, the original video is discarded. V , re-acquire the original video V ;
[0013] If the optical flow score is higher than the preset threshold, a size of W × H × C Empty map I , the empty map I Divide into The size is W × H × C of blocks; among them, k The original video obtained from V The number of frames of the first image extracted with medium spacing;
[0014] S104, will k The first image of the frame is placed in the empty image in order from left to right and from top to bottom. I of k The size is W × H ×C from the block of W × H × C to obtain a second image of
[0015] S105. Resize the second image of size W × H × C to a third image of size T × T × C ; where T represents the length and width of the third image;
[0016] S106. Input the third image of size T × T × C and the generated class text into the LVLM model to output the generated video description.
[0017] In a preferred embodiment, in S102, from the obtained original video V equidistantly extract k = 9 frames or k = 16 frames of the first image.
[0018] In a preferred embodiment, it is characterized in that, in S102, from the obtained original video V equidistantly extract k = 25 frames of the first image.
[0019] In a preferred embodiment, in step S103, calculate the optical flow fraction of the original video V by the following method:
[0020] S1031. Input the extracted k frames of the first image into a deep learning model for optical flow calculation to obtain k-1 frames of an optical flow feature map of size W × H × 2;
[0021] where W represents the length of each frame of the first image, H represents the width of each frame of the first image, and the number of channels of the optical flow feature map is 2, representing the horizontal and vertical directions of the optical flow feature map respectively;
[0022] S1032. k-1 frames of the optical flow feature map of size W × HThe optical flow feature map of ×2 is normalized to the range of [-1, 1], and for k-1 the weighted average value of the normalized optical flow feature map of each frame is calculated to obtain the optical flow fraction of the original video V .
[0023] In a preferred embodiment, in step S1032, the weighted average value of the normalized optical flow feature map is calculated by the following method:
[0024] For k-1 the positions greater than 0 in the normalized optical flow feature map of each frame are multiplied by a first coefficient α greater than 1, and for k- 1 the positions less than 0 in the normalized optical flow feature map of each frame are multiplied by a second coefficient β less than 1 to obtain k-1 the weighted average value of the normalized optical flow feature map of each frame.
[0025] In a preferred embodiment, in step S1032, the average value of the normalized optical flow feature map of each frame is directly calculated to obtain the optical flow fraction of the original video k-1 . V
[0026] In a preferred embodiment, in step S103, the positions of the empty graph are represented in the order from left to right and from top to bottom using coordinates I of k blocks with a size of W × H × C .
[0027] Another aspect of the present invention is to provide a video description generation system based on a grid graph, and the video description generation system is used to execute a video description generation method based on a grid graph provided by the present invention.
[0028] Another aspect of the present invention is to provide a computer storage medium, and the computer storage medium is used to store computer execution instructions, and the computer execution instructions are used to execute a video description generation method based on a grid graph provided by the present invention.
[0029] Compared with the prior art, the present invention has the following beneficial effects:
[0030] A video description generation method, system and storage medium based on a grid graph proposed by the present invention, by constructing an empty graph I , dividing the empty graph I into blocks with a size of W × H × C , and the extracted kThe first frame of the image is placed in an empty image from left to right and from top to bottom in sequence I in the k blocks of size W × H × C to be stitched into a second grid-like image of size W × H × C The second grid-like image of size W × H × C is adjusted to a third image (video image) of size T × T × C so that the size of the third image (video image) input into the LVLM model becomes 1× T × T × C The computing cost is reduced from the original calculation of k video images to the calculation of 1 / k video images, greatly reducing the computational complexity, computational resource consumption, and generation time of video description generation.
[0031] A method, system, and storage medium for video description generation based on a grid graph proposed by the present invention. Since the extracted k frames of the first image are placed in an empty image from left to right and from top to bottom in sequence I in the k blocks of size W × H × C to be stitched into a second grid-like image of size W × H × C even when the number of frames of the extracted first image is k relatively small, it can ensure to reduce the crossing time between frames and avoid incoherence of inter-frame information. Thus, when the second grid-like image of size W × H × C is adjusted to a third image of size T × T × C as the input of the LVLM model, while reducing the computational complexity, computational resource consumption, and generation time of video description generation, it can retain as much video information as possible and ensure the generation quality of video descriptions.
[0032] A method, system and storage medium for generating video descriptions based on grid graphs proposed by the present invention utilize the extracted k first image of each frame to calculate the optical flow fraction of the original video, and the original video without obvious motion is removed through the optical flow fraction of the original video, further improving the quality of the original video data. Description of the Drawings
[0033] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0034] Figure 1 is a flowchart of a method for generating video descriptions based on grid graphs according to the present invention. Detailed Embodiments
[0035] In order to make the above and other features and advantages of the present invention clearer, the present invention will be further described below with reference to the drawings. It should be understood that the specific embodiments given herein are for the purpose of explaining to those skilled in the art and are merely exemplary, not restrictive.
[0036] Combined with Figure 1 , according to an embodiment of the present invention, a method for generating video descriptions based on grid graphs is provided. Video description generation refers to the process in which a computer automatically generates a text description for a video.
[0037] A method for generating video descriptions based on grid graphs according to the present invention includes the following method steps:
[0038] Step S101: Obtain the original video V .
[0039] Step S102: Extract the first image of each frame at equal intervals from the obtained original video V ; where k should satisfy being able to be square-rooted. k The size of each first image is
[0040] × W × H × C ; where W represents the length of each first image, H represents the width of each first image, C represents the number of channels of the original video V . Generally, the original video VNumber of channels C= 3.
[0041] In one embodiment, from the acquired original video V Extract at equal intervals k = 9 frames or k = 16 frames of the first image. The number of frames of the extracted first image k = 9 frames or k = 16 frames all satisfy being square-rooted.
[0042] In this embodiment, exemplarily from the acquired original video V Extract at equal intervals k = 9 frames of the first image, as Figure 1 shown, the extracted k = 9 frames of the first image are respectively: the 1st frame of the first image, the 2nd frame of the first image, the 3rd frame of the first image, the 4th frame of the first image, the 5th frame of the first image, the 6th frame of the first image, the 7th frame of the first image, the 8th frame of the first image, and the 9th frame of the first image.
[0043] In another embodiment, in order to ensure reducing the crossing time between frames and avoiding incoherence of inter-frame information, it is preferred to extract from the acquired original video V Extract at equal intervals k = 25 frames of the first image. The number of frames of the extracted first image k = 25 frames satisfies being square-rooted.
[0044] Step S103, use the extracted k frames of the first image to calculate the optical flow fraction of the original video V .
[0045] Specifically, calculate the optical flow fraction of the original video through the following method V :
[0046] Step S1031, input the extracted k frames of the first image into a deep learning model for optical flow calculation to obtain k-1 frames of an optical flow feature map with a size of W × H × 2 G (( k-1 ) × W × H × 2 optical flow feature map G ).
[0047] Wherein, W represents the length of each frame of the first image (the length of each frame of the optical flow feature map G ), H represents the width of each frame of the first image (the width of each frame of the optical flow feature map GThe width), the optical flow feature map G The number of channels of is 2, respectively representing the optical flow feature map G In the horizontal and vertical directions.
[0048] The deep learning model of the present invention is a deep learning model for optical flow estimation, and those skilled in the art can reasonably select a specific deep learning model according to specific circumstances.
[0049] Step S1032, will k-1 The frame size is W × H ×2 optical flow feature map G Normalize to between [-1, 1], for k-1 The weighted average of the frame-normalized optical flow feature maps is calculated to obtain the optical flow score of the original video V .
[0050] Furthermore, the weighted average of the normalized optical flow feature maps is calculated by the following method:
[0051] For k-1 The positions greater than 0 in the frame-normalized optical flow feature map are multiplied by a first coefficient α greater than 1, and for k- 1 The positions less than 0 in the frame-normalized optical flow feature map are multiplied by a second coefficient β less than 1 to obtain k-1 The weighted average of the frame-normalized optical flow feature map.
[0052] Preferably, the first coefficient α is taken as α = 1.05; the second coefficient β is taken as β = 0.95.
[0053] Since the first images of the original video without obvious motion V Of k The frames are basically the same, lacking the temporal features of the video modality. The present invention uses optical flow to represent the motion information between frames of the original video V There are V Of k The first image of the frame has k -1 frame optical flow feature map G , the larger the optical flow value, the more the pixels at that place have motion information. By k -1 frame optical flow feature map G Calculate the optical flow score of the original video V , the larger the optical flow score, the more it shows that the original video V As a whole has certain motion features. The present invention uses the extracted k The first image of the frame calculates the optical flow score of the original video, and the original video without obvious motion is removed through the optical flow score of the original video, further improving the quality of the original video data.
[0054] The present invention calculates the weighted average of the normalized optical flow feature map, making the weight of the motion information of the original video V larger to more accurately evaluate the V motility of the original video.
[0055] In some other embodiments, when calculating the optical flow fraction of the original video V , the weighted average may not be calculated for the optical flow feature map after frame normalization, and the average value of the optical flow feature map after frame normalization is directly calculated to obtain the k-1 optical flow fraction of the original video. k-1 V According to the embodiments of the present invention, if the optical flow fraction is lower than a preset threshold
[0056] , the original video is discarded γ , and the original video is re-obtained V V。
[0057] If the optical flow fraction is higher than the preset threshold γ , an empty map with a size of W × H × C is constructed I , and the empty map I is divided into blocks with a size of W × H × C ; where k is the number of frames of the first image equally spacedly extracted from the obtained original video V .
[0058] As shown in Figure 1 , in this embodiment, V equally spacedly extracts k = 9 frames of the first image from the obtained original video, and thus constructs an empty map with a size of 3 W × 3 H × C , that is, the length of the empty map I , i.e., the empty map I is 3 W , and the width is 3 H .
[0059] The constructed empty map I is divided into 9 blocks with a size of W × H × C , and the coordinates are used to represent the I of the empty map k = 9 blocks with a size ofW × H × C Position of the block
[0060] In this embodiment, nine blocks with a size of W × H × C are respectively: block (1), block (2), block (3), block (4), block (5), block (6), block (7), block (8) and block (9), as Figure 1 shown
[0061] Step S104: Place the k first frame of the first image from left to right and from top to bottom in sequence into the empty image I in the k blocks with a size of W × H × C to obtain a second image (grid image) with a size of W × H × C
[0062] In this embodiment, place the k = 9 frames of the first image from left to right and from top to bottom in sequence into the empty image I in the k = 9 blocks with a size of W × H × C
[0063] That is, place the first image of the first frame in block (1), the first image of the second frame in block (2), the first image of the third frame in block (3), the first image of the fourth frame in block (4), the first image of the fifth frame in block (5), the first image of the sixth frame in block (6), the first image of the seventh frame in block (7), the first image of the eighth frame in block (8), and the first image of the ninth frame in block (9) to obtain a second image (grid image) with a size of 3 W × 3 H × C as Figure 1 shown
[0064] The present invention places the extracted k frames of the first image from left to right and from top to bottom in sequence into the empty image I in the k blocks with a size of W × H × C to splice into a size of W × H × C The second image (grid image) in a grid pattern, where the second image (grid image) has the same number of grids in length and width, and a single second image (grid image) contains all the information of the extracted k frames of the first image, and the second image (grid image) contains k the timing information of the frames of the first image in such a way that the number of grids in length and width is the same, so that even when the number of frames of the extracted first image k is small, it can ensure a reduction in the crossing time between frames and avoid discontinuous information between frames.
[0065] Step S105: Adjust the obtained second image with a size of W × H × C into a third image with a size of T × T × C ; where T represents the length and width of the third image.
[0066] In this embodiment, the obtained second image with a size of 3 W × 3 H × C is adjusted into a third image with a size of T × T × C ( T usually 336 / 448).
[0067] Step S106: Input the third image with a size of T × T × C and the generated text into the LVLM model to output the generated video description.
[0068] For example, in this embodiment, the generated text promptP input into the LVLM model is: Please generate the video description.
[0069] Input the third image with a size of T × T × C obtained in step S105 and the generated text promptP (for example: Please generate the video description) into the LVLM model to output the generated video description.
[0070] In the present invention, the extracted k frames of the first image are placed in order from left to right and from top to bottom in the empty image I of k number of W ×H × C in the block of, spliced into a size of W × H × C grid-shaped second image (grid map), the size of W × H × C the grid-shaped second image is adjusted to a third image (video image) with a size of T × T × C so that the size of the third image (video image) input into the LVLM model becomes 1× T × T × C , and all the information of the k frames of the first image and the k sequential information of the k frames of the first image are included in one third image (video image). The computing cost is reduced from the original calculation of k video images to the calculation of 1 /
[0071] In some embodiments, if the number of frames of the first image V equidistantly extracted from the obtained original video k cannot be square-rooted, for example k = 8, then k = 8 frames of the first image are directly sequentially spliced to obtain a size of kW × H × C or W × kH × C second image, and the second image with a size of kW × H × C or W × kH × C is adjusted to a third image (video image) with a size of T × T × C and input into the LVLM model. At this time, the computing complexity, computing resource consumption, and generation time of video description generation can also be reduced. However, since the second image is not a grid map, it may cause a long crossing time between frames and discontinuous inter-frame information.
[0072] It should be noted that the present invention is based on the obtained original video VThe number of frames of the first image extracted at medium intervals k The case where the square root cannot be taken (for example k = 8) is a fallback solution. The most preferred solution should be to obtain the original video V The number of frames of the first image extracted at medium intervals k should satisfy being able to take the square root (for example k = 9), so that the extracted k frames of the first image are placed in order from left to right and from top to bottom in an empty image I of k sized W × H × C blocks, and stitched into a second image (grid image) in a grid shape of size W × H × C so that the second image in a grid shape of size W × H × C is adjusted to a third image of size T × T × C as the input of the LVLM model. While reducing the computational complexity, computational resource consumption, and generation time of video description generation, it can retain the information of the video as much as possible and ensure the generation quality of the video description.
[0073] According to an embodiment of the present invention, a video description generation system based on a grid image is provided for implementing a video description generation method based on a grid image provided by the present invention.
[0074] According to an embodiment of the present invention, a computer storage medium is provided for storing computer execution instructions, and the computer execution instructions are used to execute a video description generation method based on a grid image provided by the present invention.
[0075] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A video description generation method based on a grid graph, characterized in that: The video description generation method comprises the following steps: S101. Obtaining original video V ; S102, from the original video obtained V Medium pitch extraction k Frame first image; in, k It should satisfy the square root; The size of the first image in each frame is W × H × C ;in, W Indicates the length of the first image in each frame, H Indicates the width of the first image in each frame, C Represents the original video V The number of channels; S103, using the extracted k Frame first image, calculate the original video V The optical flow score of If the optical flow score is lower than the preset threshold, the original video is discarded V , re-obtain the original video V ; If the optical flow score is higher than the preset threshold, a size of W × H × C Empty map I , the empty map I Divide into The size is W × H × C of blocks; among them, k The original video obtained from V The number of frames of the first image extracted with medium spacing; S104, will k The first image of the frame is placed in the empty image in order from left to right and from top to bottom. I of k The size is W × H × C In the block of W × H × C a second image of S105, the size is W × H × C The second image is resized to T × T × C A third image of ; wherein, T represents the length and width of the third image; S106, the size is T × T × C The third image and the generated class text are input into the LVLM model together, and the generated video description is output.
2. The video description generation method according to claim 1, characterized in that: In S102, the original video is obtained V Medium pitch extraction k =9 frames or k =16 frames of the first image.
3. The video description generation method according to claim 1, characterized in that: In S102, the original video is obtained V Medium pitch extraction k =25 frames of the first image.
4. The video description generation method according to claim 1, characterized in that: In step S103, the original video is calculated by the following method V The optical flow score is: S1031, extract k The first image of the frame is input into the deep learning model for optical flow calculation, and we get k-1 The frame size is W × H ×2 optical flow feature map; in, W Indicates the length of the first image in each frame, H Indicates the width of the first image of each frame. The number of channels of the optical flow feature map is 2, which respectively represents the horizontal and vertical directions of the optical flow feature map; S1032, k-1 The frame size is W × H ×2 optical flow feature map is normalized to [-1,1], k-1 The weighted average of the optical flow feature map after frame normalization is calculated to obtain the original video V The optical flow score.
5. The video description generation method according to claim 4, characterized in that: In step S1032, a weighted average value is calculated for the normalized optical flow feature map by the following method: right k-1 The positions greater than 0 in the optical flow feature map after frame normalization are multiplied by a first coefficient α greater than 1. k-1 The positions less than 0 in the optical flow feature map after frame normalization are multiplied by a second coefficient β less than 1 to obtain k-1 The weighted average of the normalized optical flow feature maps of the frames.
6. The video description generation method according to claim 4, characterized in that: In step S1032, directly calculate k- 1 The average value of the optical flow feature map after frame normalization is used to obtain the original video V The optical flow score.
7. The video description generation method according to claim 1, characterized in that: In step S103, the coordinates are used to represent the empty image from left to right and from top to bottom. I of k The size is W × H × C The location of the block.
8. A video description generation system based on a grid graph, characterized in that: The video description generation system is used to execute the video description generation method according to any one of claims 1 to 7.
9. A computer storage medium, characterized in that The computer storage medium is used to store computer-executable instructions, and the computer-executable instructions are used to execute the video description generation method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Video description generation method and device based on deep learning model
CN117292293A
Video description information generation method, video processing method, and corresponding devices
WO2020199904A1