Video generation method, device and electronic equipment
By detecting and identifying key frames and customizing the cropping process based on frame dimensions and height, the method addresses the misalignment in existing AI-generated videos, resulting in improved video presentation quality.
Patent Information
- Application Number
- CN202310246423.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-03-10
AI Technical Summary
In the prior art, when generating videos, the determination of the cropping box depends on significance detection and the center of the cropping box, resulting in poor video generation effect, especially when processing text information and object detection in the image.
By object detection of image frames, key image frames are identified, and cropping frames are generated based on the size and height information of key image frames, the image frames are cropped, and video is generated.
It improves the matching degree between image frames and object detection results, enhances the video presentation effect, reduces or avoids text information, and improves the video presentation quality.
Smart Images

Figure CN116258994B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, specifically to technical fields such as computer vision, augmented reality, virtual reality, deep learning, etc., and can be applied to scenarios such as the metaverse and artificial intelligence generated content (AIGC). In particular, it relates to a video generation method, device, and electronic device. Background Art
[0002] With the development of artificial intelligence technologies, in some scenarios, videos are also generated in an artificial intelligence manner. Currently, the main technical solution for generating videos using artificial intelligence is to perform saliency detection on each frame in the video, calculate the peak of the saliency heat map in the overall image, use it as the center of the cropping frame, and then calculate the target cropping frame based on the center of the cropping frame and the preset cropping frame size for cropping, and generate a video based on the cropped images. Summary of the Invention
[0003] The present disclosure provides a video generation method, device, and electronic device.
[0004] According to one aspect of the present disclosure, a video generation method is provided, including:
[0005] Performing object detection on multiple image frames in a first image set respectively to obtain the object detection results of the multiple image frames in the first image set, and identifying key image frames in the first image set based on the object detection results of the multiple image frames in the first image set;
[0006] Generating a first cropping frame according to the image size and height information of the key image frames, and performing a first cropping on the multiple image frames in the first image set based on the first cropping frame and the object detection results of the multiple image frames in the first image set to obtain a second image set, where the height information is the height information of the multiple image frames in the first image set;
[0007] Generating a video based on the multiple image frames in the second image set.
[0008] According to another aspect of the present disclosure, a video generation device is provided, including:
[0009] A detection module, configured to perform object detection on multiple image frames in a first image set respectively to obtain the object detection results of the multiple image frames in the first image set, and identify key image frames in the first image set based on the object detection results of the multiple image frames in the first image set;
[0010] The first cropping module is configured to generate a first cropping frame according to the image size and height information of the key image frame, and perform first cropping on multiple image frames in the first image set based on the first cropping frame and the object detection results of the multiple image frames in the first image set, to obtain a second image set, where the height information is the height information of the multiple image frames in the first image set;
[0011] The generation module is configured to generate a video based on the multiple image frames in the second image set.
[0012] According to another aspect of the present disclosure, there is provided an electronic device, including:
[0013] At least one processor; and
[0014] A memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the video generation method provided by the present disclosure.
[0016] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the video generation method provided by the present disclosure.
[0017] According to another aspect of the present disclosure, there is provided a computer program product including a computer program, where the computer program implements the video generation method provided by the present disclosure when executed by a processor.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0020] Figure 1 is a flowchart of a video generation method provided by the present disclosure;
[0021] Figure 2 is a schematic diagram of cropping of a video frame provided by the present disclosure;
[0022] Figure 3 is a schematic diagram of a video generation method provided by the present disclosure;
[0023] Figures 4a to 4f is a structural diagram of a video generation device provided by the present disclosure;
[0024] Figure 5 It is a block diagram of an electronic device for implementing an embodiment of the present disclosure. Specific embodiments
[0025] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] Please refer to Figure 1 , Figure 1 is a flowchart of a video generation method provided by the present disclosure. As Figure 1 shown, it includes the following steps:
[0027] Step S101: Perform object detection on multiple image frames in the first image set respectively to obtain the object detection results of the multiple image frames in the first image set, and identify key image frames in the first image set based on the object detection results of the multiple image frames in the first image set.
[0028] The above-mentioned first image set may be an image set obtained by frame extraction from an input image set, an image set obtained by frame extraction from a certain video, or the above-mentioned first image set is the input image set, or the above-mentioned first image set is an image set that has been cropped or not cropped. In addition, the image frames in the above-mentioned first image set may be image frames with the same image size.
[0029] The above-mentioned performing object detection on multiple image frames in the first image set respectively may be performing object detection on each image frame in the first image set respectively, or performing object detection on some image frames in the first image set respectively, such as sampling the above-mentioned first image set and performing object detection on the sampled image frames.
[0030] The above-mentioned identifying key image frames in the first image set based on the object detection results of the multiple image frames in the first image set may be comparing the object detection results of the multiple image frames respectively to generate key image frames with the most prominent objects among these multiple image frames.
[0031] Step S102: Generate a first cropping frame according to the image size and height information of the key image frame. Based on the first cropping frame and the object detection results of multiple image frames in the first image set, perform a first cropping on the multiple image frames in the first image set to obtain a second image set. The height information is the height information of the multiple image frames in the first image set.
[0032] The above height information can be the unified height information in the above first image set. For example, the height information of all image frames in the first image set is the same. Or, the above height information can be the lowest height information in the above first image set. For example, there are image frames with inconsistent heights in the first image set.
[0033] The above generating the first cropping frame according to the image size and height information of the key image frame can be based on the image size and height information of the key image frame, generating the first cropping frame with the center of the picture of the key image frame, and the vertical axis coordinate of the first cropping frame is referenced by the vertical axis coordinate corresponding to the above height information.
[0034] The above performing a first cropping on the multiple image frames in the first image set based on the first cropping frame and the object detection results of the multiple image frames in the first image set to obtain a second image set can be comparing the object detection result of each image with the above first cropping frame, performing a first cropping on the multiple image frames in the first image set based on the comparison result, and for some image frames in the first image set, performing processing such as updating and displacement on the above first cropping frame before cropping.
[0035] Step S103: Generate a video based on the multiple image frames in the second image set.
[0036] The above generating a video based on the multiple image frames in the second image set can be synthesizing the multiple image frames in the second image set to generate a video.
[0037] In the present disclosure, since the multiple image frames in the first image set are cropped based on the first cropping frame and the object detection results of the multiple image frames in the first image set, it can make the cropped image frames easier to match the object detection results, thereby improving the cropping effect of the image frames, and further making the display effect of the finally generated video better.
[0038] In the present disclosure, the above method can be applied to an electronic device, that is, all steps of the above method are executed by the electronic device. The electronic device can be an electronic device such as a server, a computer, a mobile phone, or a tablet computer.
[0039] In one embodiment, the method further includes:
[0040] Identify the text position information of multiple image frames in the third image set;
[0041] Generate a second cropping box based on the text position information, and perform a second cropping on multiple image frames in the third image set according to the second cropping box to obtain the first image set.
[0042] Among them, the above-mentioned third image set can be a set of image frames in a certain video, or an image set sampled from a certain video. For example: for an input video with N frames, frame by frame extraction is performed, such as extracting k1 frames of images at a uniform interval of f frames among all frames.
[0043] The above-mentioned text position information can be the position information of the text box representing the image frame, where the text box can also be understood as the subtitle box.
[0044] The above-mentioned generating the second cropping box based on the text position information can be to generate a cropping box for removing the text information in the image frame, that is, the image frame after being cropped by the above-mentioned second cropping box does not include text information, such as subtitle information.
[0045] In this embodiment, since the multiple image frames in the third image set are second cropped according to the second cropping box, this can reduce or avoid the first image set including the text information of the original image, and further improve the display effect of the video.
[0046] In one embodiment, the text position information includes: a first text box coordinate information set of multiple image frames included in the third image set, and the text box coordinate information in the first text box coordinate information set includes horizontal axis coordinate information and vertical axis coordinate information;
[0047] The generating the second cropping box based on the text position information includes:
[0048] Select a first vertical axis coordinate information in the first text box coordinate information set, and the distance from the first vertical axis coordinate information to the top of the image frame is less than the distance from other vertical axis coordinate information to the top of the image frame, and the other vertical axis coordinate information is all or part of the vertical axis coordinate information in the first text box coordinate information set except the first vertical axis coordinate information;
[0049] Generate a second cropping box based on the first vertical axis coordinate information.
[0050] Among them, the above-mentioned first text box coordinate information can be the coordinate information of the text box, such as {xin, ymin, xmax, ymax}, where xin, ymin, xmax, and ymax respectively represent the minimum value of the horizontal axis coordinate, the minimum value of the vertical axis coordinate, the maximum value of the horizontal axis coordinate, and the maximum value of the vertical axis coordinate of the text box.
[0051] The above-mentioned generation of the second cropping frame based on the first vertical axis coordinate information may be to generate a second cropping frame whose size in the vertical axis direction matches the size represented by the above-mentioned first vertical axis coordinate information.
[0052] In this embodiment, by generating the second cropping frame based on the first vertical axis coordinate information, the text information of the image frames in the first image set can be cropped by the second cropping frame, thereby reducing or avoiding the text information in the previous video in the finally generated video.
[0053] It should be noted that in the embodiments of the present application, it is not limited that the above-mentioned text position information includes: the first text box coordinate information set of multiple image frames included in the first image set. For example: the image frames in the first image set can be divided into multiple small regions in advance, and the text position information can also be represented by the identification information of these small regions.
[0054] In one embodiment, the sizes of the multiple image frames included in the third image set are the same, the width of the second cropping frame is the same as the width of the multiple image frames included in the third image set, and the first cropping frame is from the top of the multiple image frames included in the third image set to the coordinate position represented by the first vertical axis coordinate information in the height direction.
[0055] Taking the coordinate origin of the image frame as the upper left corner of the image frame as an example, each extracted image frame is recognized using an optical character recognition (OCR) text recognition algorithm based on deep learning to obtain the set of all detected text boxes (which can be represented as D t ), where the text box form is {xin, ymin, xmax, ymax} ∈ D t , and the set of text boxes with ymin > (0.5 * height) is taken (which can be represented as D ts ), where height represents the height in the image frame of the first image set. Calculate the minimum ymin in D ts of all frames, and generate a cropping frame {0, 0, W, ymin ts}, where W is the width of the image frame, and all the extracted frames are cropped to obtain the cropped frame set F, that is, the above-mentioned second image set. For example: as ts shown, it includes the coordinate origin 201, the text information 202, and generates a second cropping frame 203 of {0, 0, W, ymin Figure 2 ts}.}
[0056] It should be noted that in the embodiments of the present application, mainly taking the origin of the coordinate of the image frame as the upper left corner of the image frame as an example. In the embodiments of the present application, the origin of the coordinate of the image frame can also be other position points, such as the upper right corner of the image frame, etc., which is not limited herein.
[0057] In this embodiment, since the second cropping frame in the height direction is from the top of the multiple image frames included in the third image set to the coordinate position represented by the first vertical axis coordinate information, this further reduces or avoids the text information in the previous video existing in the finally generated video.
[0058] It should be noted that in the embodiments of the present application, it is not limited that the second cropping frame in the height direction is from the top of the multiple image frames included in the first image set to the coordinate position represented by the first vertical axis coordinate information. For example: for some scenarios, the second cropping frame in the height direction can be an offset from the top of the multiple image frames included in the third image set to the coordinate position represented by the first vertical axis coordinate information, and this offset can be set according to empirical values, which can also reduce or avoid the text information in the previous video existing in the finally generated video.
[0059] In one embodiment, the selecting the first vertical axis coordinate information from the first text box coordinate information set includes:
[0060] Performing a first filtering on the first text box coordinate information set to obtain a second text box coordinate information set, where the first filtering is used to filter out the text box coordinate information in the first text box coordinate information set whose distance from the top of the image frame is less than a first preset distance;
[0061] Selecting the first vertical axis coordinate information from the second text box coordinate information set, and the other vertical axis coordinate information is all or part of the vertical axis coordinate information in the second text box coordinate information set except the first vertical axis coordinate information.
[0062] Wherein, the above first preset distance is a preset empirical value, for example: 0.5 times the height of the image frame.
[0063] In this embodiment, since the text box coordinate information in the first text box coordinate information set whose distance from the top of the image frame is less than the first preset distance is filtered out, this can avoid the second cropping frame from cropping off too many regions in the image frame, so as to improve the display effect of the finally generated video.
[0064] In one embodiment, the performing target detection on the multiple image frames in the first image set respectively to obtain the target detection results of the multiple image frames in the first image set, and identifying the key image frames in the first image set based on the target detection results of the multiple image frames in the first image set includes:
[0065] Perform object detection and object trajectory tracking on multiple image frames in the first image set respectively, to obtain the object detection results and object trajectory tracking results of the multiple image frames in the first image set, where the object trajectory tracking results are used to represent the trajectory information of tracking the objects detected by the object detection;
[0066] Based on the object detection results and object trajectory tracking results of the multiple image frames in the first image set, identify the key image frames in the first image set.
[0067] The above-mentioned performing object detection and object trajectory tracking on multiple image frames in the first image set respectively can be to sample the first image set and perform object detection and object trajectory tracking on the sampled multiple image frames.
[0068] The above-mentioned object detection can be to perform object detection by using a deep learning-based object detection algorithm (such as the Yolo algorithm) to obtain a set of detection boxes of the objects in each image frame.
[0069] In some embodiments, each object detection result can be represented as an object detection box {xin, ymin, xmax, ymax, cls, s} ∈ D o , where cls is the class label, S is the confidence score, and D o represents the set of object detection boxes.
[0070] It should be noted that in the embodiments of the present application, the object detection result is not limited to the above-mentioned object detection box. For example: the image frame is pre-divided into multiple small regions, so that the object detection result can also be represented by the identification of these small regions.
[0071] The above-mentioned object trajectory tracking can be to perform trajectory tracking based on a trajectory tracking algorithm, such as performing trajectory tracking based on the Kernel Correlation Filter (KCF) algorithm.
[0072] In some embodiments, an image frame can include multiple object detection results. For example: an image frame includes multiple human images, so that each human image has a corresponding object detection result, and each object detection result has a corresponding object trajectory tracking result.
[0073] The above-mentioned identifying the key image frames in the first image set based on the object detection results and object trajectory tracking results of the multiple image frames in the first image set can be to generate the image frame with the most prominent or significant object as the key image frame based on the object detection results and object trajectory tracking results of each image frame.
[0074] It should be noted that when the above-mentioned second image set includes images of multiple scenarios, key image frame recognition can be performed separately for different scenarios. For example, each scenario corresponds to an image set.
[0075] In this embodiment, since the key image frames are recognized based on the object detection results and the object trajectory tracking results, the accuracy of the key image frames can be improved.
[0076] It should be noted that in the embodiments of the present application, it is not limited to recognizing key image frames based on object detection results and object trajectory tracking results. For example, in some embodiments, key image frames can also be recognized only based on object detection results. For example, the image frame with the highest confidence in the object detection results is used as the key image frame.
[0077] In one embodiment, performing object detection and object trajectory tracking on multiple image frames in the first image set respectively to obtain the object detection results and object trajectory tracking results of the multiple image frames in the first image set includes:
[0078] Performing object detection and object trajectory tracking on the image frames in the image subset of the first image set respectively to obtain the object detection results and object trajectory tracking results of the image frames in the image subset, where the image subset includes multiple image frames sampled at intervals in the first image set;
[0079] Among them, the object trajectory tracking result of the first image frame in the image subset includes: a first trajectory detection box, the frame number of the first image frame, and a first trajectory array. The first trajectory detection box is the object detection result of the first image frame, and the first trajectory array includes the object detection result of the first image frame. The first image frame is the first image frame in the image subset;
[0080] Among them, the object trajectory tracking result of the second image frame in the image subset includes: a second trajectory detection box, the frame number of the second image frame, and a second trajectory array. The second trajectory detection box is the object detection result of the second image frame, and the second trajectory array includes the object detection result of the first image frame and the object detection result of the second image frame, or the second trajectory array is the object detection result of the second image frame. The second image frame is the image frame in the image subset other than the first image frame.
[0081] The object trajectory tracking result of the above first image frame can be the set D of detection boxes detected in the first image frame o Construct a set of object trajectory tracking results (which can be expressed as Track), and this set is specifically expressed as T oFor each target detection box d = {xin, ymin, xmax, ymax, cls, s}, its corresponding target track append result (Track t) = {d t ,_newest_fidx,_histroy}, where d t It represents the first trajectory detection frame, _newest_fidx represents the frame number of the first image frame, specifically the frame number that makes the target trajectory tracking result updated the most recently, and _histroy represents the first trajectory array, specifically an array composed of target detection frames that have been updated for the target trajectory tracking result.
[0082] It should be noted that the target trajectory tracking result of each image frame in this embodiment can be expressed as {d t ,_newest_fidx,_histroy}, for example: for the target trajectory tracking result of the second image frame above, d t ,_newest_fidx, _histroy respectively represent the second trajectory detection frame, the frame number of the second image frame and the second trajectory array.
[0083] In the case where the second trajectory array includes the target detection result of the first image frame and the target detection result of the second image frame, if there are other image frames between the second image frame and the first image frame, the second trajectory array includes, in addition to the target detection result of the first image frame and the target detection result of the second image frame, the target detection results of the other image frames in the interval, that is, the second trajectory array of each image frame includes the target detection results of the image frame and the image frame before the image frame.
[0084] In addition, the target trajectory tracking result of the second image frame is obtained by adding the target trajectory additional result of the first image frame. o Each Track t (Track t represents any target track tracking result) in (target track tracking result set of the first image frame) is located at the position of the second image frame and the d attribute of Track t (i.e. target detection result) is updated, and the new d is added to the _histroy array, the _newest_fidx information is updated, and the corresponding d is updated t .
[0085] In this embodiment, the accuracy of the target addition result can be improved through the trajectory detection frame, frame sequence number and trajectory array of the image frame.
[0086] It should be noted that, in some embodiments, the target trajectory appending result of each image frame is not limited to include the trajectory detection box, the frame number and the trajectory array. For example, in some embodiments, the image frame is divided into multiple small areas, so that the target trajectory appending result includes a trajectory detection identifier, a frame number and a trajectory identifier array, wherein these identifiers are used for the corresponding small areas for trajectory appending.
[0087] In one embodiment, the method further comprises:
[0088] Calculating the intersection over union (IOU) of the target detection result of the second image frame and the trajectory detection frame in the special target trajectory tracking result, wherein the special target trajectory tracking result includes: a target trajectory tracking result whose frame number difference with the frame number of the second image frame is within a preset range;
[0089] Wherein, when the IOU between the target detection result of the second image frame and the matching target trajectory tracking result is greater than a preset threshold, the second trajectory array includes the target detection result of the first image frame and the target detection result of the second image frame, and when the IOU between the target detection result of the second image frame and the matching target trajectory tracking result is less than or equal to the preset threshold, the second trajectory array is the target detection result of the second image frame;
[0090] The matching target trajectory tracking result is a target trajectory tracking result having the highest matching degree with the target detection result of the second image frame among the special target trajectory tracking results.
[0091] Among them, the above-mentioned preset range can be a preset empirical value, for example: the above-mentioned special target trajectory tracking result is the target trajectory tracking result of _newest_fidx attribute>j-3k, wherein j is the frame number of the current second image frame, and k is the sampling interval of the image subset.
[0092] For example: for the target detection result set D in the second image frame o For each target detection result in (which can be expressed as d), calculate t∈T with all _newest_fidx attributes>j-3k o Medium t IOU of d and d t The bipartite graph between d and d is obtained by using the Hungarian matching algorithm. tmax (i.e., the trajectory detection box in the above matching target trajectory tracking result). tmax If the IOU of _histroy is greater than the preset threshold ε, the match is successful, and the last detection box of _histroy in tmax (i.e., the above matching target trajectory tracking result) is replaced with d, and dt Replace it with d, and at the same time update the _newest_fidx of tmax to j. If the IOU with d tmax is less than the preset threshold ε, then the current d fails to match, and a new Track is created for d and added to T o in the same way as the target trajectory tracking result of the first image frame, that is, the trajectory detection box is the target detection result of this image frame, and the trajectory array includes the target detection result of this image frame.
[0093] In this embodiment, by calculating the IOU, the target trajectory tracking result of the second image frame can be made more accurate, and since only the IOU between the target detection result and the trajectory detection box in the special target trajectory tracking result needs to be calculated, this can save computing resources.
[0094] It should be noted that in some embodiments, the IOU may not be calculated. For example, it is directly determined that the second trajectory array includes the target detection result of the first image frame and the target detection result of the second image frame, or the second trajectory array is the target detection result of the second image frame.
[0095] In some embodiments, for the multiple target trajectory tracking results of the above-mentioned first image frame, the number of frames tracked by each target trajectory tracking result may be different. For example, check the _newest_fidx of all Tracks in T o If j - newest > _max_age (_max_age is defaulted to 3k), then take out this Track from T o and add it to the set T d to represent the Tracks that have been tracked. If the current image frame f j is the last frame in the current scene (such as a certain second image set), then add all Tracks in T o to T d (the set of Tracks that have been tracked).
[0096] In one embodiment, the identifying the key image frames in the first image set based on the target detection results and target trajectory tracking results of multiple image frames in the first image set includes:
[0097] Calculate the first scores of the target trajectory tracking results of multiple image frames in the image subset respectively, where the first score of each target trajectory tracking result is calculated based on the coordinates in the trajectory array of the target trajectory tracking result;
[0098] For each image frame in the image subset, calculate the second score of the image frame based on the target detection result of the image frame and the first score;
[0099] Select the image frame with the second highest score in the image subset as the key image frame in the first image set.
[0100] Among them, the above first score can be positively correlated with the size of the target detection box in the trajectory array of the target trajectory tracking result.
[0101] For example: For each target trajectory tracking result Track in T d calculate the score
[0102] where score() is the preset weight corresponding to each class label. For example: the default weight for the face target is 25, and the rest is 1. d[ymax], [ymin], d[xmax], and d[xmin] respectively represent the values of ymax, ymin, xmax, and xmin in the target detection box d.
[0103] It should be noted that in the present disclosure, there is no limitation on calculating the above first score through the above formula. For example: calculate the above first score through the following formula:
[0104] Or
[0105]
[0106] After obtaining the first score s of each target trajectory tracking result t , Track can be sorted from large to small according to s t to obtain T d ’ (the set of Track that has been tracked and sorted).
[0107] Based on the target detection result of the image frame and the first score, the second score of the image frame can be calculated as follows:
[0108] The second score is positively correlated with the size of the detection box of the target detection result and the corresponding first score.
[0109] For example: Traverse each image frame f in the image subset j , calculate where s d is the first score of the target trajectory tracking result Track to which the target detection result d belongs. The relationship between the target detection result d and Track can be found by traversing the histroy attribute of the target trajectory tracking result Track in the target trajectory tracking set T d . Select the image frame f with the highest s fj keyAs the key image frame.
[0110] It should be noted that in this embodiment, the calculation of the above second score is not limited to the above formula. For example, the above second score can also be calculated by the following formula:
[0111] Or
[0112]
[0113] In this embodiment, the recognition accuracy of the key image frame is improved by the above second score.
[0114] It should be noted that in the embodiments of the present application, the recognition of the above key image frame is not limited to the above first score and second score. For example, in some embodiments, the image frame to which the target detection result with the most included in the trajectory array in the target trajectory tracking result belongs can be used as the key frame.
[0115] In one embodiment, generating a first cropping frame according to the image size and height information of the key image frame, and performing a first cropping on multiple image frames in the first image set based on the first cropping frame and the target detection results of the multiple image frames in the first image set to obtain a second image set, includes:
[0116] Calculating the aspect ratio of the key image frame based on the image size of the key image frame;
[0117] Generating a first cropping frame based on the aspect ratio and the height information;
[0118] Comparing the target detection results of multiple image frames in the first image set with the first cropping frame respectively to obtain comparison results;
[0119] For multiple image frames in the first image set, performing a first cropping respectively based on the comparison results and the first cropping frame to obtain a second image set.
[0120] Among them, the above aspect ratio can be expressed as r = W / H, where W and H represent the width and height of the key image frame respectively.
[0121] In some embodiments, generating the first cropping frame based on the aspect ratio and the height information may be generating a first cropping frame with the aspect ratio of the key image aspect ratio, and the height of the second cropping frame is equal to the height represented by the height information.
[0122] In some embodiments, generating the first cropping frame based on the aspect ratio and the height information may be generating the above first cropping frame through the following formula:
[0123] Crop = {(W - r × ymin ts ) / 2, 0, r × ymin ts , ymin ts )}
[0124] where W is the width of the above-mentioned key image frame, ymin ts represents the above-mentioned height information, and r is the above-mentioned aspect ratio.
[0125] The above comparison result can represent the positional relationship between each of the above object detection results and the above first cropping frame, for example: contained, not contained, partially contained, etc.
[0126] And the first cropping based on the comparison result and the first cropping frame can be to directly crop according to the first cropping frame, or to update, offset or scale the first cropping frame and then perform cropping.
[0127] In this embodiment, since the first cropping is performed based on the comparison result and the first cropping frame to obtain a second image set, the cropping effect of the image frame can be improved, such as avoiding the target being truncated.
[0128] And since the first cropping frame is generated based on the aspect ratio and height information of the key image frame, it is not necessary to dynamically generate cropping frames frame by frame, thereby effectively improving the algorithm processing speed and reducing the content production cost.
[0129] It should be noted that in the embodiments of the present application, there is no limitation on obtaining the second image set in the above manner. For example: directly cropping according to the above first cropping frame.
[0130] In one embodiment, the first cropping is respectively performed on multiple image frames in the first image set based on the comparison result and the first cropping frame, including:
[0131] For multiple image frames in the first image set, the first cropping is respectively performed based on the comparison result, the first cropping frame and the target data group, and the target data group is a multi-dimensional array generated based on the coordinates in the key image frame and the image size of the key image frame;
[0132] where the first cropping includes at least one of the following:
[0133] Cropping according to the first cropping frame;
[0134] Updating the first cropping frame and performing cropping based on the updated cropping frame, and the update is based on the target data group;
[0135] Displacing the first cropping frame and performing cropping based on the displaced cropping frame.
[0136] In some embodiments, the above-mentioned target data group may be a multi-dimensional array generated based on the coordinates of the target detection result with the highest cumulative first score in the corresponding target trajectory tracking result in the key image frame and the image size of the key image frame. For example: According to the target detection result set D key in the key image frame f o of all target detection results d belonging to T d in Track s t , sort D o from high to low, and traverse D o in order, maintaining a set of data valid_xmin = d0[xmin], valid_xmax = d0[xmax], invalid_xmin = -1, invalid_xmax = W + 1, where d0[xmin] and d0[xmax] respectively represent the xmin and xmax in the first target detection box in the sorted target detection result set D o .
[0137] In some embodiments, the above-mentioned target data group may be a multi-dimensional array generated based on the coordinates of the target detection result corresponding to the target trajectory tracking result with the highest first score in the key image frame and the image size of the key image frame.
[0138] Among them, the above-mentioned cropping according to the first cropping box, updating the first cropping box, and displacing the first cropping box may be to select direct cropping, updating, and displacement according to the comparison result between the target detection result and the first cropping box. In addition, the above-mentioned update may be to scale the first cropping box. For example: For each target detection box (Dets) in the key frame, determine whether it conflicts with the current first cropping box (which can be represented as Crop) in turn according to the priority of the corresponding target trajectory tracking result. If there is a conflict, move Crop to frame or exclude the current detection box, or scale Crop to exclude the current detection box.
[0139] An example is as follows:
[0140] The new width (which can be represented as newW) is calculated as Crop[2] - Crop[0], and the new height (which can be represented as newH) is calculated as Crop[3] - Crop[1], where Crop[0], Crop[1], Crop[2], and Crop[3] respectively represent the first, second, third, and fourth elements in the first cropping box. For each target detection result d = {xmin, ymin, xmax, ymax}, it may include at least one of the following:
[0141] If xmax < Crop[2] and xmin > Crop[0], it means that d is completely within the first cropping frame Crop and no additional processing is required. Update valid_xmin = max(valid_xmin, xmin) and valid_xmax = min(valid_xmax, xmax) in the above target data group;
[0142] Otherwise, if xmin > Crop[2], it means that d is completely on the right side outside Crop and no additional processing is required;
[0143] Otherwise, if xmax < Crop[0], it means that d is completely on the left side outside Crop and no additional processing is required;
[0144] Otherwise, if xmax > Crop[2] and xmax - valid_xmin < newW and xmax < invalid_xmax, it means that d is truncated on the right side and Crop can be shifted to the right to retain the integrity. Therefore, update Crop[2] = xmax, Crop[0] = Crop[2] - newH * r, update valid_xmin = max(valid_xmin, xmin), and update valid_xmax = min(valid_xmax, xmax);
[0145] Otherwise, if xmin < Crop[0] and valid_xmax - xmin < newW and xmax < invalid_xmax, it means that d is truncated on the left side and Crop can be shifted to the left to retain the integrity. Therefore, update Crop[0] = xmin, Crop[2] = Crop[0] + newH * r, update valid_xmin = max(valid_xmin, xmin), and update valid_xmax = min(valid_xmax, xmax);
[0146] Otherwise, if xmax > Crop[2] and invalid_xmax - newW > invalid_xmin, it means that d is truncated on the right side and d can be excluded by shifting Crop to the left. Therefore, update invalid_xmax = max(xmax, valid_xmax), Crop[2] = invalid_xmax, Crop[0] = Crop[2] - newW; if xmax > Crop[2] and invalid_xmax - newW <= invalid_xmin, it means that d is truncated on the right side and d cannot be excluded by shifting Crop to the left, and only the way of scaling Crop can be used. Therefore, update Crop[0] = invalid_xmin, Crop[2] = invalid_xmax, and at the same time update Crop[3] = newW / r; if we want to prevent the area of the finally generated Crop from being too small, a limiting condition can be added. If (invalid_xmax - invalid_xmax) > δ × r × ymin ts , then update Crop, otherwise skip this d, where δ is a preset value;
[0147] Otherwise, if xmin < Crop[0] and invalid_xmin + new < invalid_xmax, it means that d is truncated on the left side and d can be excluded by shifting Crop to the right. Therefore, update invalid_xmin = min(xmin, valid_xmin), Crop[0] = invalid_xmin, Crop[2] = Crop[0] - newW; if xmin < Crop[0] and invalid_xmin + new >= invalid_xmax, it means that d is truncated on the left side and d cannot be excluded by shifting Crop to the right, and only the way of scaling Crop can be used. Therefore, update Crop[0] = invalid_xmin, Crop[2] = invalid_xmax, and at the same time update Crop[3] = newW / r; if we want to prevent the area of the finally generated Crop from being too small, a limiting condition can be added. If (invalid_xmax - invalid_xmax) > δ × r × ymin ts , then update Crop, otherwise skip this d.
[0148] In this embodiment, by performing the above-mentioned cropping according to the first cropping frame, updating the first cropping frame, displacing the first cropping frame, and scaling the first cropping frame, the target detection result can be flexibly cropped, thereby further reducing or avoiding the target being truncated.
[0149] In one embodiment, the first image set includes N scenes, and for each scene, there are a corresponding key image frame, a first cropping frame, and a second image set, where N is an integer greater than 1;
[0150] Generating a video based on multiple image frames in the second image set includes:
[0151] Performing size correction on the second image sets of the N scenes, and synthesizing the second image sets of the N scenes after size correction to generate a video;
[0152] Among them, the size correction includes:
[0153] Performing interpolation processing on the image frames to be corrected that are cropped based on the scaled cropping frame, so that the image size of the image frames to be corrected matches the image sizes of other image frames, where the other image frames are the image frames in the second image sets of the N scenes except the image frames to be corrected.
[0154] In this embodiment, scenes can be performed on image frames. For example: performing scene detection on the first image set, and then performing subsequent processing on the image sets of each scene respectively to obtain the key image frames, the first cropping frames, and the second image sets of each scene.
[0155] For example: for each frame f starting from the second frame in the second image set i , calculate diff i = sum(Gray(f i ) - Gray(f i-1 )) / W × H × 255 × α, where Gray is to convert the frame image in RGB format into a grayscale image, sum is to sum all the pixels of the image, f i-1 is the previous frame of f i , and α is a constant coefficient (default 100 or 900). Then calculate the second-order diff2 i = diff i - diff i-1 , and then update diff i = min(diff2 i , diff i ). If diff i is greater than the threshold (default 10), then this frame is considered the first frame of a new scene, otherwise it is still the current scene.
[0156] Among them, the above-mentioned interpolation processing of the image frames to be corrected that are cropped based on the scaled cropping frame can be to perform bicubic interpolation on the image frames to be corrected, so that the image size of the image frames to be corrected matches the image sizes of other image frames.
[0157] For example: newH in the first cropping frame Crop of a certain scenario < ymin ts , which indicates that this Crop has undergone scaling. Therefore, bicubic interpolation is performed on the cropped frame to fix its resolution to [r × ymin ts , ymin ts .
[0158] In this embodiment, since the size of the second image set of the N scenarios is corrected, the image sizes of the image frames of the generated video can be ensured to be the same, so as to improve the display effect of the video.
[0159] In addition, separate processing is performed for different scenarios, which can improve the prediction speed, avoid image jitter, and adapt to a wider range of multi-scenario video types.
[0160] In one embodiment, performing object detection on multiple image frames in the first image set respectively to obtain object detection results of the multiple image frames in the first image set, and identifying key image frames in the first image set based on the object detection results of the multiple image frames in the first image set includes:
[0161] Performing object detection on multiple image frames in the first image set respectively to obtain object detection results of the multiple image frames in the first image set;
[0162] Performing a second filtering on the object detection results of the multiple image frames in the first image set, where the second filtering is used to filter out at least one of the following: object detection results with an object detection box smaller than a preset detection box, and object detection results with an object detection type not belonging to a preset detection category;
[0163] Identifying key image frames in the first image set based on the object detection results of the multiple image frames in the first image set that have undergone the second filtering.
[0164] Among them, the above preset detection box can be an empirical value. For example: filtering out detection boxes with an area (ymax - ymin) * (xmax - xmin) less than γ, and γ defaults to 16, 15, etc.
[0165] The above preset detection categories can include types such as faces, bodies, animals, etc.
[0166] In this embodiment, the above second filtering can save computational overhead.
[0167] It should be noted that in the embodiments of the present application, for the processing of object detection results, such as object trajectory tracking, the object detection results after the second filtering can be used for processing, which will not be elaborated here.
[0168] An embodiment, as Figure 3 shown, includes the following steps:
[0169] Step S301, input a video;
[0170] Step S302, perform OCR detection on subtitles;
[0171] Step S303, subtitle cropping & frame extraction. Specifically, this step can be the second cropping of the image frames in the first image set according to the second cropping frame in the above embodiment;
[0172] Step S304, scene detection. This step can perform scene detection on the video frame set obtained in step S303 to determine at least one scene;
[0173] Step S305, perform object detection and object trajectory tracking detection on the input frames;
[0174] Step S306, sort the object trajectory tracking. Specifically, it can be sorted according to the first score calculated in the above embodiment;
[0175] Step S307, determine the key frames and the first cropping frame. Specifically, in this step, the key frames and the first cropping frame of each scene are respectively determined. Among them, the first cropping frame can also be called the initialization cropping frame;
[0176] Step S308, perform the first cropping. The first cropping specifically includes directly cropping according to the first cropping frame if there is no conflict, moving the first cropping frame to frame part of the object detection frame, moving the first cropping frame to exclude part of the object detection frame, and scaling the first cropping frame to exclude part of the object detection frame;
[0177] Step S309, video encoding, specifically generating a video as in the above embodiment;
[0178] Step S310, output the video.
[0179] In some embodiments, the embodiments of the present application can be applied to the algorithm for converting text and images into video. The specific algorithm process is as follows:
[0180] For the input video, first extract frames, perform logo / watermark detection on each extracted frame, repair the detected logo / watermark detection frame, remove the logo / watermark. For the video frames after removing the logo / watermark, first perform OCR subtitle detection, determine whether there is a static subtitle frame. If so, call the above video generation method for processing, otherwise directly synthesize and output the video.
[0181] In the present disclosure, since the multiple image frames in the first image set are cropped based on the object detection results of the first cropping frame and the multiple image frames in the first image set, the cropped image frames can be more easily matched with the object detection results, thereby improving the cropping effect of the image frames and further making the display effect of the finally generated video better.
[0182] Please refer to Figure 4a , Figure 4a which is a video generation device provided by the present disclosure. As Figure 4a shown in
[0183] The detection module 401 is configured to perform object detection on multiple image frames in the first image set respectively, obtain the object detection results of the multiple image frames in the first image set, and identify the key image frames in the first image set based on the object detection results of the multiple image frames in the first image set;
[0184] The first cropping module 402 is configured to generate a first cropping frame according to the image size and height information of the key image frame, and perform a first cropping on the multiple image frames in the first image set based on the first cropping frame and the object detection results of the multiple image frames in the first image set, to obtain a second image set, where the height information is the height information of the multiple image frames in the first image set;
[0185] The generation module 403 is configured to generate a video based on the multiple image frames in the second image set.
[0186] In one embodiment, as Figure 4b shown in
[0187] The recognition module 404 is configured to recognize the text position information of the multiple image frames in the third image set;
[0188] The second cropping module 405 is configured to generate a second cropping frame based on the text position information, and perform a second cropping on the image frames in the third image set according to the second cropping frame, to obtain the first image set.
[0189] In one embodiment, the text position information includes: a first text box coordinate information set of the multiple image frames included in the third image set, and the text box coordinate information in the first text box coordinate information set includes horizontal axis coordinate information and vertical axis coordinate information;
[0190] The second cropping module 405 is configured to select first vertical axis coordinate information from the first text box coordinate information set, where the distance from the first vertical axis coordinate information to the top of the image frame is less than the distances from other vertical axis coordinate information to the top of the image frame, and the other vertical axis coordinate information is all or part of the vertical axis coordinate information in the first text box coordinate information set except the first vertical axis coordinate information; and generate a second cropping box based on the first vertical axis coordinate information.
[0191] In one embodiment, the multiple image frames included in the third image set have the same size, the width of the second cropping box is the same as the width of the multiple image frames included in the third image set, and in the height direction, the first cropping box is from the top of the multiple image frames included in the third image set to the coordinate position represented by the first vertical axis coordinate information.
[0192] In one embodiment, the second cropping module 405 is configured to perform a first filtering on the first text box coordinate information set to obtain a second text box coordinate information set, where the first filtering is used to filter out the text box coordinate information in the first text box coordinate information set whose distance to the top of the image frame is less than a first preset distance; and select first vertical axis coordinate information from the second text box coordinate information set, and the other vertical axis coordinate information is all or part of the vertical axis coordinate information in the second text box coordinate information set except the first vertical axis coordinate information.
[0193] In one embodiment, as Figure 4c shown, the detection module 401 includes:
[0194] A first detection unit 4011, configured to perform object detection and object trajectory tracking on each of the multiple image frames in the first image set, and obtain the object detection results and object trajectory tracking results of the multiple image frames in the first image set, where the object trajectory tracking results are used to represent the trajectory information of tracking the objects detected by the object detection;
[0195] A first recognition unit 4012, configured to identify key image frames in the first image set based on the object detection results and object trajectory tracking results of the multiple image frames in the first image set.
[0196] In one embodiment, the first detection unit 4011 is configured to perform object detection and object trajectory tracking on each of the image frames in an image subset of the first image set, and obtain the object detection results and object trajectory tracking results of the image frames in the image subset, where the image subset includes multiple image frames sampled at intervals in the first image set;
[0197] Among them, the target trajectory tracking result of the first image frame in the image subset includes: a first trajectory detection box, the frame sequence number of the first image frame, and a first trajectory array. The first trajectory detection box is the target detection result of the first image frame. The first trajectory array includes the target detection results of the first image frame. The first image frame is the first image frame in the image subset;
[0198] Among them, the target trajectory tracking result of the second image frame in the image subset includes: a second trajectory detection box, the frame sequence number of the second image frame, and a second trajectory array. The second trajectory detection box is the target detection result of the second image frame. The second trajectory array includes the target detection results of the first image frame and the second image frame. Alternatively, the second trajectory array is the target detection result of the second image frame. The second image frame is an image frame in the image subset other than the first image frame.
[0199] In one embodiment, as Figure 4d shown, the device further includes:
[0200] A calculation module 406, configured to calculate the overlap degree IOU between the target detection result of the second image frame and the trajectory detection box in the special target trajectory tracking result. The special target trajectory tracking result includes: the target trajectory tracking results whose difference between the frame sequence numbers and the frame sequence number of the second image frame is within a preset range;
[0201] Among them, when the IOU between the target detection result of the second image frame and the IOU in the matching target trajectory tracking result is greater than a preset threshold, the second trajectory array includes the target detection results of the first image frame and the second image frame. When the IOU between the target detection result of the second image frame and the IOU in the matching target trajectory tracking result is less than or equal to the preset threshold, the second trajectory array is the target detection result of the second image frame;
[0202] The matching target trajectory tracking result is the target trajectory tracking result with the highest matching degree with the target detection result of the second image frame in the special target trajectory tracking result.
[0203] In one embodiment, the first recognition unit 4012 is configured to calculate the first scores of the target trajectory tracking results of multiple image frames in the image subset respectively. Among them, the first score of each target trajectory tracking result is calculated based on the coordinates in the trajectory array of the target trajectory tracking result; for each image frame in the image subset, calculate the second score of the image frame based on the target detection result of the image frame and the first score; and select the image frame with the highest second score in the image subset as the key image frame in the first image set.
[0204] In one embodiment, as Figure 4e shown, the first cropping module 402 includes:
[0205] A calculation unit 4021, configured to calculate the aspect ratio of the key image frame based on the image size of the key image frame;
[0206] A generation unit 4022, configured to generate a first cropping frame based on the aspect ratio and the height information;
[0207] A comparison unit 4023, configured to compare the object detection results of multiple image frames in the first image set with the first cropping frame respectively to obtain comparison results;
[0208] A cropping unit 4024, configured to perform a first cropping on multiple image frames in the second image set respectively based on the comparison results and the first cropping frame to obtain a second image set.
[0209] In one embodiment, the cropping unit 4024 is configured to perform a first cropping on multiple image frames in the first image set respectively based on the comparison results, the first cropping frame and a target data group, and the target data group is a multi-dimensional array generated based on the coordinates in the key image frame and the image size of the key image frame;
[0210] Wherein, the first cropping includes at least one of the following:
[0211] Cropping according to the first cropping frame;
[0212] Updating the first cropping frame, cropping based on the updated cropping frame, and the update is performed based on the target data group;
[0213] Displacing the first cropping frame, and cropping based on the displaced cropping frame.
[0214] In one embodiment, the first image set includes N scenes, and for each scene, there is a corresponding key image frame, the first cropping frame and the second image set, where N is an integer greater than 1;
[0215] The generation module 403 is configured to correct the sizes of the second image sets of the N scenes, and synthesize the second image sets of the N scenes with corrected sizes to generate a video;
[0216] Wherein, the size correction includes:
[0217] Interpolate the to-be-corrected image frame that is cropped based on the scaled cropping box so that the image size of the to-be-corrected image frame matches the image sizes of other image frames, where the other image frames are the image frames in the second image set of the N scenarios except the to-be-corrected image frame.
[0218] In one embodiment, as Figure 4f , the detection module 401 includes:
[0219] A second detection unit 4013, configured to perform object detection on multiple image frames in the first image set respectively to obtain object detection results of the multiple image frames in the first image set;
[0220] A filtering unit 4014, configured to perform a second filtering on the object detection results of the multiple image frames in the first image set, where the second filtering is used to filter out at least one of the following: object detection results with an object detection box smaller than a preset detection box, and object detection results with an object detection type not belonging to a preset detection category;
[0221] A second recognition unit 4015, configured to recognize key image frames in the first image set based on the object detection results of the multiple image frames in the first image set that have undergone the second filtering.
[0222] The video generation device provided by the present disclosure can implement each process implemented by the video generation method provided by the present disclosure and achieve the same technical effects. To avoid repetition, it will not be elaborated here.
[0223] In the technical solution of the present disclosure, the acquisition, storage, and application of user personal information involved all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0224] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0225] Wherein, the above-mentioned electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the video generation method provided by the present disclosure.
[0226] The above-mentioned readable storage medium stores computer instructions, wherein the computer instructions are used to cause the computer to execute the video generation method provided by the present disclosure.
[0227] The above-mentioned computer program product includes a computer program, and the computer program implements the video generation method provided by the present disclosure when executed by a processor.
[0228] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 500 that may be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as, for example, personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0229] As Figure 5 shown, the device 500 includes a computing unit 501 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the device 500 can also be stored. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0230] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, such as, for example, a keyboard, a mouse, etc.; an output unit 507, such as, for example, various types of displays, speakers, etc.; a storage unit 508, such as, for example, a magnetic disk, an optical disk, etc.; and a communication unit 509, such as, for example, a network card, a modem, a wireless communication transceiver, etc. The communication unit 509 allows the device 500 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0231] The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 501 executes the various methods and processes described above, such as the video generation method. For example, in some embodiments, the video generation method can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded into the RAM 503 and executed by the computing unit 501, one or more steps of the video generation method described above can be executed. Alternatively, in other embodiments, the computing unit 501 can be configured to execute the video generation method in any other suitable manner (e.g., by means of firmware).
[0232] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0233] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0234] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0235] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0236] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of a communication network include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0237] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.
[0238] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present disclosure can be achieved, and no limitation is imposed herein.
[0239] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A video generation method, comprising: Performing object detection on multiple image frames in a first image set respectively to obtain object detection results of the multiple image frames in the first image set, and identifying key image frames in the first image set based on the object detection results of the multiple image frames in the first image set; Generating a first cropping frame according to the image size and height information of the key image frame, and performing a first cropping on the multiple image frames in the first image set based on the first cropping frame and the object detection results of the multiple image frames in the first image set to obtain a second image set, where the height information is the height information of the multiple image frames in the first image set; Generating a video based on the multiple image frames in the second image set; The method further comprises: Identifying text position information of multiple image frames in a third image set, where the text position information includes: a first text box coordinate information set of the multiple image frames included in the third image set, and the text box coordinate information in the first text box coordinate information set includes horizontal axis coordinate information and vertical axis coordinate information; Selecting a first vertical axis coordinate information in the first text box coordinate information set, where the distance from the first vertical axis coordinate information to the top of the image frame is less than the distances from other vertical axis coordinate information to the top of the image frame, and the other vertical axis coordinate information is all or part of the vertical axis coordinate information in the first text box coordinate information set except the first vertical axis coordinate information; Generating a second cropping frame based on the first vertical axis coordinate information, and performing a second cropping on the multiple image frames in the third image set according to the second cropping frame to obtain the first image set; wherein, the multiple image frames included in the third image set have the same size, the width of the second cropping frame is the same as the width of the multiple image frames included in the third image set, and the first cropping frame is from the top of the multiple image frames included in the third image set to the coordinate position represented by the first vertical axis coordinate information in the height direction; Wherein, the selecting the first vertical axis coordinate information in the first text box coordinate information set includes: Performing a first filtering on the first text box coordinate information set to obtain a second text box coordinate information set, where the first filtering is used to filter out the text box coordinate information in the first text box coordinate information set whose distance to the top of the image frame is less than a first preset distance; Selecting the first vertical axis coordinate information in the second text box coordinate information set, where the other vertical axis coordinate information is all or part of the vertical axis coordinate information in the second text box coordinate information set except the first vertical axis coordinate information.
2. The method according to claim 1, wherein The performing object detection on multiple image frames in the first image set respectively to obtain object detection results of the multiple image frames in the first image set, and identifying key image frames in the first image set based on the object detection results of the multiple image frames in the first image set, includes: Perform object detection and object trajectory tracking on multiple image frames in the first image set respectively to obtain the object detection results and object trajectory tracking results of the multiple image frames in the first image set, where the object trajectory tracking results are used to represent the trajectory information of tracking the objects detected by the object detection; Identify key image frames in the first image set based on the object detection results and object trajectory tracking results of the multiple image frames in the first image set.
3. The method according to claim 2, wherein The performing object detection and object trajectory tracking on multiple image frames in the first image set respectively to obtain the object detection results and object trajectory tracking results of the multiple image frames in the first image set includes: Perform object detection and object trajectory tracking on the image frames in an image subset of the first image set respectively to obtain the object detection results and object trajectory tracking results of the image frames in the image subset, where the image subset includes multiple image frames sampled at intervals in the first image set; Among them, the object trajectory tracking result of the first image frame in the image subset includes: a first trajectory detection box, the frame number of the first image frame, and a first trajectory array. The first trajectory detection box is the object detection result of the first image frame, and the first trajectory array includes the object detection result of the first image frame. The first image frame is the first image frame in the image subset; Among them, the object trajectory tracking result of the second image frame in the image subset includes: a second trajectory detection box, the frame number of the second image frame, and a second trajectory array. The second trajectory detection box is the object detection result of the second image frame, and the second trajectory array includes the object detection result of the first image frame and the object detection result of the second image frame, or the second trajectory array is the object detection result of the second image frame. The second image frame is an image frame in the image subset other than the first image frame.
4. The method according to claim 3, wherein The method further includes: Calculate the overlap degree IOU between the object detection result of the second image frame and the trajectory detection box in the special object trajectory tracking result. The special object trajectory tracking result includes: the object trajectory tracking result where the difference between the frame number and the frame number of the second image frame is within a preset range; Among them, when the IOU between the object detection result of the second image frame and the IOU in the matching object trajectory tracking result is greater than a preset threshold, the second trajectory array includes the object detection result of the first image frame and the object detection result of the second image frame. When the IOU between the object detection result of the second image frame and the IOU in the matching object trajectory tracking result is less than or equal to the preset threshold, the second trajectory array is the object detection result of the second image frame; The matching object trajectory tracking result is the object trajectory tracking result with the highest matching degree with the object detection result of the second image frame in the special object trajectory tracking result.
5. The method according to claim 3, wherein, The identifying key image frames in the first image set based on the object detection results and object trajectory tracking results of the multiple image frames in the first image set includes: Calculate the first score of the target trajectory tracking results of multiple image frames in the image subset respectively, where the first score of each target trajectory tracking result is calculated based on the coordinates in the trajectory array of the target trajectory tracking result; For each image frame in the image subset, calculate the second score of the image frame based on the target detection result of the image frame and the first score; Select the image frame with the highest second score in the image subset as the key image frame in the first image set.
6. The method according to claim 1, wherein Generating a first cropping frame according to the image size and height information of the key image frame, and performing a first cropping on multiple image frames in the first image set based on the first cropping frame and the target detection results of multiple image frames in the first image set to obtain a second image set, including: Calculate the aspect ratio of the key image frame based on the image size of the key image frame; Generate a first cropping frame based on the aspect ratio and the height information; Compare the target detection results of multiple image frames in the first image set with the first cropping frame respectively to obtain comparison results; For multiple image frames in the first image set, perform a first cropping respectively based on the comparison results and the first cropping frame to obtain a second image set.
7. The method according to claim 6, wherein, The performing a first cropping on multiple image frames in the first image set respectively based on the comparison results and the first cropping frame includes: For multiple image frames in the first image set, perform a first cropping respectively based on the comparison results, the first cropping frame and a target data group, where the target data group is a multi-dimensional array generated based on the coordinates in the key image frame and the image size of the key image frame; Among them, the first cropping includes at least one of the following: Crop according to the first cropping frame; Update the first cropping frame, and perform cropping based on the updated cropping frame, and the update is based on the target data group; Displace the first cropping frame, and perform cropping based on the displaced cropping frame.
8. The method according to claim 7, wherein, The first image set includes N scenes, and there is a corresponding key image frame, the first cropping frame and the second image set for each scene, where N is an integer greater than 1; Generating a video based on multiple image frames in the second image set, including: Correct the sizes of the second image sets of the N scenes, and synthesize the corrected second image sets of the N scenes to generate a video; Among them, the size correction includes: Perform interpolation processing on the image frame to be corrected cropped based on the scaled cropping frame, so that the image size of the image frame to be corrected matches the image sizes of other image frames, where the other image frames are the image frames in the second image sets of the N scenes except the image frame to be corrected.
9. The method according to claim 1, wherein Performing target detection on multiple image frames in the first image set respectively to obtain the target detection results of multiple image frames in the first image set, and identifying the key image frame in the first image set based on the target detection results of multiple image frames in the first image set, including: Perform object detection on multiple image frames in the first image set respectively to obtain the object detection results of the multiple image frames in the first image set; Perform a second filtering on the object detection results of the multiple image frames in the first image set, where the second filtering is used to filter out at least one of the following: object detection results with an object detection box smaller than a preset detection box, and object detection results with an object detection type not belonging to a preset detection category; Based on the object detection results of the multiple image frames in the first image set that have passed the second filtering, identify the key image frames in the first image set.
10. A video generation device, comprising: A detection module, configured to perform object detection on multiple image frames in a first image set respectively to obtain the object detection results of the multiple image frames in the first image set, and based on the object detection results of the multiple image frames in the first image set, identify the key image frames in the first image set; A first cropping module, configured to generate a first cropping box according to the image size and height information of the key image frames, and based on the first cropping box and the object detection results of the multiple image frames in the first image set, perform a first cropping on the multiple image frames in the first image set to obtain a second image set, where the height information is the height information of the multiple image frames in the first image set; A generation module, configured to generate a video based on the multiple image frames in the second image set; The device further comprises: An identification module, configured to identify the text position information of multiple image frames in a third image set; A second cropping module, configured to generate a second cropping box based on the text position information, and perform a second cropping on the image frames in the third image set according to the second cropping box to obtain the first image set; Wherein, the text position information includes: a first text box coordinate information set of multiple image frames included in the third image set, and the text box coordinate information in the first text box coordinate information set includes horizontal axis coordinate information and vertical axis coordinate information; The second cropping module is configured to select a first vertical axis coordinate information in the first text box coordinate information set, where the distance from the first vertical axis coordinate information to the top of the image frame is less than the distances from other vertical axis coordinate information to the top of the image frame, and the other vertical axis coordinate information is all or part of the vertical axis coordinate information in the first text box coordinate information set except the first vertical axis coordinate information; and generate a second cropping box based on the first vertical axis coordinate information; The sizes of the multiple image frames included in the third image set are the same, the width of the second cropping box is the same as the width of the multiple image frames included in the third image set, and the first cropping box is from the top of the multiple image frames included in the third image set to the coordinate position represented by the first vertical axis coordinate information in the height direction; The second cropping module is used to perform a first filtering on the first text box coordinate information set to obtain a second text box coordinate information set. The first filtering is used to filter out the text box coordinate information in the first text box coordinate information set that is less than a first preset distance from the top of the image frame; and select a first vertical axis coordinate information in the second text box coordinate information set, where the other vertical axis coordinate information is all or part of the vertical axis coordinate information in the second text box coordinate information set except the first vertical axis coordinate information.
11. The apparatus according to claim 10, wherein The detection module includes: A first detection unit, which is used to perform object detection and object trajectory tracking on multiple image frames in the first image set respectively, to obtain the object detection results and object trajectory tracking results of the multiple image frames in the first image set. The object trajectory tracking result is used to represent the trajectory information of the object tracked by the object detection. A first recognition unit, which is used to identify the key image frames in the first image set based on the object detection results and object trajectory tracking results of the multiple image frames in the first image set.
12. The device according to claim 11, wherein, The first detection unit is used to perform object detection and object trajectory tracking on the image frames in the image subset of the first image set respectively, to obtain the object detection results and object trajectory tracking results of the image frames in the image subset. The image subset includes multiple image frames sampled at intervals in the first image set. Among them, the object trajectory tracking result of the first image frame in the image subset includes: a first trajectory detection box, the frame number of the first image frame, and a first trajectory array. The first trajectory detection box is the object detection result of the first image frame, and the first trajectory array includes the object detection result of the first image frame. The first image frame is the first image frame in the image subset. Among them, the object trajectory tracking result of the second image frame in the image subset includes: a second trajectory detection box, the frame number of the second image frame, and a second trajectory array. The second trajectory detection box is the object detection result of the second image frame, and the second trajectory array includes the object detection result of the first image frame and the object detection result of the second image frame, or the second trajectory array is the object detection result of the second image frame. The second image frame is the image frame in the image subset except the first image frame.
13. The apparatus according to claim 12, wherein, The device further includes: A calculation module, which is used to calculate the overlap degree IOU between the object detection result of the second image frame and the trajectory detection box in the special object trajectory tracking result. The special object trajectory tracking result includes: the object trajectory tracking result whose difference between the frame number and the frame number of the second image frame is within a preset range. Wherein, when the Intersection over Union (IOU) between the object detection result of the second image frame and the matching object trajectory tracking result is greater than a preset threshold, the second trajectory array includes the object detection result of the first image frame and the object detection result of the second image frame; when the IOU between the object detection result of the second image frame and the matching object trajectory tracking result is less than or equal to the preset threshold, the second trajectory array is the object detection result of the second image frame. The matching object trajectory tracking result is the object trajectory tracking result with the highest matching degree to the object detection result of the second image frame among the special object trajectory tracking results.
14. The apparatus according to claim 12, wherein, The first recognition unit is used to calculate the first scores of the object trajectory tracking results of multiple image frames in the image subset respectively, wherein the first score of each object trajectory tracking result is calculated based on the coordinates in the trajectory array of the object trajectory tracking result; for each image frame in the image subset, calculate the second score of the image frame based on the object detection result of the image frame and the first score; and select the image frame with the highest second score in the image subset as the key image frame in the first image set.
15. The apparatus according to claim 10, wherein, The first cropping module includes: a calculation unit, configured to calculate the aspect ratio of the key image frame based on the image size of the key image frame; a generation unit, configured to generate a first cropping frame based on the aspect ratio and the height information; a comparison unit, configured to compare the object detection results of multiple image frames in the first image set with the first cropping frame respectively to obtain comparison results; a cropping unit, configured to perform first cropping on multiple image frames in the second image set respectively based on the comparison results and the first cropping frame to obtain a second image set.
16. The device according to claim 15, wherein, The cropping unit is configured to perform first cropping on multiple image frames in the first image set respectively based on the comparison results, the first cropping frame and the target data group, and the target data group is a multi-dimensional array generated based on the coordinates in the key image frame and the image size of the key image frame; Wherein, the first cropping includes at least one of the following: cropping according to the first cropping frame; updating the first cropping frame and cropping based on the updated cropping frame, and the update is based on the target data group; displacing the first cropping frame and cropping based on the displaced cropping frame.
17. The apparatus according to claim 16, wherein, The first image set includes N scenes, and there are corresponding key image frames, first cropping frames and second image sets for each scene, where N is an integer greater than 1; The generation module is configured to correct the sizes of the second image sets of the N scenes, and synthesize the second image sets of the N scenes with corrected sizes to generate a video; Wherein, the size correction includes: Interpolate the to-be-corrected image frame that is cropped based on the scaled cropping frame so that the image size of the to-be-corrected image frame matches the image size of other image frames, where the other image frames are the image frames in the second image set of the N scenarios except the to-be-corrected image frame.
18. The apparatus according to claim 10, wherein The detection module includes: A second detection unit configured to perform object detection on multiple image frames in the first image set respectively to obtain object detection results of the multiple image frames in the first image set; A filtering unit configured to perform a second filtering on the object detection results of the multiple image frames in the first image set, where the second filtering is used to filter out at least one of the following: object detection results with object detection frames smaller than a preset detection frame, object detection results with object detection types not belonging to a preset detection category; A second recognition unit configured to recognize key image frames in the first image set based on the object detection results of the multiple image frames in the first image set that have undergone the second filtering.
19. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-9.
21. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-9.
Citation Information
Patent Citations
Video generation method, device and equipment and storage medium
CN112714263A
Live video editing method and device and computer equipment
CN113542777A