A video dotting method and device, electronic equipment and storage medium
By performing scene segmentation on the original video and matching with a deep learning model, the location of the points in the target video is automatically determined and marked, solving the problem of manually repeating the marking after video replacement, thus improving user experience and video playback volume.
Patent Information
- Application Number
- CN202411020470.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-26
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2044-07-26
AI Technical Summary
In existing technologies, after a video is replaced, manual re-marking is required, which increases workload and reduces the timeliness of users' access to real-time content, affecting video playback volume and user experience.
By performing segmentation on the original video, determining the segmentation feature vector, and associating the points with the target segmentation, a deep learning model is used to match the segmentation and automatically determine the position of the points in the target video for marking.
It enables automatic marking of replacement videos, reducing the workload of repetitive marking, shortening the time difference, ensuring the timeliness of users' access to instant content and the amount of video playback, and improving the user experience.
Smart Images

Figure CN118828127B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video processing, and more particularly to a video dotting method and device, an electronic device, and a storage medium. BACKGROUND
[0002] In the whole link of video content production and operation, when the video is affected by multiple-dimensional elements such as sound and picture quality, content safety, network public opinion, policy guidance, and sensitive figures, the original video needs to be replaced by a modified target video. Since the points (such as video head and tail points, video highlight points, video star points, short video intelligent recommendation points, and member preview product points) of the original video operation are only associated with the original video, therefore, these points of the original video operation in the replaced target video will all be invalid.
[0003] In order to ensure that the replaced target video is played according to the points of the original video operation, the existing technology requires the operation personnel to manually operate and dot the replaced target video again, which not only leads to repeated dotting work investment, but also causes time difference due to repeated dotting, affecting the timeliness of users obtaining instant content, thereby affecting the video play quantity and user experience. SUMMARY
[0004] Therefore, the present application discloses a video dotting method and device, an electronic device, and a storage medium to automatically dot the target video after replacing the original video, reduce the work investment caused by repeated dotting, shorten the time difference caused by repeated dotting, ensure the timeliness of users obtaining instant content and the video play quantity, and thus improve the user experience.
[0005] A video dotting method, comprising:
[0006] performing shot processing on an original video to obtain a plurality of first shot segments, and determining a first shot feature vector corresponding to each first shot segment;
[0007] operating and dotting the original video to obtain each point;
[0008] associating each point to a corresponding first target shot segment, and determining the position of the point in the first target shot segment;
[0009] performing shot processing on a target video after replacing the original video to obtain a plurality of second shot segments, and determining a second shot feature vector corresponding to each second shot segment;
[0010] finding a second target shot segment from the plurality of second shot segments, in which the second shot feature vector matches the first shot feature vector corresponding to the first target shot segment;
[0011] Determine the dotting position of the point in the first target sub-mirror segment in the second target sub-mirror segment, and perform a dotting operation.
[0012] Optionally, the original video is subjected to sub-mirror processing to obtain a plurality of first sub-mirror segments, and a first sub-mirror feature vector corresponding to each first sub-mirror segment is determined, comprising:
[0013] The original video is subjected to sub-mirror algorithm to obtain a plurality of first sub-mirror segments;
[0014] Determine the sub-mirror start and end time of each first sub-mirror segment in the original video;
[0015] Determine the total number of frames of each first sub-mirror segment;
[0016] Extract a sub-mirror intermediate frame from each first sub-mirror segment, and process the sub-mirror intermediate frame using a deep learning pre-training model to obtain an intermediate frame feature vector;
[0017] The sub-mirror start and end time, the total number of frames, and the intermediate frame feature vector are combined into a triple, and the triple is determined as the first sub-mirror feature vector.
[0018] Optionally, the first target sub-mirror segment corresponding to each point is associated, and the position of the point in the first target sub-mirror segment is determined, comprising:
[0019] According to the point time point corresponding to each point, the point is associated with the corresponding first target sub-mirror segment;
[0020] According to the time length of the point time point from the sub-mirror start point of the first target sub-mirror segment, the position of the point in the first target sub-mirror segment is determined.
[0021] Optionally, the second target sub-mirror segment in which the second sub-mirror feature vector matches the first sub-mirror feature vector corresponding to the first target sub-mirror segment is found from a plurality of second sub-mirror segments, comprising:
[0022] Based on the total number of frames contained in each second sub-mirror feature vector, each candidate sub-mirror segment with the same target total number of frames is found from a plurality of second sub-mirror segments, wherein the target total number of frames is the total number of frames contained in the first sub-mirror feature vector corresponding to the first target sub-mirror segment.
[0023] calculate vector distances between a target intermediate frame feature vector and each candidate intermediate frame feature vector, wherein the target intermediate frame feature vector is an intermediate frame feature vector contained in the first target sub-shot feature vector corresponding to the first target sub-shot segment, and each candidate intermediate frame feature vector is an intermediate frame feature vector contained in the second sub-shot feature vector corresponding to the candidate sub-shot segment;
[0024] determine a target candidate sub-shot segment corresponding to a target vector distance satisfying a preset distance condition among all the vector distances as the second target sub-shot segment.
[0025] Optionally, the determining of the target candidate sub-shot segment corresponding to the target vector distance satisfying the preset distance condition among all the vector distances as the second target sub-shot segment comprises:
[0026] finding a minimum vector distance from all the vector distances;
[0027] in a case where the minimum vector distance is not greater than a preset distance threshold, determining the minimum vector distance as the target vector distance;
[0028] determining the target candidate sub-shot segment corresponding to the target vector distance as the second target sub-shot segment.
[0029] Optionally, the determining of the dotting position of the dot position in the first target sub-shot segment in the second target sub-shot segment and the performing of the dotting operation comprise:
[0030] determining a dot position offset based on a starting point corresponding to the second target sub-shot segment and a starting point corresponding to the first target sub-shot segment;
[0031] calculating a sum of a dot position time point of the dot position in the first target sub-shot segment and the dot position offset to obtain a target dot position time point of the dot position in the second target sub-shot segment;
[0032] taking a position of the target dot position time point in the second target sub-shot segment as the dotting position and performing the dotting operation.
[0033] Optionally, the method further comprises:
[0034] in a case where the second target sub-shot segment is not found from the plurality of second sub-shot segments, determining a first same sub-shot segment corresponding to a front adjacent sub-shot segment of the first target sub-shot segment in the target video and a second same sub-shot segment corresponding to a rear adjacent sub-shot segment of the first target sub-shot segment in the target video;
[0035] determine a minimum artificial clocking video segment interval between the first same sub-shot segment and the second same sub-shot segment;
[0036] output the minimum artificial clocking video segment interval.
[0037] A video clocking device, comprising:
[0038] a first sub-shot processing unit configured to perform sub-shot processing on an original video to obtain a plurality of first sub-shot segments and determine a first sub-shot feature vector corresponding to each of the first sub-shot segments;
[0039] a clocking unit configured to perform operation clocking on the original video to obtain each clocking point;
[0040] a clocking point association unit configured to associate each clocking point to a corresponding first target sub-shot segment and determine a position of the clocking point in the first target sub-shot segment;
[0041] a second sub-shot processing unit configured to perform sub-shot processing on a target video obtained by replacing the original video to obtain a plurality of second sub-shot segments and determine a second sub-shot feature vector corresponding to each of the second sub-shot segments;
[0042] a matching unit configured to find, from the plurality of second sub-shot segments, a second target sub-shot segment in which the second sub-shot feature vector matches the first sub-shot feature vector corresponding to the first target sub-shot segment;
[0043] a clocking position determination unit configured to determine a clocking position of the clocking point in the first target sub-shot segment in the second target sub-shot segment and perform clocking operation.
[0044] A computer storage medium storing at least one instruction, which, when executed by a processor, implements the video clocking method described above.
[0045] An electronic device, comprising a memory and a processor;
[0046] the memory is configured to store at least one instruction;
[0047] the processor is configured to execute the at least one instruction to implement the video clocking method described above.
[0048] From the technical solution, the application discloses a video dotting method and device, electronic equipment and storage medium, the original video is processed to obtain a plurality of first mirror segments, and a first mirror feature vector corresponding to each first mirror segment is determined, the original video is operated to obtain each point, each point is associated with the corresponding first target mirror segment, and the position of the point in the first target mirror segment is determined, the target video after the original video is replaced is processed to obtain a plurality of second mirror segments, and a second mirror feature vector corresponding to each second mirror segment is determined, the second target mirror segment in which the second mirror feature vector matches the first mirror feature vector corresponding to the first target mirror segment is found from the plurality of second mirror segments, the dotting position of the point in the first target mirror segment in the second target mirror segment is determined, and the dotting operation is performed. The first target mirror segment in which the point in the original video is marked is matched with each second mirror segment in the target video, the second target mirror segment matched with the first target mirror segment is found, and the latest dotting position of the point in the first target mirror segment in the second target mirror segment is determined to perform dotting, the target video after the original video is replaced is automatically dotted, thereby not only reducing the work input caused by repeated dotting, but also shortening the time difference caused by repeated dotting, ensuring the timeliness of instant content acquisition and video play volume of the user, and further improving the user experience. BRIEF DESCRIPTION OF DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are only embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor based on the disclosed drawings.
[0050] Figure 1 A flow chart of a video dotting method disclosed by the embodiment of the present application;
[0051] Figure 2 A structural schematic diagram of a video dotting device disclosed by the embodiment of the present application;
[0052] Figure 3 A structural schematic diagram of an electronic equipment disclosed by the embodiment of the present application. DETAILED DESCRIPTION
[0053] With reference to the drawings of the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments of the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of the present application.
[0054] The embodiment of the present application discloses a video marking method and device, electronic equipment and storage medium, by matching the first target split lens segment marked with a point in the original video with each second split lens segment in the target video, finding the second target split lens segment matched with the first target split lens segment, and determining the latest marking position of the point in the first target split lens segment in the second target split lens segment to mark, realizing automatic marking of the target video after replacing the original video, thereby not only reducing the work input caused by repeated marking, but also shortening the time difference caused by repeated marking, ensuring the timeliness of instant content acquisition and video play volume of the user, and further improving the user experience.
[0055] Referring to Figure 1 The embodiment of the present application discloses a video marking method flow chart, and the method comprises:
[0056] Step S101, performing split lens processing on the original video to obtain a plurality of first split lens segments, and determining a first split lens feature vector corresponding to each first split lens segment.
[0057] In the embodiment, the original video can be a media asset video initially stored in a media asset system, and specifically can be a media asset long video.
[0058] In the embodiment, the original video is divided into split lens level video units by split lens processing, and a plurality of first split lens segments are obtained.
[0059] In actual application, the original video can be processed by using a PySceneDetect library for split lens processing to obtain a plurality of first split lens segments. PySceneDetect is a powerful video scene detection library, which is based on Python and OpenCV, and can detect scene changes in a video, including lens switching, fade-in and fade-out, and automatically divide the video into a plurality of independent scenes.
[0060] In the embodiment, the first split lens feature vector corresponding to each first split lens segment is a triple, which includes split lens start and end time, split lens total frame number and intermediate frame feature vector. This way can represent the video at minimum cost, without the need to calculate and store a large number of feature vectors.
[0061] Step S102, performing operation marking on the original video to obtain each point.
[0062] The operation dotting of the original video generally refers to marking the highlight, the interesting point, the star appearance time point, the advertisement insertion point, the song start and end point in the variety show, the link start and end point, etc.
[0063] After the operation dotting of the original video is completed, the three tuple information (point position title, point position type, and point position time point) corresponding to each point position can be obtained by browsing the video marking time point position and the type of the operation point position.
[0064] The operation dotting of the original video can specifically include: media operation, member operation, advertisement operation, etc. which need to be performed on the original video (such as media long video) initially stored, and the three tuple information (point position title, point position type, and point position time point) can be simplified by browsing the video marking time point position and the type of the operation point position.
[0065] In step S103, each point position is associated with the corresponding first target split shot segment, and the position of the point position in the first target split shot segment is determined.
[0066] In actual application, according to the point position time point corresponding to each point position, the first split shot segment containing the point position time point is determined, which is recorded as the first target split shot segment, so as to associate the point position with the corresponding first target split shot segment. The position of the point position in the first target split shot segment can be determined according to the time length of the point position time point from the split shot start point of the first target split shot segment.
[0067] It should be noted that the first target split shot segment in the embodiment is actually the first split shot segment marked with the point position.
[0068] In step S104, the original video after replacement is subjected to split shot processing to obtain a plurality of second split shot segments, and the second split shot feature vector corresponding to each second split shot segment is determined.
[0069] Similarly, in actual application, the target video can be subjected to split shot processing by using the PySceneDetect library to obtain a plurality of second split shot segments.
[0070] In the embodiment, the second split shot feature vector corresponding to each second split shot segment is a three tuple, which includes: split shot start and end time, split shot total frame number, and intermediate frame feature vector. This way can represent the video at the minimum cost without calculating and storing a large number of feature vectors.
[0071] In step S105, the second target split shot segment in which the second split shot feature vector matches the first split shot feature vector corresponding to the first target split shot segment is found from the plurality of second split shot segments.
[0072] The first target sub-shot segment is matched with the second target sub-shot segment by matching the first sub-shot feature vector corresponding to the first target sub-shot segment with the second sub-shot feature vector corresponding to each second sub-shot segment.
[0073] In step S106, the position of the point in the first target sub-shot segment is determined in the second target sub-shot segment, and a dotting operation is performed.
[0074] In actual application, the offset of the point in the second target sub-shot segment relative to the point in the first target sub-shot segment can be determined first, then the new dotting position of the point in the second target sub-shot segment is determined according to the offset, and the dotting operation is performed, so that full-automatic point synchronization between the target video and the original video in the replacement scene is realized.
[0075] In summary, the present application discloses a video dotting method, which comprises the following steps: performing sub-shot processing on an original video to obtain a plurality of first sub-shot segments, and determining a first sub-shot feature vector corresponding to each first sub-shot segment; performing dotting on the original video to obtain each point, associating each point to the corresponding first target sub-shot segment, and determining the position of the point in the first target sub-shot segment; performing sub-shot processing on a target video obtained by replacing the original video to obtain a plurality of second sub-shot segments, and determining a second sub-shot feature vector corresponding to each second sub-shot segment; finding a second target sub-shot segment from the plurality of second sub-shot segments, in which the second sub-shot feature vector matches the first sub-shot feature vector corresponding to the first target sub-shot segment; determining the dotting position of the point in the first target sub-shot segment in the second target sub-shot segment, and performing a dotting operation. The first target sub-shot segment in which the point is marked in the original video is matched with each second sub-shot segment in the target video, the second target sub-shot segment matching the first target sub-shot segment is found, and the latest dotting position of the point in the first target sub-shot segment in the second target sub-shot segment is determined for dotting, so that automatic dotting is realized for the target video obtained by replacing the original video, thereby not only reducing the work input caused by repeated dotting, but also shortening the time difference caused by repeated dotting, ensuring the timeliness of instant content acquisition and video play volume for users, and further improving user experience.
[0076] To further optimize the above embodiment, step S101 can specifically comprise:
[0077] (1) performing lens cutting on the original video by using a sub-shot algorithm to obtain a plurality of first sub-shot segments;
[0078] The shot algorithm, namely, a video shot technology (Shot Boundary Detection), mainly divides a video into multiple shot segments according to video pictures. The basic idea of video shot is to compare the similarity of each frame of a shot picture. If the picture similarity is high, it belongs to the same shot; if the similarity is very low, it belongs to the next shot. This technology can be used to cut long video content into shot granularity segments.
[0079] The shot algorithm in the embodiment can be a PySceneDetect library.
[0080] (2) Determine the shot start and end time of each first shot segment in the original video.
[0081] When the original video is shot cut, the shot start and end time of each first shot segment in the original video can be determined.
[0082] (3) Determine the total number of frames of each first shot segment.
[0083] When the shot start and end time of each first shot segment in the original video is determined, the total number of frames within the shot start and end time can be determined.
[0084] (4) Extract a shot intermediate frame from each first shot segment, and process the shot intermediate frame using a deep learning pre-trained model to obtain an intermediate frame feature vector.
[0085] The shot intermediate frame extracted from the first shot segment can be a shot intermediate frame picture.
[0086] The deep learning pre-trained model used when processing the shot intermediate frame can be a Resnet (Residual Network), VGGnet, EfficientNet, MocoV3, etc.
[0087] Resnet (Residual Network) is a very popular deep learning model, especially in image recognition, classification and other tasks.
[0088] VGGnet is a deep convolutional neural network model that explores the relationship between the depth of convolutional neural networks and their performance, and successfully builds a 16-19 layer deep convolutional neural network.
[0089] EfficientNet is a high-efficiency and powerful convolutional neural network architecture proposed by Google Brain team, and the core idea of EfficientNet is to balance the depth, width and resolution of the network to achieve efficient model design through compound scaling.
[0090] MoCoV3 (Momentum Contrast v3) is the latest version of the MOCO series, which focuses on enabling computers to learn powerful visual feature representations from large amounts of unlabeled data through the design of intelligent contrast learning strategies.
[0091] The deep learning pre-training model in the embodiment preferentially selects a MocoV3 pre-training model, and the MocoV3 pre-training model is an open source and widely used deep image pre-training model in business.
[0092] (5) The shot start and end time, the total number of shot frames and the intermediate frame feature vector are combined into a triple, and the triple is determined as the first shot feature vector.
[0093] To further optimize the above embodiment, step S103 can specifically include:
[0094] (1) According to the point time corresponding to each point, the point is associated with the corresponding first target shot segment.
[0095] After obtaining each point by operating the original video, the point type and point time corresponding to each point are recorded. Since the first shot feature vector contains the shot start and end time, the first shot segment associated with the point can be determined by judging the shot start and end time to which the point time belongs. For convenience, the first shot segment containing the point is recorded as the first target shot segment in the embodiment.
[0096] (2) According to the time length of the point time from the shot start point of the first target shot segment, the position of the point in the first target shot segment is determined.
[0097] In actual application, when determining the position of the point time in the first target shot segment, the position of the point in the first target shot segment is usually the time length of the point time from the shot start point, that is, the position of the point in the first target shot segment = point time - shot start point.
[0098] For example, the starting point of a certain song is 13:34.09. According to the point time, the point can be automatically associated with the corresponding first target sub-shot segment, which is assumed to be ((13:33.07, 13:36.08), 125 frames, <intermediate frame feature vector>). The position of the point in the first target sub-shot segment is calculated as follows: position = point time - sub-shot starting point. In this example, the position of the point in the associated first target sub-shot segment is 13:34.09-13:33.07 = 00:01.02.
[0099] To further optimize the above embodiment, step S105 can specifically include:
[0100] (1) Based on the total number of frames contained in each second sub-shot feature vector, find each candidate sub-shot segment with the same total number of frames as the target sub-shot from the plurality of second sub-shot segments.
[0101] Wherein, the total number of frames of the target sub-shot is the total number of frames contained in the first sub-shot feature vector corresponding to the first target sub-shot segment.
[0102] In actual application, in the case that the total duration and sub-shot information of the target video and the original video are inconsistent, the second sub-shot feature vectors corresponding to each second sub-shot segment in the target video are traversed to find each candidate sub-shot segment with the same total number of frames as the first target sub-shot. The sub-shot information can include:
[0103] Specifically, each first target sub-shot segment with a marked point in the original video is denoted as Scene. For each Scene in Scenes in the original video, find all candidate sub-shot segments Same_frame_nums_scenes with Scene_frame_nums frames from each second sub-shot segment of the target video, where Scene_frame_nums is the total number of frames corresponding to the first target sub-shot segment.
[0104] (2) Calculate the vector distance between the target intermediate frame feature vector and each candidate intermediate frame feature vector.
[0105] Wherein, the target intermediate frame feature vector is the intermediate frame feature vector contained in the first sub-shot feature vector corresponding to the first target sub-shot segment. Each candidate intermediate frame feature vector is the intermediate frame feature vector contained in the second sub-shot feature vector corresponding to the candidate sub-shot segment.
[0106] Specifically, the vector distance between the <target intermediate frame feature vector> and each <intermediate frame feature vector> in Same_frame_nums_scenes is calculated. The vector distance between two intermediate frame features with the same sub-shot feature is usually 0.
[0107] (3) determining a target candidate sub-shot corresponding to a target vector distance satisfying a preset distance condition in all the vector distances as the second target sub-shot.
[0108] Specifically, a minimum vector distance is found from all the vector distances;
[0109] In a case where the minimum vector distance is not greater than a preset distance threshold, the minimum vector distance is determined as the target vector distance;
[0110] The target vector distance corresponding target candidate sub-shot is determined as the second target sub-shot.
[0111] The second target sub-shot can be marked as Scene_new.
[0112] The second target sub-shot in the embodiment is a second sub-shot in each second sub-shot of the target video, which needs to be marked with a point position.
[0113] The preset distance threshold is determined according to actual needs, for example, 0.1, which is not limited in the present application.
[0114] It should be noted that when the total time length and the sub-shot information (including: the number of sub-shots, the start and end time points of sub-shots) of the target video after replacement are consistent with those of the original video, it indicates that the target video is consistent with the original video, and at this time, all point position information in the original video can be automatically synchronized to the target video after replacement.
[0115] To further optimize the above embodiment, step S106 can specifically include:
[0116] (1) determining a point position offset based on the start point corresponding to the second target sub-shot and the start point corresponding to the first target sub-shot.
[0117] The point position offset Scene_bias = start point of Scene_new - start point of Scene.
[0118] Scene_new represents the second target sub-shot, and Scene represents the first target sub-shot.
[0119] (2) calculating a sum of the point position time point of the point position in the first target sub-shot and the point position offset to obtain a target point position time point of the point position in the second target sub-shot.
[0120] In practical applications, all point information corresponding to the first target sub-shot segment is acquired and marked as Scene_points, each point is marked as Scene_point by traversing Scene_points, and the position of Scene_point in Scene_new is calculated, as follows:
[0121] The point time point of Scene_point_new = the point time point of Scene_point + Scene_bias.
[0122] (3) The position of the target point time point in the second target sub-shot segment is taken as the dotting position, and the dotting operation is performed.
[0123] To further optimize the above embodiment, after step S104, the video dotting method can further include:
[0124] (1) In the case that the second target sub-shot segment is not found from the plurality of second sub-shot segments, the first same sub-shot segment corresponding to the front adjacent sub-shot segment of the first target sub-shot segment in the target video is determined, and the second same sub-shot segment corresponding to the rear adjacent sub-shot segment of the first target sub-shot segment in the target video is determined.
[0125] (2) The sub-shot segment between the first same sub-shot segment and the second same sub-shot segment is determined as a minimum artificial dotting video segment interval.
[0126] (3) The minimum artificial dotting video segment interval is output.
[0127] If Scene does not find a matching Scene_new, the same sub-shot segment needs to be found forward and backward, so as to determine a minimum artificial dotting video segment interval for artificial re-dotted in the interval, and the above process is repeated from Scene forward and backward until the first same sub-shot segment corresponding to the front adjacent sub-shot segment in the target video and the second same sub-shot segment corresponding to the rear adjacent sub-shot segment in the target video are found. The sub-shot segment between the first same sub-shot segment and the second same sub-shot segment is marked as a minimum artificial dotting video segment interval for re-dotted in Scene, which is marked as Scene_checkarea, and Scene_checkarea is displayed on the front-end web system for the operation personnel to check the minimum suspicious segment and re-dotted or give up dotting in the minimum suspicious segment. After clicking synchronization, the point synchronization of the target video is completed.
[0128] Corresponding to the above method embodiment, the application further discloses a video dotting device.
[0129] Reference is made toFigure 2 The structural diagram of the video dotting device is disclosed in the embodiment of the application, and the device can include:
[0130] The first split shot processing unit 201 is configured to perform split shot processing on the original video to obtain a plurality of first split shot segments, and determine a first split shot feature vector corresponding to each first split shot segment.
[0131] In the embodiment, the original video can be a media video initially stored in a media system, and specifically can be a media long video.
[0132] In the embodiment, the original video is divided into split shot level video units by performing split shot processing on the original video, and a plurality of first split shot segments are obtained.
[0133] In actual application, the original video can be processed by using the PySceneDetect library to obtain a plurality of first split shot segments. PySceneDetect is a powerful video scene detection library based on Python and OpenCV, which can detect scene changes in a video, including shot switching, fade-in and fade-out, and automatically segment the video into a plurality of independent scenes.
[0134] In the embodiment, the first split shot feature vector corresponding to each first split shot segment is a triple, which includes split shot start and end time, total number of split shot frames, and intermediate frame feature vector. This way can represent the video at minimum cost without the need to calculate and store a large number of feature vectors.
[0135] The dotting unit 202 is configured to perform operation dotting on the original video to obtain each dot position.
[0136] The operation dotting on the original video usually refers to marking the highlight, highlight, star appearance time point, advertisement insertion point, song start and end point in a variety of, and link start and end point in a variety of.
[0137] After the operation dotting on the original video is completed, the triple information (dot position title, dot position type, and dot position time point) corresponding to each dot position can be obtained by browsing the video marking time point position and the type of operation dot position.
[0138] The operation dotting on the original video can specifically include media operation, member operation, advertisement operation, and other dotting operations that need to be performed on the original video (such as media long video) initially stored. By browsing the video marking time point position and the type of operation dot position, it can be simplified to a triple information (dot position title, dot position type, and dot position time point).
[0139] The dot position association unit 203 is configured to associate each dot position to a corresponding first target split shot segment, and determine the position of the dot position in the first target split shot segment.
[0140] In actual application, according to the point time point corresponding to each point, a first sub-lens segment containing the point time point is determined, denoted as a first target sub-lens segment, so as to associate the point with the corresponding first target sub-lens segment. The position of the point in the first target sub-lens segment can be determined according to the time length of the point time point from the sub-lens starting point of the first target sub-lens segment.
[0141] It should be noted that the first target sub-lens segment in this embodiment is actually the first sub-lens segment marked with the point.
[0142] The second sub-lens processing unit 204 is configured to perform sub-lens processing on the target video obtained by replacing the original video to obtain a plurality of second sub-lens segments, and determine a second sub-lens feature vector corresponding to each second sub-lens segment.
[0143] Similarly, in actual application, the target video can be processed by using the PySceneDetect library to obtain a plurality of second sub-lens segments.
[0144] In this embodiment, the second sub-lens feature vector corresponding to each second sub-lens segment is a three-tuple, which includes: sub-lens start and end time, total number of sub-lens frames, and intermediate frame feature vector. This way can represent the video at minimum cost, without the need to calculate and store a large number of feature vectors.
[0145] The matching unit 205 is configured to find, from the plurality of second sub-lens segments, a second target sub-lens segment in which the second sub-lens feature vector matches the first sub-lens feature vector corresponding to the first target sub-lens segment.
[0146] In this embodiment, by matching the first sub-lens feature vector corresponding to the first target sub-lens segment with the second sub-lens feature vector corresponding to each second sub-lens segment, a second target sub-lens segment matching the first target sub-lens segment is found.
[0147] The dot position determination unit 206 is configured to determine the dot position of the point in the first target sub-lens segment in the second target sub-lens segment, and perform a dotting operation.
[0148] In actual application, the point position offset of the point in the second target sub-lens segment relative to the point in the first target sub-lens segment can be determined first, and then the new dot position of the point in the second target sub-lens segment is determined according to the point position offset, and the dotting operation is performed, so as to realize full-automatic point synchronization between the target video and the original video in the replacement scene.
[0149] In summary, the application discloses a kind of video dotting device, and the original video is treated to obtain a plurality of first mirror segments, and the first mirror feature vector corresponding to each first mirror segment is determined, and the original video is operated to obtain each point, and each point is associated into corresponding first target mirror segment, and the position of point in first target mirror segment is determined, and the original video is replaced target video is treated to obtain a plurality of second mirror segments, and the second mirror feature vector corresponding to each second mirror segment is determined, and the second target mirror segment that second mirror feature vector and the first mirror feature vector corresponding to first target mirror segment are matched is found from a plurality of second mirror segments, and the dotting position of point in first target mirror segment in second target mirror segment is determined, and dotting operation is executed.The first target mirror segment of marking point in original video is matched with each second mirror segment in target video, and the second target mirror segment that matches first target mirror segment is found, and the latest dotting position of point in first target mirror segment in second target mirror segment is determined to dot, realize that the target video of original video replacement is automatically dotted, so as to not only reduce the work input brought by repeated dotting, but also shorten the time difference caused by repeated dotting, guarantee the timeliness of instant content acquisition and video play volume of user, and then improve user experience.
[0150] To further optimize the above embodiment, the first mirror processing unit 201 can be specifically used for:
[0151] The original video is cut into multiple first mirror segments using a mirror algorithm.
[0152] The mirror start and end time of each first mirror segment in the original video is determined.
[0153] The total frame number of each first mirror segment is determined.
[0154] The mirror intermediate frame is extracted from each first mirror segment, and the intermediate frame feature vector is obtained by processing the mirror intermediate frame using a deep learning pre-training model.
[0155] The mirror start and end time, the total frame number and the intermediate frame feature vector are combined into a triple, and the triple is determined as the first mirror feature vector.
[0156] In the embodiment, the shot segmentation algorithm, i.e., video shot segmentation technology, mainly segments a video into multiple shot segments according to video pictures. The basic idea of video shot segmentation is to compare the similarity of each frame of a shot picture. If the similarity of the picture is high, it belongs to the same shot; if the similarity is very low, it belongs to the next shot. This technology can be used to cut long video content into shot granularity segments.
[0157] The deep learning pre-training model used when processing the shot intermediate frame can be a Resnet (Residual Network), VGGnet, EfficientNet, MocoV3, etc.
[0158] The deep learning pre-training model in the embodiment preferably uses a MocoV3 pre-training model. The MocoV3 pre-training model is an open source and widely used deep image pre-training model.
[0159] To further optimize the above embodiment, the point association unit 203 can be specifically configured to:
[0160] According to the point time point corresponding to each point, the point is associated with the corresponding first target shot segment;
[0161] According to the time length of the point time point from the shot start point of the first target shot segment, the position of the point in the first target shot segment is determined.
[0162] In the embodiment, after the original video is operated and dotted to obtain each point, the point type and point time point corresponding to each point are recorded. Since each first shot feature vector contains shot start and end time, by judging the shot start and end time to which the point time point belongs, the first shot segment associated with the point can be determined. For convenience of distinction, the first shot segment containing the point is recorded as the first target shot segment.
[0163] In actual application, when determining the position of the point time point in the first target shot segment, the position of the point time point from the first target shot segment is usually the position of the point in the first target shot segment, i.e., the position of the point in the first target shot segment = point time point - shot start point.
[0164] To further optimize the above embodiment, the matching unit 205 can be specifically configured to:
[0165] find each candidate shot segment with the same target total shot frame number from the plurality of second shot segments based on total shot frame numbers contained in each second shot feature vector, wherein the target total shot frame number is a total shot frame number contained in the first shot feature vector corresponding to the first target shot segment;
[0166] calculate vector distances between a target intermediate frame feature vector and each candidate intermediate frame feature vector, wherein the target intermediate frame feature vector is an intermediate frame feature vector contained in the first shot feature vector corresponding to the first target shot segment, and each candidate intermediate frame feature vector is an intermediate frame feature vector contained in the second shot feature vector corresponding to the candidate shot segment;
[0167] determine a target candidate shot segment corresponding to a target vector distance satisfying a preset distance condition from all the vector distances as the second target shot segment.
[0168] To further optimize the above embodiment, the matching unit 205 can be specifically used for:
[0169] find a minimum vector distance from all the vector distances;
[0170] determine the minimum vector distance as the target vector distance in a case where the minimum vector distance is not greater than a preset distance threshold;
[0171] determine a target candidate shot segment corresponding to the target vector distance as the second target shot segment.
[0172] To further optimize the above embodiment, the dotting position determination unit 206 can be specifically used for:
[0173] determine a dot position offset based on a starting point corresponding to the second target shot segment and a starting point corresponding to the first target shot segment;
[0174] calculate a sum of a dot position time point of the dot position in the first target shot segment and the dot position offset to obtain a target dot position time point of the dot position in the second target shot segment;
[0175] take a position of the target dot position time point in the second target shot segment as the dotting position, and perform a dotting operation.
[0176] To further optimize the above embodiment, the video dotting device can further include:
[0177] An adjacent sub-clip determining unit is configured to, in a case where the second target sub-clip is not found from the plurality of second sub-clips, determine a first same sub-clip corresponding to a front adjacent sub-clip of the first target sub-clip in the target video and a second same sub-clip corresponding to a rear adjacent sub-clip of the first target sub-clip in the target video.
[0178] A segment interval determining unit is configured to determine a sub-clip between the first same sub-clip and the second same sub-clip as a minimum artificial time-stamping video segment interval.
[0179] An output unit is configured to output the minimum artificial time-stamping video segment interval.
[0180] It should be noted that the specific working principles of the components in the device embodiment are described in the corresponding parts of the method embodiment, which will not be described here.
[0181] Corresponding to the above-mentioned embodiments, the application further discloses a computer storage medium, which stores at least one instruction, and the at least one instruction is executed by a processor to realize the steps shown in the video time-stamping method embodiment.
[0182] The computer storage medium can be a tangible medium, which can contain or store a program for use by or in connection with an instruction execution system, apparatus or device. The computer storage medium can be a machine-readable signal medium or a machine-readable storage medium. The computer storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the machine-readable storage medium include one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0183] Corresponding to the above-mentioned embodiments, the application further discloses an electronic device, as shown in the figure. Figure 3 The electronic device can include a processor 1 and a memory 2.
[0184] The processor 1 and the memory 2 can complete mutual communication through a communication bus 3.
[0185] The processor 1 is configured to execute at least one instruction.
[0186] The memory 2 is configured to store at least one instruction.
[0187] The processor 1 can be a central processing unit CPU, or an application specific integrated circuit ASIC, or one or more integrated circuits configured to carry out embodiments of the application.
[0188] The memory 2 can comprise a high speed RAM memory, and possibly also a non-volatile memory, such as at least one disk memory.
[0189] The processor executes at least one instruction to implement the steps of the video marking method embodiment.
[0190] Finally, it has to be noted that the terms like first and second etc., in the present text are used only to distinguish one entity or operation from another entity or operation, without necessarily requiring or implying any such actual relationship or order between these entities or operations. Also, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.
[0191] The various embodiments described in this specification are presented by way of example, and the details of the embodiments are not intended to limit the scope of the application. The various embodiments can be implemented in hardware, software, firmware, or any combination thereof.
[0192] The foregoing description of the disclosed embodiments enables a person skilled in the art to implement or use the application. Numerous modifications to these embodiments will be readily apparent to those skilled in the art, and the generic principles defined herein can be applied to other embodiments without departing from the spirit or scope of the application. Thus, the present application is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A video timing method, characterized by, The method comprises the following steps: performing shot processing on the original video to obtain a plurality of first shot segments, and determining a first shot feature vector corresponding to each first shot segment; performing operation dotting on the original video to obtain each dot; associating each dot to the corresponding first target shot segment, and determining the position of the dot in the first target shot segment; performing shot processing on the target video obtained by replacing the original video to obtain a plurality of second shot segments, and determining a second shot feature vector corresponding to each second shot segment; finding a second target shot segment from a plurality of second shot segments, in which the second shot feature vector matches the first shot feature vector corresponding to the first target shot segment; determining the dotting position of the dot in the first target shot segment in the second target shot segment, and performing dotting operation.
2. The video punching method according to claim 1, wherein, The method of performing shot processing on the original video to obtain a plurality of first shot segments, and determining a first shot feature vector corresponding to each first shot segment, comprises the following steps: performing shot cutting on the original video by using a shot algorithm to obtain a plurality of first shot segments; determining the shot start and end time of each first shot segment in the original video; determining the total number of frames of each first shot segment; extracting a shot intermediate frame from each first shot segment, and processing the shot intermediate frame by using a deep learning pre-training model to obtain an intermediate frame feature vector; composing a triple consisting of the shot start and end time, the total number of frames and the intermediate frame feature vector, and determining the triple as the first shot feature vector.
3. The video punching method of claim 1, wherein, The method of associating each dot to the corresponding first target shot segment, and determining the position of the dot in the first target shot segment, comprises the following steps: associating each dot to the corresponding first target shot segment according to the dot time point corresponding to each dot; determining the position of the dot in the first target shot segment according to the time length of the dot time point from the shot start point of the first target shot segment.
4. The video timing method according to any one of claims 1 to 3, wherein, The method of finding a second target shot segment from a plurality of second shot segments, in which the second shot feature vector matches the first shot feature vector corresponding to the first target shot segment, comprises the following steps: based on the total number of frames contained in each second shot feature vector, finding each candidate shot segment with the same target total number of frames from a plurality of second shot segments, wherein the target total number of frames is the total number of frames contained in the first shot feature vector corresponding to the first target shot segment; calculating the vector distance between the target intermediate frame feature vector and each candidate intermediate frame feature vector, wherein the target intermediate frame feature vector is the intermediate frame feature vector contained in the first shot feature vector corresponding to the first target shot segment, and each candidate intermediate frame feature vector is the intermediate frame feature vector contained in the second shot feature vector corresponding to the candidate shot segment. The target candidate sub-mirror segment corresponding to the target vector distance satisfying the preset distance condition among all the vector distances is determined as the second target sub-mirror segment.
5. The video strobing method of claim 4, wherein, The second target sub-mirror segment is determined from the target vector distance satisfying the preset distance condition among all the vector distances. The minimum vector distance is found from all the vector distances. In a case that the minimum vector distance is not greater than a preset distance threshold, the minimum vector distance is determined as the target vector distance. The target candidate sub-mirror segment corresponding to the target vector distance is determined as the second target sub-mirror segment.
6. The video strobing method of claim 1, wherein, The point position in the second target sub-mirror segment is determined based on the starting point of the second target sub-mirror segment and the starting point of the first target sub-mirror segment. The sum of the point position time point of the point in the first target sub-mirror segment and the point position offset is calculated to obtain a target point position time point of the point in the second target sub-mirror segment. The position of the target point position time point in the second target sub-mirror segment is taken as the point position, and the point operation is performed. Further comprising:
7. The video strobing method of claim 1, wherein, In a case that the second target sub-mirror segment is not found from the plurality of second sub-mirror segments, a first same sub-mirror segment corresponding to a front adjacent sub-mirror segment of the first target sub-mirror segment in the target video and a second same sub-mirror segment corresponding to a rear adjacent sub-mirror segment of the first target sub-mirror segment in the target video are determined. The sub-mirror segment between the first same sub-mirror segment and the second same sub-mirror segment is determined as a minimum artificial point video segment interval. The minimum artificial point video segment interval is output. Comprising:
8. A video punching device, characterized by comprising: A first sub-mirror processing unit is configured to perform sub-mirror processing on an original video to obtain a plurality of first sub-mirror segments, and determine a first sub-mirror feature vector corresponding to each first sub-mirror segment. A point unit is configured to perform operation point on the original video to obtain each point. A point association unit is configured to associate each point to a corresponding first target sub-mirror segment, and determine a position of the point in the first target sub-mirror segment. A second sub-mirror processing unit is configured to perform sub-mirror processing on a target video obtained by replacing the original video to obtain a plurality of second sub-mirror segments, and determine a second sub-mirror feature vector corresponding to each second sub-mirror segment. A matching unit is configured to find a second target sub-mirror segment from the plurality of second sub-mirror segments, in which the second sub-mirror feature vector matches the first sub-mirror feature vector corresponding to the first target sub-mirror segment. A point position determination unit is configured to determine a point position of the point in the first target sub-mirror segment in the second target sub-mirror segment, and perform a point operation. The computer storage medium stores at least one instruction, which is executed by a processor to implement the video point method of any one of claims 1-7.
9. A computer storage medium, characterized in that The electronic device comprises a memory and a processor.
10. An electronic device, comprising: The computer storage medium stores at least one instruction, which is executed by a processor to implement the video point method of any one of claims 1-7. The electronic device comprises a memory and a processor. The memory is configured to store at least one instruction; The processor is configured to execute the at least one instruction to implement the video dotting method according to any one of claims 1-7.
Citation Information
Patent Citations
Method and device for detecting target video fragment in video and electronic equipment
CN108769731A
Method and device for determining label of target video, computing equipment and storage medium
CN112163122A