Video explosion point migration method and device, electronic equipment and computer storage medium
By receiving video pop-up migration requests, extracting target time-series frame groups based on the pop-up locations of the original video, and determining the pop-up locations of the replacement video based on the video frame sequence of the replacement video and the target time-series frame groups, this method solves the problems of low efficiency and low matching accuracy in traditional methods, and achieves efficient and accurate video pop-up migration.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
- Filing Date
- 2026-02-12
- Publication Date
- 2026-05-19
AI Technical Summary
Traditional methods for locating viral moments in videos are inefficient, have low matching accuracy, and cannot adapt to changes in video content, leading to repetitive manual tracking and soaring operating costs.
By receiving video pop-up migration requests, the target time frame group is extracted based on the pop-up position of the original video. Based on the video frame sequence of the replacement video and the target time frame group, the pop-up position of the replacement video is determined. A hash matching and time offset constraint algorithm, combined with fingerprint feature matching, is used to achieve accurate time mapping of the replacement video.
It improves the efficiency and accuracy of video pop-up migration, reduces operating costs, enhances user experience, and automates the pop-up migration process.
Smart Images

Figure CN122069376A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video technology, and in particular to a method, apparatus, electronic device, and computer storage medium for migrating video pop-ups. Background Technology
[0002] In the process of video content operation, a large number of videos need to be marked with exciting moments to improve user experience. When video content is updated or replaced, the original timestamps of exciting moments become invalid and need to be repositioned and marked.
[0003] Traditional methods for locating viral moments in videos often employ techniques such as manual re-marking, single-frame feature matching, or fixed timeline mapping. These methods suffer from low efficiency, low matching accuracy, and an inability to adapt to content changes. Especially in complex scenarios such as video replacement, content editing, and multi-version operation, traditional methods cannot effectively solve the problem of automatically migrating the timestamps of viral moments, leading to repetitive manual marking and a surge in operational costs. Summary of the Invention
[0004] In view of this, the present invention provides a method, apparatus, electronic device and computer storage medium for migrating video pop-ups, which effectively improves the efficiency and accuracy of video pop-up migration.
[0005] The first aspect of this invention provides a method for migrating video pop-ups, comprising:
[0006] Receive a video clip shift request; wherein, the video clip shift request includes information about the original video and information about the replacement video; the information about the original video includes the location of the clip in the original video;
[0007] Based on the location of the explosion point in the original video, the target time frame group is extracted from the original video;
[0008] Based on the video frame sequence of the replacement video and the target time frame group, the location of the explosion point in the replacement video is determined;
[0009] The target time frame group is moved to the pop-up position of the replacement video.
[0010] Optionally, determining the pop-up location of the replacement video based on the video frame sequence of the replacement video and the target time-series frame group includes:
[0011] For each replacement video frame in the video frame sequence of the replacement video, the replacement video frame is matched with each key frame in the target time frame group to obtain the first key frame matching result.
[0012] Based on all the matching results of the first keyframe and the time offset constraint, determine the best matching result for each keyframe.
[0013] The location of the explosion point in the replacement video is determined based on the best matching result corresponding to all the keyframes.
[0014] Optionally, determining the pop-up location of the replacement video based on the video frame sequence of the replacement video and the target time-series frame group includes:
[0015] For each key frame in the target temporal frame group, a candidate temporal frame group in the video frame sequence of the replacement video is determined according to the time point of the key frame.
[0016] For each candidate frame in each of the candidate time-series frame groups, the candidate frame is matched with the key frame to obtain the second key frame matching result;
[0017] Based on all the matching results of the second keyframe and the time offset constraint, determine the best matching result for each keyframe;
[0018] The location of the explosion point in the replacement video is determined based on the best matching result corresponding to all the keyframes.
[0019] Optionally, before determining the burst location of the replacement video based on the video frame sequence of the replacement video and the target temporal frame group, the method further includes:
[0020] The fingerprint features of the target time-series frame group are determined based on the duration of the target time-series frame group and the hash value of each frame in the target time-series frame group.
[0021] The fingerprint features of the target time frame group are matched with the fingerprint features in the feature library to obtain the fingerprint feature matching result;
[0022] If the fingerprint feature matching result indicates a successful match, then the burst point position corresponding to the fingerprint feature that successfully matches the target time frame group in the feature library is taken as the burst point position of the replacement video.
[0023] If the fingerprint feature matching result indicates that a match is unsuccessful, then the step of determining the burst location of the replacement video based on the video frame sequence of the replacement video and the target time frame group is executed.
[0024] Optionally, after determining the burst location of the replacement video based on the video frame sequence of the replacement video and the target time-series frame group, the method further includes:
[0025] Based on the duration of the time frame group corresponding to the explosion point of the replacement video and the hash value of each frame in the time frame group corresponding to the explosion point of the replacement video, the fingerprint features of the time frame group corresponding to the explosion point of the replacement video are determined.
[0026] The fingerprint features of the time-series frame group corresponding to the explosion point position of the replacement video are stored in the feature library.
[0027] A second aspect of the present invention provides a device for migrating video pop-ups, comprising:
[0028] A receiving unit is configured to receive a video pop-up migration request; wherein the video pop-up migration request includes information about the original video and information about the replacement video; the information about the original video includes the pop-up locations of the original video;
[0029] The target temporal frame group determination unit is used to extract the target temporal frame group from the original video based on the location of the burst point in the original video.
[0030] The explosion location determination unit is used to determine the explosion location of the replacement video based on the video frame sequence of the replacement video and the target time frame group;
[0031] The migration unit is used to migrate the target time frame group to the burst position of the replacement video.
[0032] Optionally, the explosion point location determination unit includes:
[0033] The first matching unit is used to match each replacement video frame in the video frame sequence of the replacement video with each key frame in the target time frame group to obtain the first key frame matching result.
[0034] The first constraint unit is used to determine the best matching result for each key frame based on all the matching results of the first key frame and the time offset constraint.
[0035] The first burst point location determination subunit is used to determine the burst point location of the replacement video based on the best matching result corresponding to all the keyframes.
[0036] Optionally, the explosion point location determination unit includes:
[0037] The candidate temporal frame group determination unit is used to determine, for each key frame in the target temporal frame group, a candidate temporal frame group in the video frame sequence of the replacement video according to the time point where the key frame is located.
[0038] The second matching unit is used to match each candidate frame in each candidate time frame group with the key frame to obtain the second key frame matching result.
[0039] The second constraint unit is used to determine the best matching result for each key frame based on all the matching results of the second key frame and the time offset constraint.
[0040] The second burst point location determination subunit is used to determine the burst point location of the replacement video based on the best matching result corresponding to all the keyframes.
[0041] Optionally, the video pop-up migration device further includes:
[0042] The first fingerprint feature determination unit is used to determine the fingerprint features of the target time frame group based on the duration of the target time frame group and the hash value of each frame in the target time frame group.
[0043] The third matching unit is used to match the fingerprint features of the target time frame group with the fingerprint features in the feature library to obtain the fingerprint feature matching result;
[0044] The burst location determination unit is further configured to, if the fingerprint feature matching result indicates a successful match, take the burst location corresponding to the fingerprint feature that successfully matches the target time frame group in the feature library as the burst location of the replacement video; if the fingerprint feature matching result indicates a failed match, determine the burst location of the replacement video based on the video frame sequence of the replacement video and the target time frame group.
[0045] Optionally, the video pop-up migration device further includes:
[0046] The second fingerprint feature determination unit is used to determine the fingerprint feature of the time sequence frame group corresponding to the explosion point position of the replacement video based on the duration of the time sequence frame group corresponding to the explosion point position of the replacement video and the hash value of each frame in the time sequence frame group corresponding to the explosion point position of the replacement video.
[0047] The storage unit is used to store the fingerprint features of the time-series frame group corresponding to the burst position of the replacement video into the feature library.
[0048] A third aspect of the present invention provides an electronic device, comprising:
[0049] One or more processors;
[0050] A storage device on which one or more programs are stored;
[0051] When the one or more programs are executed by the one or more processors, the one or more processors implement the video pop-up migration method as described in any one of the first aspects.
[0052] A fourth aspect of the present invention provides a computer storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the video burst point migration method as described in any one of the first aspects.
[0053] As can be seen from the above scheme, the present invention provides a method, apparatus, electronic device, and computer storage medium for migrating video pop-points. After receiving a video pop-point migration request, the method extracts a target temporal frame group from the original video based on the pop-point position. Then, based on the video frame sequence of the replacement video and the target temporal frame group, it achieves accurate time mapping under the replacement video, thereby determining the pop-point position of the replacement video, replacing the traditional manual calibration method. Finally, the target temporal frame group is migrated to the pop-point position of the replacement video. This effectively improves the migration efficiency and accuracy of video pop-points. Attached Figure Description
[0054] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0055] Figure 1 A flowchart illustrating a method for migrating video pop-up points provided in an embodiment of the present invention;
[0056] Figure 2 A flowchart illustrating a method for determining the location of a clip in a replacement video, as provided in another embodiment of the present invention;
[0057] Figure 3 A flowchart illustrating a method for determining the location of a clip in a replacement video, as provided in another embodiment of the present invention;
[0058] Figure 4 A flowchart illustrating a method for migrating video pop-ups, as provided in another embodiment of the present invention;
[0059] Figure 5 A schematic diagram of a video pop-up migration device provided in another embodiment of the present invention;
[0060] Figure 6 This is a schematic diagram of an electronic device for implementing a method for migrating video pop-ups, as provided in another embodiment of the present invention. Detailed Implementation
[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0062] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0063] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this invention are all information and data authorized by the user or fully authorized by all parties.
[0064] It should be noted that the concepts of "first" and "second" mentioned in this invention are only used to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0065] It should be noted that the terms "a" and "a plurality of" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0066] This invention provides a method for migrating video pop-up points, such as... Figure 1 As shown, the specific steps include:
[0067] S101, Receive video pop-up migration request.
[0068] The video clip migration request includes information about the original video and the replacement video; the original video information includes the location of the clip's explosion point. The explosion point location includes at least the start and end points of the explosion.
[0069] It should be noted that a replacement video refers to replacing the original video content with a new video, while retaining the original video's file identifier, metadata, or playback link. Replacement videos can be used in scenarios such as video editing, platform content updates, and copyright replacement, and are not limited to these here.
[0070] In the practical application of this invention, the original video can be obtained through the path information in the original video information, and the replacement video can be obtained through the path information in the replacement video information. Of course, the original video and the replacement video can also be directly carried in the video breakout request, which is not limited here.
[0071] S102. Extract the target time frame group from the original video based on the explosion point location.
[0072] A temporal frame group is a set of video frames arranged in chronological order to represent dynamic changes over a period of time. The target temporal frame group is a set of video frames corresponding to the burst points in the original video. A video frame refers to a single still image in a video and is the smallest unit of video content.
[0073] In the practical application of this invention, the target time frame group can be extracted from the original video based on the start and end points of the explosion points.
[0074] After extracting the target time frame group, it can be stored on disk, in memory, etc., and then migrated after the explosion point of the replacement video is confirmed.
[0075] S103. Based on the video frame sequence of the replacement video and the target time frame group, determine the explosion point location of the replacement video.
[0076] The replacement video video sequence includes multiple replacement video frames, and the target time frame group includes multiple target video frames.
[0077] Specifically, the location of the explosion point in the replacement video is determined by matching the replacement video frames in the replacement video frame sequence with the target video frames in the target time frame group.
[0078] In another embodiment of the present invention, multiple keyframes can be determined in the target time frame group, and the explosion point position of the replacement video can be determined by matching the replacement video frames in the video frame sequence of the replacement video with the keyframes in the target time frame group.
[0079] The keyframe settings can be predefined or directly incorporated into the original video information; there are no restrictions here.
[0080] The method for determining predefined keyframes can be to use certain points with obvious characteristics as keyframes, such as the video frames corresponding to the start point, high energy point, and end point of the explosion. There is no limitation here.
[0081] Optionally, in another embodiment of the present invention, one implementation of step S103 is as follows: Figure 2 As shown, it includes:
[0082] S201. For each replacement video frame in the video frame sequence of the replacement video, match the replacement video frame with each key frame in the target time frame group to obtain the first key frame matching result.
[0083] In the practical application of this invention, the specific matching method can be, but is not limited to, calculating the hash distance difference between the two, that is, calculating the hash of the replacement video frame and the hash of the key frame in the target time frame group, and then calculating the distance difference between the two hashes. This can be achieved using, but is not limited to, the Minghan distance; no limitation is made here. It is understood that the smaller the distance difference, the more similar the two are.
[0084] In practical applications of this invention, a backup algorithm can also be set. If the main algorithm fails to match, the backup algorithm is switched to effectively improve the overall processing success rate. For example, the main algorithm is the Perceptual Hash Algorithm (pHash), and the backup algorithm is the Structural Similarity Index Measure (SSIM).
[0085] Taking the video frame corresponding to the explosion point as an example, each replacement video frame in the replacement video's video frame sequence is matched with the key frame corresponding to the explosion point to obtain the first key frame matching result, that is, the matching result between the replacement video frame and the key frame corresponding to the explosion point.
[0086] The following is a diagram illustrating the matching results of the first keyframe:
[0087] frame=4301, phash_dist=28;
[0088] frame=4302, phash_dist=28;
[0089] frame=4303, phash_dist=28;
[0090] frame=4304, phash_dist=0;
[0091] frame=4305, phash_dist=0;
[0092] frame=4306, phash_dist=0;
[0093] frame=4307, phash_dist=0;
[0094] frame=4308, phash_dist=10;
[0095] frame=4309, phash_dist=10;
[0096] Where frame is the frame number, and phash_dist is the distance between the replacement video frame (the replacement video frame corresponding to frame) and the keyframe corresponding to the explosion start point.
[0097] S202. Based on the matching results of all first keyframes and the time offset constraint, determine the best matching result for each keyframe.
[0098] The time offset constraint is a preset maximum tolerable time difference, such as 0.5 seconds or 1 second, which is not limited here. Taking the explosion point position of the original video, which includes the start point, high energy point, and end point, as an example, the time offset constraint is to ensure that the time length between the high energy point and the start point, and between the end point and the high energy point, is within 0.5 seconds of the corresponding time difference in the original video.
[0099] Continuing with the example above, the first keyframe matching result contains frame and phash_dis. To calculate the time difference between them, we need to know the time point. For this purpose, we can obtain the time point of the current frame by dividing the frame number by the frame rate, but this is not limited to this.
[0100] In the practical application of this invention, the top N replacement video frames with the smallest distance from the key frame can be selected from the first key frame matching results as target replacement video frames.
[0101] The optimal matching result for each keyframe is determined based on the target replacement video frame corresponding to that keyframe and the time offset constraint.
[0102] Taking the original video's explosion point location as including the start point, high-energy point, and end point as an example, that is, the key frame is the video frame corresponding to the start point, the video frame corresponding to the high-energy point, and the video frame corresponding to the end point, then three groups of target replacement video frames will be obtained, called group A1, group A2, and group A3. The target replacement video frame that appears earlier in each group is selected first.
[0103] Assuming that the time difference between the first target replacement video frame A21 in group A2 and the first target replacement video frame A11 in group A1 satisfies the time offset constraint, and the time difference between the first target replacement video frame A31 in group A3 and the first target replacement video frame A21 in group A2 satisfies the time offset constraint, then, taking A11 as the best matching result corresponding to the starting point video frame, taking A21 as the best matching result corresponding to the high-energy point video frame, and taking A31 as the best matching result corresponding to the ending point video frame.
[0104] Assuming the time difference between the first target replacement video frame A21 in group A2 and the first target replacement video frame A11 in group A1 satisfies the time offset constraint, and the time difference between the first target replacement video frame A31 in group A3 and the first target replacement video frame A21 in group A2 does not satisfy the time offset constraint, then continue querying the next target replacement video frame A32 in group A3 to see if its time difference with the first target replacement video frame A21 in group A2 satisfies the time offset constraint, until a target replacement video frame in group A3 that satisfies the time offset constraint with the first target replacement video frame A21 in group A2 is found. If no target replacement video frame in group A3 satisfies the time offset constraint with the first target replacement video frame A21 in group A2, then... If the time difference between the first target replacement video frame A21 and the target replacement video frame A22 in group A2 satisfies the time offset constraint, then the next target replacement video frame A22 in group A2 can be selected. At this point, it is necessary to re-evaluate whether the time difference between A22 and A11 satisfies the time offset constraint. If it does, then perform time offset constraint analysis between A22 and the target replacement video frames in group A3. If not, then continue searching for the next target replacement video frame in group A1, and so on, until the target replacement video frame A1X in group A1 satisfies the time offset constraint with the target replacement video frame A2Y in group A2, and the target replacement video frame A2Y in group A2 satisfies the time offset constraint with the target replacement video frame A3Z in group A3. Then, A1X is taken as the best matching result for the starting point video frame, A2Y as the best matching result for the high-energy point video frame, and A3Z as the best matching result for the ending point video frame.
[0105] In practical applications of this invention, the time difference between the best matching results corresponding to keyframes can also be used to judge the matching quality. Taking a time offset constraint of 0.5 seconds as an example, a deviation within 0.5 seconds can be considered a high-quality match, 1 second is considered medium-quality, greater than 1 second but less than 5 seconds is considered a low-quality match, and a deviation greater than 5 seconds is completely unusable. In current practical use, a matching success rate of 99.9% can be achieved within 0.5 seconds.
[0106] S203. Based on the best matching results corresponding to all keyframes, determine the location of the explosion point in the replacement video.
[0107] Continuing with the above examples, the trigger points in the replacement video include A1X, A2Y, and A3Z.
[0108] To improve matching speed, in another embodiment of the present invention, one implementation of step S103 is as follows: Figure 3 As shown, it includes:
[0109] S301. For each key frame in the target time frame group, determine the candidate time frame group in the video frame sequence of the replacement video according to the time point of the key frame.
[0110] Specifically, based on the time point of the key frame in the target time frame group of the original video, N frames before and after the corresponding time point in the video frame sequence of the replacement video are selected as candidate time frame groups.
[0111] For example, if the keyframe in the target time frame group of the original video is at 20 minutes and 18 seconds, then the N frames before and after the 20 minutes and 18 seconds position in the replacement video's video frame sequence are considered as candidate time frame groups.
[0112] S302. For each candidate frame in each candidate time frame group, match the candidate frame with the key frame to obtain the second key frame matching result.
[0113] S303. Based on the matching results of all second keyframes and the time offset constraint, determine the best matching result for each keyframe.
[0114] S304. Based on the best matching results corresponding to all keyframes, determine the location of the explosion point in the replacement video.
[0115] It should be noted that the specific implementation methods of steps S302 to S304 can refer to the implementation methods of steps S201 to S203, and will not be repeated here.
[0116] S104. Move the target time frame group to the burst point position of the replacement video.
[0117] In the practical application of this invention, the user or AI can send the original video clip data and the address of the new video after replacement to the clip migration engine. The calling interface can be a variety of communication modes such as RESTful API, gRPC, and GraphQL, which are not limited here.
[0118] As can be seen from the above scheme, the present invention provides a method for migrating video pop-points. After receiving a video pop-point migration request, the method determines a target temporal frame group in the original video based on the pop-point position. Then, based on the video frame sequence of the replacement video and the target temporal frame group, it achieves accurate time mapping under the replacement video, thereby determining the pop-point position of the replacement video, replacing the traditional manual calibration method. Finally, the target temporal frame group is migrated to the pop-point position of the replacement video. This effectively improves the efficiency and accuracy of video pop-point migration.
[0119] Another embodiment of the present invention provides a method for migrating video pop-ups, such as... Figure 4 As shown, it specifically includes:
[0120] S401, Receive video pop-up migration request.
[0121] The video clip migration request includes information about the original video and the replacement video; the information about the original video includes the location of the clip's clip.
[0122] S402. Extract the target time frame group from the original video based on the explosion point location.
[0123] It should be noted that the specific implementation methods of steps S401 to S402 can be referred to the implementation methods of steps S101 to S102, and will not be repeated here.
[0124] S403. Determine the fingerprint features of the target time frame group based on the duration of the target time frame group and the hash value of each frame in the target time frame group.
[0125] Specifically, in the practical application of this invention, the fingerprint features of the target temporal frame group can be calculated using, but is not limited to, the following calculation formula. :
[0126] ;
[0127] in, Represents the first in the time sequence frame group The perceptual hash value of the frame; The total number of frames in the target time frame group; Representative bestowed upon the first Frame weights; This means that the weighted hash of all frames is normalized to obtain a stable aggregate hash value, which can improve the robustness of the system. Represents the first frame from the target time frame group To the last frame The total time span, which is the duration of the climax.
[0128] It should be noted that a certain tolerance range can also be set for the time interval parameter. For example, a timing deviation of ±2 frames is considered valid during matching, which can significantly improve the matching success rate in video editing scenarios.
[0129] In the practical application of this invention, a feature library can be constructed to store the fingerprint features of the target temporal frame group. The data format for storage can be as follows:
[0130] {
[0131] "fileId": "65b7bbc8898f80bdff67ec76c0970e4e",
[0132] "frame": 1,
[0133] "feature":"70968f80be4dbc885b70f96cbf67ec7e",
[0134] "images":[
[0135] {
[0136] "frame_num": 1,
[0137] "frame": "frame_1.png",
[0138] },{
[0139] "frame_num": 528,
[0140] "frame": "frame_2.png",
[0141] },{
[0142] "frame_num": 1628,
[0143] "frame": "frame_3.png",
[0144] } ]
[0146] }
[0147] Wherein, fileId is the globally unique ID of the video file, frame is the frame number of the burst timestamp frame or scene frame in the video, feature is the fingerprint of the corresponding frame, and image is the storage address of the video frame.
[0148] S404. Match the fingerprint features of the target time frame group with the fingerprint features in the feature library to obtain the fingerprint feature matching result.
[0149] In the practical application of this invention, the fingerprint features of the target time frame group can be compared one by one with the fingerprint features in the feature library. If the fingerprint features of the target time frame group are consistent with the fingerprint features in the feature library, it means that the fingerprint feature matching is successful. If the fingerprint features of the target time frame group are inconsistent with the fingerprint features in the feature library, it means that the fingerprint feature matching is unsuccessful.
[0150] S405. If the fingerprint feature matching result indicates a successful match, the burst point position corresponding to the fingerprint feature that successfully matches the target time frame group in the feature library shall be used as the burst point position of the replacement video.
[0151] S406. If the fingerprint feature matching result shows that the match is unsuccessful, the location of the explosion point in the replacement video is determined based on the video frame sequence of the replacement video and the target time frame group.
[0152] S407. Move the target time frame group to the burst position of the replacement video.
[0153] It should be noted that the specific implementation methods of steps S406 to S407 can refer to the implementation methods of steps S103 to S104, and will not be repeated here.
[0154] Optionally, in another embodiment of the present invention, after determining the location of the burst point in the replacement video based on the video frame sequence and the target temporal frame group, one implementation of the video burst point migration method further includes:
[0155] Based on the duration of the time frame group corresponding to the explosion point of the replacement video and the hash value of each frame in the time frame group corresponding to the explosion point of the replacement video, the fingerprint features of the time frame group corresponding to the explosion point of the replacement video are determined; the fingerprint features of the time frame group corresponding to the explosion point of the replacement video are stored in the feature database.
[0156] In the specific implementation of this invention, the implementation method for determining fingerprint features can be referred to the implementation method of step S403, which will not be repeated here.
[0157] In practical applications of this invention, the aforementioned video highlight migration method can be deployed within an engine (such as the VideoMarkerMigrationEngine) and integrated into the video processing pipeline. The traditional processing path, such as: replacement video → manual re-marking → publishing, is changed in this solution to: replacement video → VideoMarkerMigrationEngine → automatic highlight migration → publishing. Thus, the entire system significantly reduces operating costs and manpower input while improving video content update efficiency without affecting highlight accuracy.
[0158] The parameters and descriptions required for the engine can be found in Table 1:
[0159] parameter illustrate --src_video The original video file path, such as: / data / video / a.mp4 --dst_video The path to the replaced video file, such as: / data / video / b.mp4 --point_time A list of timestamps for breaking news, such as: [10.04, 20.88, 30.68] --threshold Similarity threshold (default 0.9) controls matching accuracy. --max_diff_duration Maximum tolerable time difference, e.g., 0.5s --resolution Use a uniform resolution to avoid differences in video scaling, such as 848*480. --output Post-migration burst timestamp output file address
[0160] Table 1
[0161] Namely: VideoMarkerMigrationEngine:
[0162] --src_video / data / video / a.mp4;
[0163] --dst_video / data / video / b.mp4 --point_time '[10.04,20.88,30.68]';
[0164] --threshold 0.9;
[0165] --max_diff_duration 0.5;
[0166] --resolution 848*480;
[0167] --output / output / markers.json;
[0168] This tool supports batch processing, taking the video and pop data passed by --src_video, dynamically matching parameters, and outputting the optimal migration result to the address specified by -output.
[0169] In practical applications of this invention, data such as matching similarity, processing time, and offset during the migration process can be reported to the analysis system, facilitating subsequent algorithm optimization (automatic adjustment of matching parameters). This implements a self-evolutionary update mechanism for the burst point migration strategy, enabling the migration accuracy to continuously improve with accumulated processing experience.
[0170] By employing the video pop-point migration method proposed in this invention, the following practical benefits are achieved while maintaining the pop-point positioning accuracy (time error ≤ time offset constraint (e.g., 0.5 seconds) and matching accuracy ≥ 90%):
[0171] 1) The efficiency of hotspot migration is improved by 288 times, significantly reducing labor costs;
[0172] After verification in a large-scale production environment, the explosion point migration system of this invention reduces the processing time of a single video from 24 hours of traditional manual processing to less than 5 minutes, improving efficiency by more than 288 times.
[0173] For 1-hour video content, the cost of migrating viral content has been reduced from 50 yuan per video to 0.1 yuan per video. Based on the platform processing an average of 20 video replacements per day, this can save more than 365,000 yuan in operating costs annually.
[0174] 2) The accuracy rate of scene replacement and highlight transfer reaches 99.9%, ensuring user experience;
[0175] The temporal frame group matching algorithm effectively solves the ambiguity problem of single frame matching, and improves the accuracy of burst point migration to 99.9% in various replacement scenarios such as video content deletion, editing, and reordering.
[0176] The bounce rate of mobile users decreased due to the failure of the content's impact, and the average viewing time of users increased, significantly improving the content consumption experience.
[0177] 3) Seamless integration with existing video processing pipelines, enabling quick and easy deployment;
[0178] The migration engine is deployed as a microservice, eliminating the need to modify existing video production and content management systems.
[0179] It supports API interface calls and batch processing mode. By simply adding a migration service module to the video processing pipeline, the automatic migration of the highlights in the replacement scene can be realized, with an integration cycle of less than 3 person-days.
[0180] 4) Intelligent learning and evolution mechanism to continuously optimize transfer performance;
[0181] The system has the ability to provide feedback on migration results and learn from them. By analyzing historical migration data, it can automatically optimize matching parameters and offset algorithms.
[0182] 5) Multi-algorithm fusion ensures a 99.9% success rate;
[0183] The system employs a dual-algorithm fusion architecture of SSIM and pHash, automatically switching to a backup algorithm when the main algorithm fails to match, achieving an overall system success rate of 99.9%.
[0184] It supports 4K / 8K ultra-high-definition video processing, and in the case of an RTX 3090 graphics card, the processing speed can reach 200 frames per second, meeting the needs of large-scale production.
[0185] As shown above, this invention provides a method for migrating video pop-up points. Based on the duration of the target time-series frame group and the hash value of each frame within the target time-series frame group, the fingerprint features of the target time-series frame group are determined. Then, the fingerprint features of the target time-series frame group are matched with fingerprint features in a feature library to obtain fingerprint feature matching results. If the fingerprint feature matching result indicates a successful match, the pop-up point position corresponding to the fingerprint feature in the feature library that successfully matches the target time-series frame group is used as the pop-up point position in the replacement video. This further improves the efficiency of video pop-up point migration.
[0186] Another embodiment of the present invention provides a device for migrating video pop-ups, such as... Figure 5 As shown, it specifically includes:
[0187] The receiving unit 501 is used to receive video pop-up migration requests.
[0188] The video clip migration request includes information about the original video and the replacement video; the information about the original video includes the location of the clip's clip.
[0189] The target temporal frame group determination unit 502 is used to extract the target temporal frame group from the original video based on the burst point position of the original video.
[0190] The explosion location determination unit 503 is used to determine the explosion location of the replacement video based on the video frame sequence of the replacement video and the target time frame group.
[0191] Optionally, in another embodiment of the present invention, one implementation of the detonation point location determination unit 503 includes:
[0192] The first matching unit is used to match each replacement video frame in the replacement video frame sequence with each key frame in the target time frame group to obtain the first key frame matching result.
[0193] The first constraint unit is used to determine the best matching result for each keyframe based on all the matching results of the first keyframe and the time offset constraint.
[0194] The first burst point location determination subunit is used to determine the burst point location of the replacement video based on the best matching result corresponding to all keyframes.
[0195] For details on the specific operation of the units disclosed in the above embodiments of the present invention, please refer to the corresponding method embodiments, such as... Figure 2 As shown, it will not be elaborated further here.
[0196] Optionally, in another embodiment of the present invention, one implementation of the detonation point location determination unit 503 includes:
[0197] The candidate temporal frame group determination unit is used to determine the candidate temporal frame group in the video frame sequence of the replacement video based on the time point of the key frame in each key frame in the target temporal frame group.
[0198] The second matching unit is used to match each candidate frame with a key frame for each candidate frame in each candidate time frame group to obtain the second key frame matching result.
[0199] The second constraint unit is used to determine the best matching result for each keyframe based on all the matching results of the second keyframe and the time offset constraint.
[0200] The second burst point location determination subunit is used to determine the burst point location of the replacement video based on the best matching result corresponding to all keyframes.
[0201] For details on the specific operation of the units disclosed in the above embodiments of the present invention, please refer to the corresponding method embodiments, such as... Figure 3 As shown, it will not be elaborated further here.
[0202] The migration unit 504 is used to migrate the target time frame group to the burst position of the replacement video.
[0203] For details on the specific operation of the units disclosed in the above embodiments of the present invention, please refer to the corresponding method embodiments, such as... Figure 1 As shown, it will not be elaborated further here.
[0204] Optionally, in another embodiment of the present invention, one implementation of the video pop-up migration device further includes:
[0205] The first fingerprint feature determination unit is used to determine the fingerprint features of the target time frame group based on the duration of the target time frame group and the hash value of each frame in the target time frame group.
[0206] The third matching unit is used to match the fingerprint features of the target time frame group with the fingerprint features in the feature library to obtain the fingerprint feature matching result.
[0207] The burst location determination unit is also used to determine the burst location of the replacement video if the fingerprint feature matching result is successful, by taking the burst location corresponding to the fingerprint feature that successfully matches the target time frame group in the feature library as the burst location of the replacement video; if the fingerprint feature matching result is unsuccessful, the burst location of the replacement video is determined based on the video frame sequence of the replacement video and the target time frame group.
[0208] For details on the specific operation of the units disclosed in the above embodiments of the present invention, please refer to the corresponding method embodiments, such as... Figure 4 As shown, it will not be elaborated further here.
[0209] Optionally, in another embodiment of the present invention, one implementation of the video pop-up migration device further includes:
[0210] The second fingerprint feature determination unit is used to determine the fingerprint features of the time sequence frame group corresponding to the explosion point of the replacement video based on the duration of the time sequence frame group corresponding to the explosion point of the replacement video and the hash value of each frame in the time sequence frame group corresponding to the explosion point of the replacement video.
[0211] The storage unit is used to store the fingerprint features of the time-series frame group corresponding to the explosion point position of the replacement video into the feature library.
[0212] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, and will not be repeated here.
[0213] As can be seen from the above scheme, the present invention provides a video pop-point migration device. After receiving a video pop-point migration request, the receiving unit 501 extracts the target time-series frame group from the original video based on the pop-point position of the original video. Then, the pop-point position determination unit 503 achieves accurate time mapping under the replacement video based on the video frame sequence of the replacement video and the target time-series frame group, thereby determining the pop-point position of the replacement video, replacing the traditional manual calibration method. Finally, the migration unit 504 migrates the target time-series frame group to the pop-point position of the replacement video. This effectively improves the migration efficiency and accuracy of video pop points.
[0214] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0215] Another embodiment of the present invention provides an electronic device, such as... Figure 6 As shown, it includes:
[0216] One or more processors 601.
[0217] Storage device 602, on which one or more programs are stored.
[0218] When the one or more programs are executed by the one or more processors 601, the one or more processors 601 implement the video pop-up migration method as described in the above embodiments.
[0219] Another embodiment of the present invention provides a computer storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the video burst point migration method as described in the above embodiments.
[0220] In the context of this invention, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable media can include, but are not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0221] It should be noted that the computer-readable medium described above in this invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this invention, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0222] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0223] Another embodiment of the present invention provides a computer program product, which, when executed, is used to perform the above-described video pop-up migration method.
[0224] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, it performs the functions defined in the methods of the embodiments of the present invention.
[0225] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in this invention is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely exemplary forms for implementing the invention.
[0226] While several specific implementation details are included in the foregoing discussion, these should not be construed as limiting the scope of the invention. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0227] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention is not limited to the specific combination of the above-described technical features, but also includes other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting the above-described features with technical features of the present invention (but not limited to) that have similar functions.
Claims
1. A method for migrating video pop-up points, characterized in that, include: Receive a video clip shift request; wherein, the video clip shift request includes information about the original video and information about the replacement video; the information about the original video includes the location of the clip in the original video; Based on the location of the explosion point in the original video, the target time frame group is extracted from the original video; Based on the video frame sequence of the replacement video and the target time frame group, the location of the explosion point in the replacement video is determined; The target time frame group is moved to the pop-up position of the replacement video.
2. The video pop-up migration method according to claim 1, characterized in that, Determining the pop-up location of the replacement video based on the video frame sequence of the replacement video and the target time-series frame group includes: For each replacement video frame in the video frame sequence of the replacement video, the replacement video frame is matched with each key frame in the target time frame group to obtain the first key frame matching result. Based on all the matching results of the first keyframe and the time offset constraint, determine the best matching result for each keyframe. The location of the explosion point in the replacement video is determined based on the best matching result corresponding to all the keyframes.
3. The video pop-up migration method according to claim 1, characterized in that, Determining the pop-up location of the replacement video based on the video frame sequence of the replacement video and the target time-series frame group includes: For each key frame in the target temporal frame group, a candidate temporal frame group in the video frame sequence of the replacement video is determined according to the time point of the key frame. For each candidate frame in each of the candidate time-series frame groups, the candidate frame is matched with the key frame to obtain the second key frame matching result; Based on all the matching results of the second keyframe and the time offset constraint, determine the best matching result for each keyframe; The location of the explosion point in the replacement video is determined based on the best matching result corresponding to all the keyframes.
4. The video pop-up migration method according to claim 1, characterized in that, Before determining the burst location of the replacement video based on the video frame sequence of the replacement video and the target temporal frame group, the method further includes: The fingerprint features of the target time-series frame group are determined based on the duration of the target time-series frame group and the hash value of each frame in the target time-series frame group. The fingerprint features of the target time frame group are matched with the fingerprint features in the feature library to obtain the fingerprint feature matching result; If the fingerprint feature matching result indicates a successful match, then the burst point position corresponding to the fingerprint feature that successfully matches the target time frame group in the feature library is taken as the burst point position of the replacement video. If the fingerprint feature matching result indicates that a match is unsuccessful, then the step of determining the burst location of the replacement video based on the video frame sequence of the replacement video and the target time frame group is executed.
5. The method for migrating video pop-up points according to claim 1, characterized in that, After determining the burst location of the replacement video based on the video frame sequence of the replacement video and the target time frame group, the method further includes: Based on the duration of the time frame group corresponding to the explosion point of the replacement video and the hash value of each frame in the time frame group corresponding to the explosion point of the replacement video, the fingerprint features of the time frame group corresponding to the explosion point of the replacement video are determined. The fingerprint features of the time-series frame group corresponding to the explosion point position of the replacement video are stored in the feature library.
6. A device for migrating video pop-up points, characterized in that, include: A receiving unit is configured to receive a video pop-up migration request; wherein the video pop-up migration request includes information about the original video and information about the replacement video; the information about the original video includes the pop-up locations of the original video; The target temporal frame group determination unit is used to extract the target temporal frame group from the original video based on the location of the burst point in the original video. The explosion location determination unit is used to determine the explosion location of the replacement video based on the video frame sequence of the replacement video and the target time frame group; The migration unit is used to migrate the target time frame group to the burst position of the replacement video.
7. The video pop-up migration device according to claim 6, characterized in that, The detonation point location determination unit includes: The first matching unit is used to match each replacement video frame in the video frame sequence of the replacement video with each key frame in the target time frame group to obtain the first key frame matching result. The first constraint unit is used to determine the best matching result for each key frame based on all the matching results of the first key frame and the time offset constraint. The first burst point location determination subunit is used to determine the burst point location of the replacement video based on the best matching result corresponding to all the keyframes.
8. The video pop-up migration device according to claim 6, characterized in that, The detonation point location determination unit includes: The candidate temporal frame group determination unit is used to determine, for each key frame in the target temporal frame group, a candidate temporal frame group in the video frame sequence of the replacement video according to the time point where the key frame is located. The second matching unit is used to match each candidate frame in each candidate time frame group with the key frame to obtain the second key frame matching result. The second constraint unit is used to determine the best matching result for each key frame based on all the matching results of the second key frame and the time offset constraint. The second burst point location determination subunit is used to determine the burst point location of the replacement video based on the best matching result corresponding to all the keyframes.
9. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the video burst migration method as described in any one of claims 1 to 5.
10. A computer storage medium, characterized in that, It stores a computer program, wherein the computer program, when executed by a processor, implements the video burst point migration method as described in any one of claims 1 to 5.