Automatic Highlight Editing Method for Movies Based on Deep Learning Audio Classification

By using a deep learning-based audio classification method, highlight clips in movies can be automatically edited, solving the problem of time-consuming movie editing and achieving efficient highlight clip generation and processing.

CN117612516BActive Publication Date: 2026-05-26HANGZHOU ARCVIDEO TECHNOLOGY CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HANGZHOU ARCVIDEO TECHNOLOGY CO LTD
Filing Date
2023-11-28
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Film editing requires a lot of manpower, time and energy. Selecting highlights is extremely time-consuming, and current technology is inefficient.

Method used

This paper adopts a deep learning-based audio classification method. By training a deep learning model for audio recognition, the audio signal is extracted and sampled. The model inference and thresholding are used to generate highlight clips and automatically edit the movie.

Benefits of technology

It enables automated generation of highlight clips, reduces manual labor, improves highlight clip processing efficiency, and allows users to obtain highlight clips simply by uploading movie files.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117612516B_ABST
    Figure CN117612516B_ABST
Patent Text Reader

Abstract

This invention discloses an automatic highlight editing method for movies based on deep learning audio classification, comprising the following steps: S1, training an audio recognition deep learning model based on the AudioSet public dataset; S2, extracting audio signals from the movie to be processed at a sampling rate of 16000; S3, sampling the extracted audio signals with a sampling window of 64000 and a sampling interval of 32000; S4, performing inference on each sampling interval using the audio recognition deep learning model, and saving the results according to classification; S5, integrating the results by category, mapping them to the movie frame count, and averaging the overlapping parts; S6, assigning 0-1 values ​​to the results of S5 according to a threshold; S7, calculating the start and end times of the corresponding highlight segments as the part with a response of 1, and averaging the scores of S5 within each segment as the highlight score; S8, editing and outputting the segments based on the highlight segment time information.
Need to check novelty before this filing date? Find Prior Art