Video processing method and electronic equipment
By analyzing the sound content in the material video and automatically identifying and generating highlight clips, the user's misoperation problem during video editing on electronic devices is solved, and the accuracy and efficiency of video editing is improved.
Patent Information
- Application Number
- CN202311867485.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
AI Technical Summary
用户在电子设备上进行视频剪辑时,难以精准定位保留片段的起始和结束位置,导致误操作频繁。
By analyzing the sound content in the material video, automatically identifying and generating highlight clips, reducing users' operations during video editing.
It reduces the possibility of misoperation during video editing and improves the accuracy and efficiency of video editing.
Smart Images

Figure CN120281960A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of terminal technology, and in particular to a video processing method and an electronic device. Background Art
[0002] As the video shooting functions of electronic devices such as mobile phones become more and more powerful, more and more users begin to use electronic devices to shoot videos to record their daily lives.
[0003] In actual use, after the user finishes shooting a video, he or she will often edit the video material he or she has shot in order to facilitate subsequent sharing or appreciation. In the process of video editing, the user often needs to edit a retained segment from a relatively long video material, and then connect the retained segments to generate a finished film.
[0004] To meet the needs of users to edit video materials, electronic devices will provide corresponding video editing applications, which display the video timeline and related editing functions in a graphical way, so that users can edit the required retained segments from the video materials. However, when users edit videos on electronic devices, it is often difficult to accurately locate the start and end positions of the retained segments due to the size of the electronic device screen and the operation method, which makes it easy for users to make mistakes. Summary of the invention
[0005] The present application provides a multi-video processing method and an electronic device to reduce the problem of user's prone to misoperation when editing videos on an electronic device.
[0006] According to the first aspect of the embodiments of the present invention, a video processing method is provided for generating a target video. The method can be applied to electronic devices. The method comprises: obtaining at least one piece of material video; determining a preliminary selection segment in the material video, wherein the preliminary selection segment refers to a video segment in the video segment contained in the material video that meets a predetermined screening condition; determining a highlight segment corresponding to the preliminary selection segment based on the analysis result of the sound content contained in the preliminary selection segment, wherein at least one of the highlight segments includes a preliminary selection segment containing a sound event that meets a predetermined condition; and generating a target video using the highlight segment. With this implementation, the highlight segment that needs to be retained in the material video can be automatically generated based on the image content of the material video and the sound in the material video, and then the target video can be generated using the highlight segment, thereby reducing the user operation in the video editing process and reducing the possibility of misoperation.
[0007] In combination with the first aspect, in a possible implementation manner, the target video contains at least one sound content corresponding to the sound event, and the sound event is obtained by performing sound event detection on the material video.
[0008] In combination with the first aspect, in a possible implementation, the obtaining of the material video includes: after receiving a specified operation of the user within a predetermined application, obtaining at least one segment of material video selected by the user through the operation process of the specified operation. By adopting this implementation, the complexity of importing the material video can be reduced, and the cumbersome process of importing video materials can be avoided.
[0009] In combination with the first aspect, in another possible implementation, determining the primary selection segment in the material video includes: selecting a video segment containing a sound event in the material video as the primary selection segment; or, selecting a video segment containing both high-score frame-extracted images and sound events in the material video as the primary selection segment, where the high-score frame-extracted image refers to a frame-extracted image with an image score higher than a predetermined score threshold, and the image score is obtained based on a scoring rule predetermined by a predetermined score, and the frame-extracted image refers to an image obtained by frame extraction of the material video. A video segment containing a sound event is often the content that the user wants to retain, and using the audio content corresponding to the sound event when generating the target video can make the content of the target video more rich.
[0010] In combination with the first aspect, in another possible implementation, the primary selection segment contains at least one complete sound event; or, the primary selection segment contains at least one sound event, and the start position of at least one sound event is outside the primary selection segment while the end position is inside the primary selection segment; or, the primary selection segment contains at least one sound event, and the start position of at least one sound event is inside the primary selection segment while the end position is outside the primary selection segment; or, the primary selection segment contains at least one sound event, and the start position and end position of the sound event are both outside the primary selection segment.
[0011] In combination with the first aspect, in another possible implementation, the at least one highlight segment includes a primary selection segment containing a sound event meeting a predetermined condition, including: the highlight segment contains a complete sound event; or, the highlight segment contains a sound event starting from the start of the sound event and having a length exceeding a predetermined retention duration, so as to facilitate subsequent connection or editing processing of the highlight segment.
[0012] In combination with the first aspect, in another possible implementation, the highlight segment contains a complete sound event, including: the start position of the highlight segment in the material video is before the start position of the sound event in the material video, and there is an interval of at least a predetermined interval duration, and the end position of the highlight segment in the material video is after the end position of the sound event in the material video, and there is an interval of at least a predetermined interval duration, so as to leave a part that can be further edited for subsequent processing.
[0013] In combination with the first aspect, in another possible implementation, generating the target video using the highlight segments includes: screening, editing, or clipping the highlight segments to obtain material segments; and generating the target video using the material segments.
[0014] In combination with the first aspect, in another possible implementation, generating the target video using the material segments includes: determining the retention mode of the audio content in the material segments; determining the high-value audio content in the material segments, where the high-value content is the audio content that needs to be retained determined based on the retention mode; and generating the target video using the image content and the high-value audio content included in the material segments. Generating the target video using the image content and the high-value audio content included in the material segments can make the content of the target video more abundant.
[0015] In combination with the first aspect, in another possible implementation, the retention mode of the audio content includes: intelligent retention mode, all retention mode, all non-retention mode, and user-defined mode.
[0016] In combination with the first aspect, in another possible implementation, when the retention mode of the audio content is the intelligent retention mode, determining the high-value audio content in the material segments includes: filtering out the high-value audio content from the audio content based on the loudness of the audio content; or, taking the audio content with a signal-to-noise ratio higher than the signal-to-noise ratio threshold as the high-value audio content; taking the audio content with a duration length exceeding the duration length threshold as the high-value audio content. With this implementation, the electronic device can automatically determine the high-value audio content in the material segments without the need for the user to set or operate.
[0017] In combination with the first aspect, in another possible implementation, filtering out the high-value audio content from the audio content based on the loudness includes: taking the audio content with a loudness higher than the loudness threshold as the high-value audio content; or, taking the audio content with a loudness change exceeding the loudness change threshold as the high-value audio content.
[0018] In combination with the first aspect, in another possible implementation, generating the target video using the image content and the high-value audio content included in the material segments includes: obtaining background dubbing; and generating the target video using the image content, the high-value audio content, and the background dubbing. The image part of the target video is obtained by connecting the image content of at least two material segments, and the audio part of the target video is obtained by superimposing the high-value audio content of the material segments and the background dubbing. With this implementation, the content of the target video can be made more abundant.
[0019] In combination with the first aspect, in another possible implementation, a transition effect for connecting two of the material segments is provided within the special effect interval between two adjacent material segments, and the special effect interval covers the connection interval between the two material segments and does not coincide with the intervals where the audio content of the two material segments is located.
[0020] In combination with the first aspect, in another possible implementation, in the audio part, starting from the starting position of the interval corresponding to the audio content, the background dubbing is faded out from the first volume value to the second volume value, and the volume of the audio content is faded in from the third volume value to the fourth volume value; and before the end position of the interval corresponding to the audio content, the background dubbing is faded in from the second volume value to the first volume value, and the volume of the audio content is faded out from the fourth volume value to the third volume value; wherein, the first volume value is higher than the second volume value, and the third volume value is lower than the fourth volume value. By adopting this implementation, the audio content from the source video can be made more prominent in the target video.
[0021] In a second aspect, the present application provides an electronic device, which includes: a memory, a display screen, and one or more processors, and the above-mentioned memory, display screen are coupled to the above-mentioned processor; the memory is used to store computer program code, and the computer program code includes computer instructions; when the computer instructions are executed by the above-mentioned one or more processors, the electronic device executes the method as described in the first aspect and any of its possible implementations.
[0022] In a third aspect, the present application provides a chip system, which is applied to an electronic device. The chip system includes one or more interface circuits and one or more processors. The interface circuits and the processors are interconnected by lines. The interface circuit is used to receive a signal from the memory of the electronic device and send the signal to the processor, and the signal includes the computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device executes the method as described in the first aspect and any of its possible implementations.
[0023] In a fourth aspect, the present application provides a computer-readable storage medium, which includes computer instructions, and when the computer instructions are run on an electronic device, the electronic device is caused to execute the method as described in the first aspect and any of its possible implementations.
[0024] In a fifth aspect, the present application provides a computer program product, and when the computer program product is run on a computer, the computer is caused to execute the method as described in the first aspect and any of its possible implementations.
[0025] The video processing method and electronic device provided in this application can automatically generate and determine the highlight segments to be retained in the material video based on the image content and sound in the material video, and then use the highlight segments to generate the target video, thereby reducing the user operations during video editing and reducing the possibility of incorrect operations. It can be understood that for the beneficial effects that can be achieved by the electronic device described in the second aspect, the chip system described in the third aspect, the computer-readable storage medium described in the fourth aspect, and the computer program product described in the fifth aspect, reference can be made to the beneficial effects in the first aspect and any of its possible implementation manners, which will not be elaborated here. Description of Drawings
[0026] Figure 1A It is a schematic structural diagram of an electronic device in an embodiment of this application;
[0027] Figure 1B It is a schematic software structure diagram of an electronic device in an embodiment of this application;
[0028] Figure 1C It is a schematic software structure diagram of an electronic device in an embodiment of this application;
[0029] Figure 2 It is a schematic flowchart of an embodiment of the video processing method in an embodiment of this application;
[0030] Figure 3A It is a schematic diagram of an example of the material video selection interface in an embodiment of this application;
[0031] Figure 3B It is a schematic diagram of another example of the material video selection interface in an embodiment of this application;
[0032] Figure 4 It is a schematic diagram of the composition of the material video in an embodiment of this application;
[0033] Figure 5A It is a schematic diagram of a way to determine the initial selected segments in an embodiment of this application;
[0034] Figure 5B It is a schematic diagram of another way to determine the initial selected segments in an embodiment of this application;
[0035] Figure 5C It is a schematic diagram of yet another way to determine the initial selected segments in an embodiment of this application;
[0036] Figure 6A It is a schematic diagram of the relationship between the initial selected segments and the sound events in an embodiment of this application;
[0037] Figure 6B It is a schematic diagram of the relationship between the material segments and the sound events in an embodiment of this application;
[0038] Figure 7A It is a schematic diagram of the correspondence between the volume curve of the audio content and the material segment in the embodiment of the present application;
[0039] Figure 7B It is another schematic diagram of the correspondence between the volume curve of the audio content and the material segment in the embodiment of the present application;
[0040] Figure 7C It is still another schematic diagram of the correspondence between the volume curve of the audio content and the material segment in the embodiment of the present application;
[0041] Figure 8 It is a schematic structural diagram of a chip system in the implementation of the present application. Detailed implementation manners
[0042] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.
[0043] The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in the specification and appended claims of the present application, the singular forms "a", "an", "the", "above-mentioned", "said", and "this" are also intended to include the expression form such as "one or more", unless there is a clear indication to the contrary in the context. It should also be understood that in the following embodiments of the present application, "at least one" and "one or more" mean one, two or more than two. The term "and / or" is used to describe the association relationship of associated objects and means that three relationships can exist; for example, A and / or B can mean: A exists alone, A and B exist simultaneously, and B exists alone, where A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship.
[0044] Referring to "one embodiment" or "some embodiments" described in this specification means that specific features, structures, or characteristics described in combination with the embodiment are included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different parts of this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0045] For the sake of clear and concise description of the following embodiments, first, a brief introduction to the relevant technologies and concepts in each implementation of the present application is made:
[0046] In various embodiments of the present application, the source video refers to a multimedia file containing images and audio, which can be generated by a user using an electronic device or received by the electronic device through channels such as the network. The images contained in the source video can be referred to as image content, and the audio contained in the source video can be referred to as audio content.
[0047] In various embodiments of the present application, a video clip refers to a part of a video. The video clip can contain both image content and audio content corresponding to the image content. The video clip can exist in the form of an independent file or as a part of a video. When the clip exists as a part of a video, it can be determined by its start and end timestamps in the source video, and the part between the start and end timestamps in the source video is the video clip.
[0048] In various embodiments of the present application, a highlight segment (Highlight) generally refers to the most exciting or important part extracted from a longer video. These segments are usually the parts that can attract the audience's attention the most in the video and can represent the core content or the most appealing moments of the entire video. For example, in a video shot by a user, the highlight segment may be the most intense or emotionally rich segment in the video.
[0049] In various embodiments of the present application, a source segment refers to a video segment used to generate a target video. The source segment can be a highlight segment or a part clipped from a highlight segment, that is, the source segment can be a segment of the source video. In some embodiments, the source segment can also include, for example, preset transition source segments, special effect source segments, etc. in an electronic device or a video editing APP. For the convenience of describing the solution of the present application, in the embodiments of the present application, only the example that the source segment is a segment of the source video is used for illustration, but it does not mean that it must be a segment of the source video.
[0050] The target video refers to a video generated based on the source video. In different application scenarios, the target video can also be referred to as a short video, a short film, a microfilm, etc. In addition to the image content and audio content of the source segment, the target video may also contain other source content, such as background dubbing, font special effects, color grading filters, transition sources, etc. Among them, the background dubbing can be background music (BGM), or it can be human voice or other sounds.
[0051] In various embodiments of the present application, the position refers to the time point in the video. The video mentioned here includes the source video, the target video, and various video clips mentioned in the embodiments. For the convenience of description, in the subsequent embodiments, the start position will also be referred to as the start or start moment, and the end position will also be referred to as the end or end moment.
[0052] In various embodiments of the present application, the image quality score refers to the evaluation result obtained by evaluating an image through image quality assessment (IQA). In the present application, the image quality evaluation may be an objective image quality evaluation, that is, the result obtained by using various algorithms and calculation models to evaluate the quality of an image.
[0053] In various embodiments of the present application, a sound event may also be referred to as an audio event, which refers to a specific type of sound or sound pattern that has a specific meaning or relevance. Sound events can be human-generated sounds (such as speech, laughter, crying), natural sounds (such as bird calls, rain sounds), mechanical sounds (such as vehicle, machine operation sounds), etc. The sound events in the source video can be determined through sound event detection, and the specific implementation process will not be elaborated here. The sound events in various embodiments of the present application may refer to all detectable sound events, or only a part of specific sound events, such as human-generated sounds.
[0054] To solve the above problems, the present application provides a video processing method, which can be executed by an electronic device. The electronic device may be a wireless terminal, an in-vehicle wireless terminal, a portable device, a wearable device, a mobile phone (or referred to as a "cellular" phone), a portable, pocket-sized, handheld terminal, etc., which exchange voice and / or data with a radio access network. For example, personal communication service (PCS) phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), and other devices. The wireless terminal may also be a subscriber unit, an access terminal, a user terminal, a user agent, a user device, or a user equipment (UE), etc. The present application does not limit the type of the electronic device.
[0055] Taking a mobile phone as an example of the above electronic device, in this embodiment, the structure of the electronic device may be as Figure 1A shown, where Figure 1A is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0056] As Figure 1AAs shown in the figure, the electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, etc. Among them, the sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.
[0057] Furthermore, when the electronic device is a mobile phone, the electronic device may further include: an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, and a subscriber identification module (SIM) card interface 195, etc.
[0058] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than those shown in the figure, or combine certain components, or split certain components, or have different component arrangements. The components shown in the figure may be implemented in hardware, software, or a combination of software and hardware.
[0059] The processor 110 may include one or more processing units. Among them, different processing units may be independent devices or integrated in one or more processors. A memory may also be provided in the processor 110 for storing instructions and data.
[0060] In some embodiments, the processor 110 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0061] The charging management module 140 is configured to receive a charging input from a charger. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 to supply power to the processor 110, the internal memory 121, the display screen 194, the camera 193, the wireless communication module 160, etc.
[0062] The mobile communication module 150 may provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device. The wireless communication module 160 may provide solutions for wireless communications including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the electronic device. In some embodiments, the antenna 1 of the electronic device is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device can communicate with the network and other devices through wireless communication technologies.
[0063] The electronic device realizes the display function through a graphics processing unit (GPU), a display screen 194, an application processor, etc. The GPU is a microprocessor for image processing, connecting the display screen 194 and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.
[0064] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. A series of graphical user interfaces (GUIs) can be displayed on the display screen 194 of the electronic device, and these GUIs are all the main screens of the electronic device.
[0065] The electronic device can realize the shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc.
[0066] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 110 through the external memory interface 120 to realize the data storage function. The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 121.
[0067] The electronic device can realize the audio function through an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone interface 170D, an application processor, etc. Such as music playing, recording, etc.
[0068] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be set in the processor 110, or some functional modules of the audio module 170 can be set in the processor 110.
[0069] The speaker 170A, also known as the "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device can listen to music or a hands-free call through the speaker 170A. The receiver 170B, also known as the "earpiece", is used to convert an audio electrical signal into a sound signal. The microphone 170C, also known as the "microphone", "transmitter", is used to convert a sound signal into an electrical signal. When making a call or sending a voice message, the user can speak close to the microphone 170C with their mouth to input the sound signal into the microphone 170C. The headphone jack 170D is used to connect a wired headphone.
[0070] The pressure sensor 180A is used to sense a pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 180A can be disposed on the display screen 194. There are many types of pressure sensors 180A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The gyroscope sensor 180B can be used to determine the motion posture of the electronic device. The barometric pressure sensor 180C is used to measure the barometric pressure. The magnetic sensor 180D includes a Hall sensor. The electronic device can use the magnetic sensor 180D to detect the opening and closing of a flip leather case. The acceleration sensor 180E can detect the magnitude of the acceleration of the electronic device in various directions (generally three axes). The distance sensor 180F is used to measure the distance. The proximity light sensor 180G can include a light-emitting diode (LED) and a light detector. The ambient light sensor 180L is used to sense the ambient light brightness. The fingerprint sensor 180H is used to collect fingerprints. The temperature sensor 180J is used to detect the temperature. The touch sensor 180K, also known as the "touch control device". The touch sensor 180K can be disposed on the display screen 194, and the touch sensor 180K and the display screen 194 form a touch screen, also known as the "touch panel". The bone conduction sensor 180M can obtain a vibration signal. The keys 190 include a power-on key, volume keys, etc. The keys 190 can be mechanical keys, touch keys, or virtual keys. The motor 191 can generate a vibration prompt. The indicator 192 can be an indicator light and can be used to indicate the charging state, battery level change, or can also be used to indicate messages, missed calls, notifications, etc. The SIM card interface 195 is used to connect a SIM card.
[0071] In addition, an operating system runs on the above components. For example, the iOS operating system developed by Apple Inc., the Android open-source operating system developed by Google Inc., the Windows operating system developed by Microsoft Corporation, etc. Application programs (applications, APPs) can be installed and run on this operating system.
[0072] To facilitate the description of the functional operations performed by each software architecture within an electronic device when executing the solution disclosed in this application, the embodiments of this application also disclose the software structure of the electronic device. Still taking a mobile phone as an example of the above-mentioned electronic device, the software system of the mobile phone can adopt a layered architecture, an event-driven architecture, a microkernel architecture, a microservices architecture, or a cloud architecture. The embodiments of this application take the Android system with a layered architecture as an example to exemplarily illustrate the software structure of the mobile phone.
[0073] As Figure 1B shown, the layered architecture divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system can be divided into five layers, from top to bottom are the application layer, the application framework layer, Android runtime and system libraries, the hardware abstraction layer (HAL), and the kernel layer.
[0074] The application layer may include a series of application packages.
[0075] As Figure 1B shown, the application packages may include applications such as the camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, etc. The video processing method of the embodiments of this application can be applied to the camera, gallery, video, or video editing application. Among them, the video editing application can be built-in with a one-key blockbuster automatic video generation engine to implement the solutions shown in each example of this application.
[0076] The application framework layer provides application programming interfaces (APIs) and programming frameworks for the applications in the application layer. The application framework layer includes some predefined functions.
[0077] As Figure 1B shown, the application framework layer may include a window manager, a content provider, a view system, a phone manager, a resource manager, a notification manager, etc.
[0078] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.
[0079] The content provider is used to store and obtain data, and make this data accessible to applications. The data may include video, image, audio, dialed and answered calls, browsing history and bookmarks, phone book, etc.
[0080] The view system includes visual controls, such as controls for displaying text, controls for displaying pictures, etc. The view system can be used to build applications. The display interface can be composed of one or more views. For example, a display interface including a text message notification icon can include a view for displaying text and a view for displaying pictures.
[0081] The phone manager is used to provide the communication function of the electronic device 100. For example, the management of call status (including answering, hanging up, etc.).
[0082] The resource manager provides various resources for applications, such as localized strings, icons, pictures, layout files, video files, and so on.
[0083] The notification manager enables applications to display notification information in the status bar. It can be used to convey informative messages, which can automatically disappear after a short stay without user interaction. For example, the notification manager is used to inform that the download is complete, message reminders, etc. The notification manager can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background running application, or a notification that appears on the screen in the form of a dialog window. For example, it prompts text information in the status bar, emits a prompt tone, the electronic device vibrates, the indicator light flashes, etc.
[0084] The media processing middle platform is used to perform various types of processing on videos. In various embodiments of the present application, the primary selection segment and the highlight segment can be determined by the media processing middle platform.
[0085] After the media processing middle platform obtains the source video from the camera, gallery, video, or video editing application, it can determine the highlight segment according to the method shown in the embodiments of the present application, and then return the highlight segment or the indication information for indicating the highlight segment to the camera, gallery, video, or video editing application. Then, the camera, gallery, video, or video editing application uses the highlight video to generate the target video. In a specific implementation, a media analysis algorithm monitoring and scheduling module can be set in the media processing middle platform. The primary selection segment and the highlight segment are determined through the media analysis algorithm monitoring and scheduling module, and the highlight segment is provided to the video editing application. Then, the video editing application generates the target video based on the highlight segment. During the generation process of the target video, the interaction mode between the media processing middle platform and the video editing application is as Figure 1C shown.
[0086] Android Runtime includes core libraries and virtual machines. Android runtime is responsible for the scheduling and management of the Android system.
[0087] The core libraries include two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android.
[0088] The application layer and the application framework layer run in the virtual machine. The virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.
[0089] The system libraries can include multiple functional modules. For example: surface manager, Media Libraries, 3D graphics processing library (e.g., OpenGL ES), 2D graphics engine (e.g., SGL), etc.
[0090] The surface manager is used to manage the display subsystem and provides the fusion of 2D and 3D layers for multiple applications.
[0091] The media library supports the playback and recording of multiple common audio and video formats, as well as static image files, etc. The media library can support multiple audio and video coding formats, such as: MPEG4, H.264, MP3, AAC, AMR, JPG, PNG, etc.
[0092] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, synthesis, and layer processing, etc.
[0093] The 2D graphics engine is the graphics engine for 2D drawing.
[0094] HAL is the abstract interface of the device kernel driver, which realizes the application programming interface for accessing the underlying device to a higher-level Java API framework. The hardware abstraction layer can include multiple library modules, such as the display module, audio module, Bluetooth module, Wi-Fi module, etc. Each module can implement an interface for a specific type of hardware component. When the framework API requests access to the device hardware, the Android system will load the library module for this hardware component.
[0095] The kernel layer is the layer between the hardware and the software. The kernel layer at least includes the display driver, camera driver, audio driver, sensor driver. Among them, the hardware can include Figure 1A the hardware structure shown.
[0096] Exemplarily, the SAT algorithm module in the software architecture can interact with the ISP in the hardware to implement the video processing method provided by the embodiments of the present application.
[0097] The video processing method provided by the embodiments of the present application will be described in detail below in combination with the software and hardware of the electronic device.
[0098] Figure 2 It is a schematic flow chart in an embodiment of the video processing method of the present application. AsFigure 2 As shown, this embodiment may include the following steps:
[0099] 101. Obtain a source video.
[0100] The source video can be a video captured by the camera of an electronic device, a video recorded using the screen recording function, or a video obtained through other means. For example, it can be a video imported from another device into the electronic device.
[0101] It should be noted that in various embodiments of the present application, the source video can be one segment or multiple segments. That is, the electronic device can generate one target video based on one source video, or generate one target video based on multiple source videos. The source video can be obtained by a video editing APP in the electronic device. The video editing APP can automatically select the source video, or the user can manually select the source video through the video editing APP. After the electronic device receives a specified operation of the user within a predetermined application, it obtains at least one segment of the source video selected by the user during the operation process of the specified operation, thereby avoiding the complexity of importing the source video and avoiding the cumbersome process of importing video materials.
[0102] For example, the user can select the source video on the video playback interface. As Figure 3A shown, the user can select the video being previewed as the source video by clicking the "One - click Blockbuster" button in the controls displayed on the video playback interface. Another example is that the user can also select the source video on the video thumbnail preview interface. As Figure 3B shown, the user can select one or more videos in the Photo APP, and then select the "One - click Movie" button in the pop - up control to select one or more videos as the source video. It should be noted that there are many ways to select the source video, which will not be listed one by one here.
[0103] 102. Determine the primary selected segments in the source video.
[0104] The primary selected segments refer to one or more of the video segments included in the source video that meet the predetermined screening conditions. The electronic device can divide the source video into several video segments according to a predetermined processing method or processing rule, and then select a predetermined number of primary selected segments from them; or, the electronic device can directly select a predetermined number of primary selected segments from the source video according to a predetermined selection rule or selection method.
[0105] The number of initial selection segments and the time length of each initial selection segment can be set as needed. The time lengths of different initial selection segments can be the same or different. The number of initial selection segments and the time length of each initial selection segment can be preset values. For example, it can be preset to separately identify 8 initial selection segments from each source video or a total of 8 initial selection segments from all source videos, and the duration of each initial selection segment is 10 seconds. Alternatively, the number of initial selection segments and the time length of each initial selection segment can also be adaptively determined according to the length of the source video or processing time requirements, etc. When the duration of the source video is relatively long, the number or length of the initial selection segments can be adaptively increased.
[0106] The electronic device can select at least one initial selection segment from the source video segments according to a preset selection method.
[0107] In some embodiments, the electronic device can extract frames from the source video to obtain a series of frame-extracted images, and then divide the source video into several video segments according to the content similarity of the frame-extracted images; then select the initial selection segments from among the video segments based on their content. For example, Figure 4 As shown, it can be determined according to the frame extraction result of the video that the source video includes three video segments: a landscape segment, a close-up segment of Person 1, and a close-up segment of Person 2. Then, the close-up segment of Person 1 and the close-up segment of Person 2 are selected as the initial selection segments.
[0108] In some embodiments, the electronic device can determine the initial selection segments according to the image value of the frame-extracted images in the source video, where the frame-extracted images refer to the images obtained by extracting frames from the source video.
[0109] In one implementation, the electronic device can extract frames from the source video to obtain a series of frame-extracted images; then perform image quality scoring on each frame-extracted image according to a preset scoring rule; then select the video segments containing high-scoring frame-extracted images as the initial selection segments, or determine the initial selection segments from the source video based on the high-scoring frame-extracted images, where the high-scoring frame-extracted images refer to the frame-extracted images with an image score higher than a preset score threshold. For example, the video segments with a certain duration before and after the high-scoring frame-extracted images can be used as the initial selection segments; or when two or more consecutive frame-extracted images are high-scoring frame-extracted images, the video segments between these frame-extracted images can be used as the initial selection segments. Among them, the high-scoring frame-extracted images can also be referred to as high-value images.
[0110] As Figure 5A shown, when the frame-extracted images 2, 4, 5, 6, and 7 are high-scoring frame-extracted images, the segments 1, 2, and 3 can be correspondingly selected as the initial selection segments.
[0111] Among them, the electronic device can determine a scoring rule according to evaluation criteria such as the image quality of the sampled frame image, the content value of the sampled frame image, and the content similarity of the sampled frame image, and then determine the score of each sampled frame image according to the scoring rule. When the score of the sampled frame image is higher than a predetermined standard, it can be considered that the sampled frame image is a high-score sampled frame image. The high-score sampled frame image can also be called a high-value image. The specific scoring rule and the determination method of the scoring rule will not be elaborated here.
[0112] In some embodiments, the electronic device can select a video segment containing a sound event in the source video as a primary selection segment. A video segment containing a sound event is often the content that the user wants to retain, and using the audio content corresponding to the sound event when generating the target video can make the content of the target video more rich.
[0113] In one implementation, the electronic device can perform interval sampling on the audio content included in the source video to obtain a series of sampled segments; then detect whether the sampled segments contain sound events; if the sampled segment contains a sound event, then the video segment corresponding to the sampled segment can be used as a primary selection segment. Among them, the sampled segment containing a sound event means that the sampled segment contains at least a part of the sound event.
[0114] It should be noted here that for the convenience of subsequent editing, when selecting a video segment containing a sound event as a primary selection segment, the start position of the primary selection segment needs to be before the start position of the sound event and at least a predetermined interval duration apart, and the end position needs to be after the end position of the sound event and also at least a predetermined interval duration apart.
[0115] As Figure 5B shown, when sampled segment 2, sampled segment 4, and sampled segment 5 contain sound events, segment 4 and segment 5 can be correspondingly selected as primary selection segments.
[0116] In some embodiments, the electronic device can also comprehensively use multiple methods to determine the primary selection segment in the source video.
[0117] In one implementation, the electronic device can select a video segment that contains both high-score sampled frame images and sound events as a primary selection segment. For example, the electronic device can first determine the video segments containing high-score sampled frame images, and then analyze whether these video segments contain sound events; if the video segment contains a sound event, then the video segment is used as a primary selection segment; if the video segment does not contain a sound event, then the video segment is not used as a primary selection segment. Alternatively, the electronic device can also first determine the video segments containing sound events, and then select the video segments containing high-score sampled frame images from the video segments containing sound events.
[0118] In another implementation, the electronic device may first determine some initial selected segments based on the image value of the source video, and then, in the remaining part of the source video, select some video segments containing sound events as another part of the initial selected segments. Alternatively, the electronic device may first determine some initial selected segments based on the sound events, and then, in the remaining part of the source video, determine some video segments as another part of the initial selected segments according to the image value of the source video. The specific implementation process of this implementation may refer to the foregoing related implementation manners, and will not be elaborated herein. It should be noted that, in various embodiments of the present application, for the same source video, the frame extraction frequency of image frame extraction and the sampling interval frequency of audio sampling may be the same or different, and the positions of image frame extraction and audio sampling may be the same or different.
[0119] As Figure 5C shown, when sampling segments 2, 4, and 5 contain sound events and frame images 4, 5, 6, and 7 are high-score frame images, segments 6, 8, and 9 may be selected as the initial selected segments accordingly, or segments 6, 7, and 9 may be selected as the initial selected segments, or only segment 7 or segment 8 may be selected as the initial selected segment. The specific selection process will not be elaborated herein.
[0120] It should be noted here that the above-mentioned methods for determining the initial selected segments are only examples, and do not represent all implementation manners of the embodiments of the present application. The embodiments of the present application do not limit the specific methods for determining the initial selected segments either. For example, the electronic device may also use a pre-trained analysis model to determine the initial selected segments from the source video, or determine the initial selected segments according to the manual selection of the user, etc. The specific implementation process will not be elaborated herein. It should be further noted that the electronic device may determine the initial selected segments only through one selection method; or may determine the initial selected segments through multiple different selection methods respectively; or may also combine different selection methods to select the initial selected segments that meet multiple selection methods at the same time. The present application does not limit this either.
[0121] 103. Determine the highlight segments corresponding to the initial selected segments according to the analysis result of the sound content included in the initial selected segments.
[0122] Since there are various ways to determine the primary selection segments, for a primary selection segment, it may or may not contain a highlight segment. Moreover, it may contain other content in addition to the complete highlight segment, or it may not contain all the content of the complete highlight segment. Therefore, after obtaining the primary selection segment, it is necessary to further analyze the primary selection segment to determine the corresponding highlight segment through analysis. The highlight segment can be further screened, edited, or added based on the primary selection segment. Further analysis of the primary selection segment can include analyzing the sound content contained in the primary selection segment to determine whether the sound content contains a sound event.
[0123] In various embodiments of the present application, there are multiple ways to determine the highlight segment. Depending on the different ways to determine the highlight segment, the highlight segment may or may not contain a sound event. When the highlight segment contains a sound event, for the convenience of subsequent connection or editing processing of the highlight segment, the following conditions need to be satisfied simultaneously between the highlight segment and the sound event: (1) The highlight segment must contain the start part of the sound event; (2) When the duration of a sound event does not exceed the predetermined retention duration, the highlight segment must contain the complete sound event; (3) When the duration of a sound event exceeds the predetermined retention duration, the highlight segment must contain at least the sound event of the predetermined retention duration. It should be noted that multiple sound events can be included in the same highlight segment, but the first sound event among them should not be an incomplete tail of a sound event, that is, the first sound event in the highlight segment needs to be a complete sound event. Among them, the predetermined retention duration can be set in advance according to needs. For example, it can be set to 3.5 seconds; or it can also be set by the user each time the target video is generated, so that target videos with different editing rhythms can be generated. For example, a highlight segment can be 5 seconds, and the sound event in it can last from the 1st second to the 2nd second of the highlight segment; or it can also last from the 1st second to the end of the highlight segment.
[0124] And depending on the different ways to determine the primary selection segment, the primary selection segment determined by the electronic device may or may not contain a sound event. When the primary selection segment contains a sound event, the corresponding relationship between the primary selection segment and the sound event may be as follows:
[0125] The first case is that the primary selection segment contains at least one complete sound event. As Figure 6A shown, the primary selection segment 2 contains the complete sound event 2 and the sound event 3.
[0126] The second case is that a sound event lasts throughout the primary selection segment, and the start and end time points are unknown, that is, both the start position and the end position of the sound event are outside the primary selection segment. As Figure 6AAs shown, the initial selection segment 3 contains a part of the sound event 4, but both the start and end of the sound event 4 are outside the initial selection segment 3.
[0127] In the third case, the start time point of the sound event is unknown and the end time point is known, that is, the start position of the sound event is outside the initial selection segment and the end position is inside the initial selection segment. The sound event starts before the initial selection segment and ends inside the initial selection segment. As Figure 6A shown, the initial selection segment 1 contains a part of the sound event 1, but the start of the sound event 1 is outside the initial selection segment 1.
[0128] In the fourth case, the sound event has a start time point and the end time point is unknown, that is, the start position of the sound event is inside the initial selection segment and the end position is outside the initial selection segment. The sound event starts inside the initial selection segment but ends after the initial selection segment. As Figure 6A shown, the initial selection segment 4 contains a part of the sound event 5, but the start of the sound event 5 is outside the initial selection segment 4.
[0129] It should be noted that when there are more than one sound event in the initial selection segment, all the sound events may be complete sound events, or it is possible that the start position of one sound event is outside the initial selection segment and the end position is inside the initial selection segment, or it is also possible that the start position of one sound event is inside the initial selection segment and the end position is outside the initial selection segment. In some embodiments, the initial selection segment may also not contain any sound event.
[0130] For different corresponding relationships between the initial selection segment and the sound event, different methods can be used to determine the highlight segment.
[0131] For the first case above, the initial selection segment can be directly used as the highlight segment; or the initial selection segment can also be split into multiple highlight segments so that each highlight segment contains a complete sound event. As Figure 6A shown, the initial selection segment 2 can be directly used as the highlight segment. In this case, the sound events included in the highlight segment are all complete sound events. For the convenience of subsequent processing, the start position of the highlight segment in the source video is before the start position of the sound event in the source video and is separated by at least a predetermined interval duration, and the end position of the highlight segment in the source video is after the end position of the sound event in the source video and is separated by at least a predetermined interval duration.
[0132] For the second case above, the initial selection segment can be discarded, that is, the initial selection segment is not used as a highlight segment; alternatively, the parts before and after the initial selection segment in the source video can be combined with the initial selection segment to obtain a highlight segment that includes the complete sound event; or, the part before the initial selection segment in the source video can be combined with the initial selection segment to obtain a highlight segment that includes the start of the audio and the sound event lasts longer than a predetermined retention duration; or the initial selection segment can be used as a highlight segment that does not include a sound event. As Figure 6A shown, the initial selection segment 3 can be directly discarded and not used as a highlight segment, or the initial selection segment 3 can be used as a highlight segment that does not include a speech event; or, the part between the end of the initial selection segment 2 and the start of the initial selection segment 4 can be used as the highlight segment corresponding to the initial selection segment 3, or the part between the start of the sound event 4 and the end of the sound event 4 can be used as the highlight segment corresponding to the initial selection segment 3, so that the highlight segment can completely include the sound event 4.
[0133] For the third case above, the initial selection segment can be discarded, that is, the initial selection segment is not used as a highlight segment; or an addition can be made on the basis of the initial selection segment, and the part before the initial selection segment in the source video can be combined with the initial selection segment to obtain a highlight segment that includes the complete sound event. As Figure 6A shown, the initial selection segment 1 can be directly discarded and not used as the highlight segment 1; or the part between the start of the sound event 1 and the end of the initial selection segment 1 can be used as the highlight segment corresponding to the initial selection segment 1.
[0134] For the fourth case above, the initial selection segment can be discarded, that is, the initial selection segment is not used as a highlight segment; or an addition can be made on the basis of the initial selection segment, and the part after the initial selection segment in the source video can be combined with the initial selection segment to obtain a highlight segment that includes the complete sound event; if the sound event in the initial selection segment lasts longer than a predetermined retention duration, the initial selection segment can also be directly used as a highlight segment. As Figure 6A shown, the initial selection segment 4 can be directly discarded and not used as a highlight segment; or the part between the start of the initial selection segment 4 and the end of the sound event 51 can be used as the highlight segment corresponding to the initial selection segment 1; or, if the duration from the start of the sound event 5 to the end of the initial selection segment 4 exceeds the predetermined retention duration, then the initial selection segment 4 can also be directly used as a highlight segment.
[0135] For the fifth case above, the initial selection segment can be discarded, or the initial selection segment can be used as a highlight segment that does not include a sound event. As Figure 6A shown, the initial selection segment 5 does not include any sound events. In this case, the initial selection segment 5 can be directly discarded and not used as a highlight segment; or, the initial selection segment 5 can be used as a highlight segment that does not include a speech event.
[0136] 104. Generate a target video using highlight segments.
[0137] After obtaining the highlight segments, the electronic device can generate a target video based on the highlight segments. The electronic device can directly use the highlight segments to generate the target video; or it can further process the highlight segments to obtain corresponding material segments, and then use the material segments to generate the target video.
[0138] It should be noted that when generating the target video, in addition to using the material segments, other material contents may also be used. For example, background dubbing, font effects, color correction filters, transition materials, etc. may be used. These material contents can be provided by the video editing APP for use when generating the target video. When generating the target video, in addition to directly connecting the material segments, transition effects can also be added between different material segments to make the target video more vivid.
[0139] In some embodiments, to make the target video more visually appealing, the electronic device often organizes the video segments according to a certain rhythm. However, when the number of highlight segments is large or the duration is long, the highlight segments need to be screened, edited, or clipped to obtain corresponding material segments, and then the processed highlight segments are used as material segments to generate the target video. When using the material segments to generate the target, both the image content and the audio content in the material segments can be used; or, only the image content or the audio content in the material segments can be used. The audio content can also be referred to as the original sound.
[0140] For example, the electronic device may accurately align the video segments in the target video with the rhythm of the background music according to the rhythm of the background dubbing. When generating the target video, it may occur that the duration of the highlight segment does not match the rhythm of the background dubbing. Therefore, the highlight segment needs to be clipped, and only the part that can match the rhythm of the background dubbing is retained as the material segment.
[0141] When clipping the highlight segment, if the highlight segment contains a sound event, the material segment obtained by clipping can contain the sound event or not. When the material segment contains a sound event, the following relationship conditions need to be satisfied simultaneously between the material segment and the sound event: (1) The material segment must contain the starting part of the sound event; (2) When the duration of a sound event does not exceed the predetermined retention duration, the material segment must contain the complete sound event; (3) When the duration of a sound event exceeds the predetermined retention duration, the material segment must contain at least the predetermined retention duration of the sound event.
[0142] Such as Figure 6BAs shown, the material segment 1 completely contains the sound event 1, the material segment 2 completely contains the sound events 2 and 3, the material segment 3 contains the beginning of the sound event 4 and the part of the sound event 4 that exceeds the predetermined retention duration.
[0143] Since the material segment contains both image content and audio content, when using the material segment to generate the target video, the audio content of the material segment can be used, or the audio content of the material segment can be not used, or only part of the audio content of the material segment can be used.
[0144] Which audio content the electronic device uses to generate the target video can be determined by the retention mode of the audio content. Generally speaking, there are the following several retention modes for the audio content in the material segment:
[0145] The first mode: intelligent retention mode. That is, intelligently retain part of the audio content in the material segment, and the part of the audio content to be retained is determined by the electronic device according to the preset retention rules. In this mode, only the valuable or meaningful part is retained, that is, only the high-value audio content is retained. Among them, the high-value audio content refers to the audio content that meets the predetermined screening conditions, and the high-value audio content can be audio content of a specific type, with specific characteristics or with specific meanings. The electronic device will intelligently decide which audio content to retain according to the content of the material segment. For example, retain the human voice in the dialogue scene, and remove the human voice in the natural landscape picture to highlight the picture atmosphere. Taking Figure 6B the shown material segment 2 as an example, only the original sound corresponding to the sound events 2 and 3 can be retained, and the other parts do not retain the original sound.
[0146] The second mode: all retention mode. That is, retain all the audio content in the material segment, and use all the audio content in the material segment to generate the target video. In this mode, whether it is dialogue, background music or environmental sound, they will all be retained. This mode is suitable for those scenarios where all the original audio in the video wants to be retained, and is suitable for the production of recording or documentary-style videos. Taking Figure 6B the shown material segment 1 as an example, all the original sound can be retained.
[0147] The third mode: all non-retention mode. That is, do not retain any audio content in the material segment. In this mode, all the audio content in the material segment will be removed. This mode is suitable for the scenarios where dubbing or post-added sound effects are used completely, and provides more creative space for the video.
[0148] Fourth mode: User-defined mode that can be adjusted by the user. In this mode, the user can set the original sound retention method for each material segment according to their own needs, and the electronic device can determine which audio content to retain according to the user's settings. This mode provides greater flexibility and is suitable for professional users who have special requirements for the audio effects of videos.
[0149] When intelligently retaining the original sound, it is necessary to filter the audio content in the material segment so as to only retain the valuable or meaningful part as high-value audio content and avoid retaining worthless or meaningless audio content. Since there are multiple audio content retention modes, in actual use, the retention mode of the audio content in the material segment can be selected according to the content contained in the material segment or according to the user's operation. There are various implementation methods for filtering the audio content to obtain high-value audio content, and the following will be described with some examples. By adopting this implementation method, the electronic device can also automatically determine the high-value audio content in the material segment without the need for user settings or operations.
[0150] In some embodiments, the electronic device can filter the audio content based on loudness, and then filter out the high-value audio content from the audio content.
[0151] In one implementation method, when filtering the audio content based on loudness, the audio content with a loudness higher than the loudness threshold can be regarded as high-value audio content. The loudness threshold can be set according to the situation. For example, it can be set to 50 dB. When the loudness of the audio content exceeds the loudness threshold, it is retained as high-value audio content; if the loudness of the audio content does not exceed the loudness threshold, then it is not retained as high-value audio content. Filtering the audio content based on loudness can screen out high-value audio content, so as to retain the voice of the protagonist intentionally filmed and not retain the weak voice of passers-by in the distance.
[0152] Take Figure 6B the shown material segment 2 as an example. If the loudness of the audio content corresponding to sound event 2 exceeds 50 dB, while the loudness of the audio content corresponding to sound event 3 does not exceed 50 dB, then only the original sound corresponding to sound event 2 can be retained, and other parts including the part corresponding to sound event 3 are not retained with the original sound.
[0153] In one implementation method, when filtering the audio content based on loudness, it is also possible to filter the audio content based on the loudness change, and regard the audio content with a significant loudness change as high-value audio content. When the loudness change of the audio content exceeds the loudness change threshold, it is retained as high-value audio content; if the loudness change of the audio content does not exceed the loudness change threshold, then it is not retained as high-value audio content.
[0154] For example, in a multi-person banquet scenario, there are usually continuous conversations with relatively low volume, and there are also special contents such as "cheers" with a significantly increased loudness. By filtering the audio content based on the loudness change, these contents are screened out from the longer audio content. Filtering the audio content based on the loudness change can achieve the mining of high-value audio content from the audio content with a longer duration.
[0155] In some embodiments, the electronic device can also filter the audio content based on the signal-to-noise ratio filtering, and then filter out the high-value audio content from the audio content.
[0156] When filtering the audio content based on the signal-to-noise ratio, the audio content with a signal-to-noise ratio higher than the signal-to-noise ratio threshold can be regarded as high-value audio content. The signal-to-noise ratio threshold can be set according to the situation. For example, it can be set to 4 dB. When the signal-to-noise ratio of the audio content exceeds the signal-to-noise ratio threshold, it is retained as high-value audio content. If the signal-to-noise ratio of the audio content does not exceed the signal-to-noise ratio threshold, then it is not retained as high-value audio content. This filtering method can be applied to the scenario where the background noise of the audio content is relatively large. When the background noise of the audio content is relatively large, the corresponding audio content needs to be filtered out and not retained as high-value audio content. However, it should be noted that in order to restore the real sense of the scene, normal ambient sounds including slightly noisy ones should be allowed, and only very large and discomforting or almost drowning out normal speech background noises need to be filtered out, such as in a busy road or beside a train, in a scenic area waterfall or a strong background sound environment scene such as by the sea with strong winds and big waves.
[0157] In some embodiments, the electronic device can also filter the audio content based on the duration length of the audio content, and then filter out the high-value audio content from the audio content.
[0158] When filtering the audio content based on the duration length of the audio content, the audio content with a duration length exceeding the duration length threshold can be regarded as high-value audio content. The duration length threshold can be set according to the situation. For example, it can be set to 1 second or 0.5 second. When the time length of the audio content exceeds the duration length threshold, it is retained as high-value audio content; if the time length of the audio content does not exceed the duration length threshold, then it is not retained as high-value audio content. Filtering the audio content based on the time length can achieve a rough screening of high-value audio content through the time length filtering, so as to retain meaningful audio content and avoid retaining meaningless audio content.
[0159] In some embodiments, the electronic device can also filter the audio content based on the speech recognition result, and then filter out the high-value audio content from the audio content.
[0160] When filtering audio content based on semantics, speech recognition can be performed on the audio content first, and the parts with specific semantics can be regarded as high-value audio content. For example, in a scenario of multi-person conversation, the audio content related to a specific person can be recognized.
[0161] When intelligently preserving the original sound, since the audio content may be filtered, in order to avoid content fragmentation, it is necessary to limit the time length of the high-value audio content according to certain rules.
[0162] In some implementations, when the time length of the audio content is within a predetermined length range, the audio content needs to be completely preserved. For example, when the audio content contains "XXX Happy Birthday", all the audio content should be completely preserved to avoid truncation of the audio content. Among them, the predetermined length range can be set as needed. For example, it can be set to 3.5 seconds and below, or it can also be set to between 0.5 seconds and 3.5 seconds.
[0163] When the audio content is long, its length exceeds the predetermined length range, and it cannot all be preserved as high-value audio content, at least the predetermined time length starting from the starting position should be preserved. Among them, the predetermined length is within the predetermined length range. Usually, it can be the maximum value of the predetermined length range.
[0164] It should be noted here that for a material segment, after filtering the audio content, it may retain one or more segments of high-value audio content, or it may not retain any high-value audio content. When the material segment contains multiple segments of high-value audio content that meet the conditions, the first segment of high-value audio content should not be an incomplete tail of an audio content, that is, the first segment of high-value audio content should have a complete beginning.
[0165] When generating the target video, the electronic device can use only the content included in the material segment, or in addition to the content included in the material segment, it can also use other material content such as background dubbing, and can also add transition special effects and volume fade-in and fade-out effects.
[0166] In some embodiments, to make the target video content more rich, when generating the target video, in addition to using the audio content of the material segment itself, background dubbing is often added to the target video during the video generation process. Background dubbing can also be called background sound, background music or background score.
[0167] When the material segment contains audio content and there is background dubbing, in order to highlight the audio content in the target video, the volume of the background dubbing in the interval corresponding to the audio content can be reduced, and fade-in and fade-out can also be set for the background dubbing and the audio content.
[0168] In one implementation, in the audio part, starting from the starting position of the interval corresponding to the audio content, the background dubbing is faded out from the first volume value to the second volume value, and the volume of the audio content is faded in from the third volume value to the fourth volume value; and before the end position of the interval corresponding to the audio content, the background dubbing is faded in from the second volume value to the first volume value, and the volume of the audio content is faded out from the fourth volume value to the third volume value; wherein, the first volume value is higher than the second volume value, and the third volume value is lower than the fourth volume value.
[0169] For example, when the material segment contains audio content, within the time interval corresponding to the audio content, starting from the start position of the audio content, the volume of the background dubbing can be first reduced to 75% or 50% of the predetermined volume using a fade-out effect, and before the end position of the audio content, the volume of the background dubbing can be restored to the predetermined volume using a fade-in effect.
[0170] For example, when setting the fade-in and fade-out effects for the audio content, the duration of the fade-in and fade-out and the minimum volume of the fade-in and fade-out can be controlled as needed. The duration of the fade-in and fade-out extends forward and backward centered on the minimum interval boundary value. For example, the duration of the fade-in and fade-out can be set to 1S, and the minimum volume of the fade-in and fade-out effect can be set to not be lower than 75% of the original volume. That is, at the start of the audio content, it fades in from 75% of the original volume of the audio content and reaches 100% of the original volume after 1 second. When the audio content ends, it fades out starting from 1 second before the end of the audio content and reaches 75% of the original volume after 1 second.
[0171] As Figure 7A shown, when the material segment contains three pieces of audio content, namely audio content 1, audio content 2, and audio content 3, in the target video, the volume change of the audio content can be as shown by the curve of the audio content volume in the figure, and the volume change of the audio content can be as shown by the curve of the audio content volume in the figure, where the dashed part represents no audio content volume.
[0172] In some embodiments, if there are two or more material segments, the electronic device needs to connect the material segments. In some embodiments, to make the target video more aesthetically pleasing, when generating the target video, that is, when connecting the material segments, a transition effect is often set in the connection interval between the material segments to cover the connection interval. The transition effect can also be called a transition special effect, including mask, cross dissolve, overlay dissolve, fade-in and fade-out, etc.
[0173] Since the transition effect lasts for a certain length of time, the time interval during which the transition effect lasts can be referred to as the effect interval. The effect interval may cover the ending part of the previous material segment and / or the starting part of the subsequent material segment. That is, a transition effect for connecting the two material segments is provided within the effect interval between two adjacent material segments. The effect interval covers the connection interval between the two material segments and does not overlap with the intervals where the audio content of the two material segments is located.
[0174] When it is necessary to add a certain transition effect in the connection interval between material segments, it is necessary to first determine whether the corresponding effect interval will overlap with the audio content. If the audio content does not fall within the effect interval, then the transition effect can be enabled in this effect interval.
[0175] If all or part of the audio content falls within the effect interval, that is, the effect interval overlaps with the interval where the audio content is located, then it is necessary to shorten the effect interval of this transition effect, or replace it with a transition effect with a shorter effect interval, so that the effect interval of the selected transition effect does not overlap with the audio content and the audio content does not fall within the effect interval. If the effect of non - overlapping between the effect interval and the audio content cannot be achieved, then the transition effect may not be enabled. By determining the transition effect in this way, it is possible to prevent the audio content from entering the effect interval, resulting in the mismatch between the audio content and the video content in the target video.
[0176] As Figure 7B shown, when material segment 1 contains audio content 4 and material segment 2 contains audio content 5, if it is necessary to set a transition effect between material segment 1 and material segment 2, then the effect interval 1 where the transition effect is located needs to be set within the interval from the end of audio content 4 to the start of audio content 5, to prevent audio content 4 and audio content 5 from falling within the effect interval. If the time length from the end of audio content 4 to the start of audio content 5 is shorter than the shortest time required for the effect interval, then it is necessary to avoid setting a transition effect between material segment 1 and material segment 2.
[0177] Taking the generation of a target video using two material segments as an example, the effect of the target video can be as Figure 7C shown. The target video includes an image part and a sound part. The sound part is the superposition of the audio content and the background dubbing. The part where the volume curve of the audio content is a dotted line indicates that the volume of the corresponding audio content in this part is 0, that is, there is no audio content in this part.
[0178] In some embodiments, the method steps of the foregoing embodiments can be completed by the same device or by different devices collaborating. When each method step is completed by two different collaborating devices, it can be completed by the collaboration of an electronic device and a cloud server, or can be completed by the collaboration of two electronic devices. For example, the foregoing steps 101 and 104, that is, the steps related to obtaining the source video and generating the target video, can be completed by one device, while steps 102 and 103, that is, the steps of determining the preliminary selected segments and determining the highlight segments, can be executed by another device.
[0179] When each method step is completed by the same electronic device, different steps can be completed by different modules located in different layers of the electronic device hierarchical architecture. For example, the foregoing steps 101 and 104, that is, the steps related to obtaining the source video and the steps related to generating the target video, can be completed by the video editing application in the application layer, while steps 102 and 103, that is, the steps of determining the preliminary selected segments and determining the highlight segments, can be completed by the media processing middle platform. Or, step 104, that is, the steps related to generating the target video, can be completed by the video editing application in the application layer, while steps 101 to 103, that is, the steps of obtaining the source video, determining the preliminary selected segments, and determining the highlight segments, can be completed by the media processing middle platform, so as to achieve the background silent processing effect of target video generation.
[0180] Corresponding to the foregoing video processing method, the present application also correspondingly provides an embodiment of a video processing device. The following is an embodiment of the device of the present application, which can be used to execute the method embodiment of the present application. For details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.
[0181] Corresponding to the foregoing embodiment of the device interaction method, the present application also provides an electronic device. Refer to Figure 1A the structural schematic diagram shown, the electronic device includes:
[0182] a processor 110 and a memory, the memory is used to store program instructions;
[0183] The processor 110 is used to call and execute the program instructions stored in the memory. When the program instructions stored in the memory are executed by the processor 1101, the electronic device is enabled to execute Figure 2 all or part of the steps in the corresponding embodiment.
[0184] The processor can be coupled to the transceiver, random access memory and read-only memory through the bus. When the electronic device needs to be operated, it is started by the basic input and output system solidified in the read-only memory or the bootloader boot system in the embedded system to guide the electronic device into a normal operating state. After the electronic device enters the normal operating state, the application program and the operating system are run in the random access memory, so that the electronic device executes Figure 2 All or part of the steps in the corresponding embodiments.
[0185] The electronic device according to the embodiment of the present invention may correspond to the above Figure 2 The electronic device in the corresponding embodiment, and the processor and storage in the electronic device can implement Figure 2 For the sake of brevity, the functions of the electronic device in the corresponding embodiment and / or the various steps and methods implemented are not described in detail here.
[0186] In a specific implementation, the embodiment of the present application further provides a computer storage medium, wherein the computer storage medium stores a computer program or instruction, and when the computer program or instruction is executed, the computer can implement the following steps: Figure 2 All or part of the steps in the corresponding embodiments. The computer-readable storage medium is set in any device, and the arbitrary device may be a random access memory (RAM), and the memory may also include a non-volatile memory (non-volatile memory), such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD) or a solid-state drive (SSD); the memory may also include a combination of the above-mentioned types of memory, etc.
[0187] The present application also provides a chip system, such as Figure 8 As shown, the chip system includes a processor 1201 and an interface circuit 1202. The processor 1201 and the interface circuit 1202 can be one or more. The processor is coupled to a memory and is used to execute a computer program or instruction stored in the memory. When the computer program or instruction is executed, the chip system can implement the following operations: Figure 2 All or part of the steps in the corresponding embodiments. The chip system can be composed of a chip, or can include a chip and other discrete devices.
[0188] In the embodiments of the present application, the various illustrative logical units and circuits can be implemented or operated by a general-purpose processor, a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or a design of any combination of the above. The general-purpose processor can be a microprocessor. Optionally, the general-purpose processor can also be any conventional processor, controller, microcontroller or state machine. The processor can also be implemented by a combination of computing devices, such as a digital signal processor and a microprocessor, multiple microprocessors, one or more microprocessors combined with a digital signal processor core, or any other similar configuration.
[0189] The steps of the methods or algorithms described in the embodiments of the present application can be directly embedded in hardware, software units executed by a processor, or a combination of both. The software units can be stored in a random access memory (RAM), flash memory, read-only memory (ROM), erasable programmable read-only memory (EPROM), registers, hard disk, removable disk, portable compact disc read-only memory (CD-ROM) or any other form of storage medium in the art. Exemplarily, the storage medium can be connected to the processor so that the processor can read information from the storage medium and write information to the storage medium. Optionally, the storage medium can also be integrated into the processor. The processor and the storage medium can be disposed in an ASIC, and the ASIC can be disposed in a user equipment (UE). Optionally, the processor and the storage medium can also be disposed in different components of the UE.
[0190] It should be understood that in the various embodiments of the present application, the magnitudes of the sequence numbers of the various processes do not mean the order of execution. The order of execution of the various processes should be determined by their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0191] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid state disk (SSD)).
[0192] For the same or similar parts among the various embodiments of this specification, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the description in the method embodiment part for the relevant parts.
[0193] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general-purpose hardware platform. Based on this understanding, the technical solution in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a storage medium such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions to enable a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments of the present invention.
[0194] For the same or similar parts among the various embodiments of this specification, reference can be made to each other. In particular, for the embodiments of the road constraint determination device disclosed in the present application, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the description in the method embodiments for the relevant parts.
[0195] It should be noted that those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only regarded as exemplary, and the true scope of the present application is pointed out by the following claims.
[0196] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A video processing method, characterized in that, Applied to an electronic device, the method includes: Obtain at least one piece of source video; Determine the primary selected segments in the source video, where the primary selected segments refer to the video segments in the source video that meet the predetermined screening conditions; Determine the highlight segments corresponding to the primary selected segments according to the analysis result of the sound content included in the primary selected segments, where at least one of the highlight segments includes a primary selected segment containing a sound event that meets the predetermined conditions; Generate a target video using the highlight segments.
2. The method according to claim 1, wherein The obtaining of the source video includes: After receiving a specified operation of the user in a predetermined application, obtain at least one piece of source video selected by the user through the operation process of the specified operation.
3. The method according to claim 1 or 2, characterized in that, Determining the primary selected segments in the source video includes: Select the video segments in the source video that contain sound events as the primary selected segments, where the sound events are obtained by performing sound event detection on the source video; or, Select the video segments in the source video that contain both high-score frame-extracted images and sound events as the primary selected segments, where the high-score frame-extracted images refer to the frame-extracted images with an image score higher than a predetermined score threshold, and the image score is obtained based on a scoring rule predetermined by a predetermined score, and the frame-extracted images refer to the images obtained by frame extraction of the source video.
4. The method according to claim 3, wherein, The primary selected segment contains at least one complete sound event; or, The primary selected segment contains at least one sound event, and the start position of at least one sound event is outside the primary selected segment while the end position is inside the primary selected segment; or, The primary selected segment contains at least one sound event, and the start position of at least one sound event is inside the primary selected segment while the end position is outside the primary selected segment; or, The primary selected segment contains at least one sound event, and the start position and the end position of the sound event are both outside the primary selected segment.
5. The method according to any one of claims 1 to 4, characterized in that The at least one of the highlight segments includes a primary selected segment containing a sound event that meets the predetermined conditions, including: The highlight segment contains a complete sound event; or, The highlight segment contains a sound event starting from the start of the sound event and having a length exceeding a predetermined retention duration.
6. The method according to claim 5, characterized in that The highlight segment contains a complete sound event, including: The start position of the highlight segment in the source video is before the start position of the sound event in the source video, and there is an interval of at least a predetermined interval duration, and the end position of the highlight segment in the source video is after the end position of the sound event in the source video, and there is an interval of at least a predetermined interval duration.
7. The method according to any one of claims 1 to 6, characterized in that, The generating of the target video using the highlight segments includes: Screen, edit or clip the highlight segments to obtain material segments; Generate a target video using the material segments.
8. The method according to claim 7, wherein The generating of the target video using the material segments includes: Determine the retention mode of the audio content in the material segments; Determine the high-value audio content in the material segments, where the high-value content is the audio content that needs to be retained determined based on the retention mode; Generate a target video using the image content and the high-value audio content included in the material segments.
9. The method according to claim 8, wherein The retention mode of the audio content includes: Intelligent retention mode, full retention mode, full non-retention mode, and user-defined mode.
10. The method according to claim 9, characterized in that When the retention mode of the audio content is the intelligent retention mode, determining the high-value audio content in the material segment includes: Filtering out the high-value audio content from the audio content based on the loudness of the audio content; or, Regarding the audio content with a signal-to-noise ratio higher than the signal-to-noise ratio threshold as the high-value audio content; Regarding the audio content with a duration length exceeding the duration length threshold as the high-value audio content.
11. The method according to claim 10, wherein Filtering out the high-value audio content from the audio content based on the loudness includes: Regarding the audio content with a loudness higher than the loudness threshold as the high-value audio content; or, Regarding the audio content with a loudness change exceeding the loudness change threshold as the high-value audio content.
12. The method according to claim 7, characterized in that, Using the image content and the high-value audio content included in the material segment to generate a target video includes: Obtaining background dubbing; Using the image content, the high-value audio content, and the background dubbing to generate a target video, where the image part of the target video is obtained by connecting the image contents of at least two material segments, and the audio part of the target video is obtained by superimposing the high-value audio content of the material segment and the background dubbing.
13. The method according to claim 12, wherein A transition effect for connecting the two material segments is provided in the special effect interval between two adjacent material segments, and the special effect interval covers the connection interval between the two material segments and does not overlap with the interval where the audio content in the two material segments is located.
14. The method according to claim 12, wherein In the audio part, starting from the starting position of the interval corresponding to the audio content, fade out the background dubbing from the first volume value to the second volume value, and fade in the volume of the audio content from the third volume value to the fourth volume value; and before the end position of the interval corresponding to the audio content, fade in the background dubbing from the second volume value to the first volume value, and fade out the volume of the audio content from the fourth volume value to the third volume value; where the first volume value is higher than the second volume value, and the third volume value is lower than the fourth volume value.
15. The method according to any one of claims 1 to 14, wherein The target video contains at least the sound content corresponding to one of the sound events.
16. An electronic device, characterized in that, Includes: A processor and a memory; the memory stores program instructions, and when the program instructions are executed by the processor, the electronic device executes the method according to any one of claims 1-15.
17. A computer storage medium, characterized in that, The computer storage medium stores a computer program or instructions, and when the computer program or instructions are executed, the method according to any one of claims 1-15 is executed.
18. A chip system, characterized in that, The chip system includes a processor, the processor is coupled with the memory, and is used to execute the computer program or instructions stored in the memory, and when the computer program or instructions are executed, the method according to any one of claims 1-15 is executed.
Citation Information
Patent Citations
Detect sports video highlights based on voice recognition
CN105912560A
Video processing method and device and electronic equipment
CN110798735A
Video data processing method and device, electronic equipment and computer storage medium
CN113992970A
Video score processing method, electronic equipment and computer readable storage medium
CN117119266A
Systems and methods for implementing cross-fading, interstitials and other effects downstream
US20140316789A1