Transition position determination method, device, medium, and electronic device

By segmenting and extracting features from the background music in the video, and using a multilayer perceptron to determine the transition position, the adaptiveness problem of determining the transition position in video editing is solved, improving the video transition effect and the creator's work efficiency.

CN116405616BActive Publication Date: 2025-10-17BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310553345.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-05-16
Publication Date
2025-10-17
Estimated Expiration
2043-05-16

AI Technical Summary

Technical Problem

Existing technologies are unable to adaptively determine appropriate transition positions in video clips, which increases the workload of video creators.

Method used

By segmenting the background music of the video to be edited, extracting the sound features of the music slices, and using a multilayer perceptron to determine the transition position based on the sound features, including local and global feature processing, and using a self-attention mechanism to optimize computational complexity.

Benefits of technology

It enables adaptive determination of appropriate transition positions based on background music, reducing the workload of video creators and ensuring the smoothness of video transitions and content matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116405616B_ABST
    Figure CN116405616B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a transition position determination method, device, medium and electronic equipment, belonging to the technical field of electronics, which can adaptively determine a suitable transition position. A transition position determination method comprises: segmenting background music of a video to be edited to obtain music clips; extracting sound features of the music clips; and determining a transition position of the video to be edited based on the sound features.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the field of electronic technology, and in particular, to a transition position determination method and device, medium and electronic equipment. BACKGROUND

[0002] Video clips refer to the formation of a required video by combining background music, video materials, special effects, subtitles, transitions and other contents. However, there is currently no technology that can adaptively determine a suitable transition position. SUMMARY

[0003] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0004] In a first aspect, the present disclosure provides a transition position determination method, comprising: segmenting background music of a video to be clipped to obtain music clips; extracting sound features of the music clips; and determining a transition position of the video to be clipped based on the sound features.

[0005] In a second aspect, the present disclosure provides a transition position determination device, comprising: a segmentation module configured to segment background music of a video to be clipped to obtain music clips; an extraction module configured to extract sound features of the music clips; and a transition determination module configured to determine a transition position of the video to be clipped based on the sound features.

[0006] In a third aspect, the present disclosure provides a computer-readable medium having stored thereon a computer program, which, when executed by a processing device, implements the steps of the method of any one of the first aspect of the present disclosure.

[0007] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having stored thereon a computer program; and a processing device configured to execute the computer program in the storage device to implement the steps of the method of any one of the first aspect of the present disclosure.

[0008] By adopting the above technical solution, since the background music of the video to be clipped is first segmented to obtain music clips, then the sound features of the music clips are extracted, and then the transition position of the video to be clipped is determined based on the sound features, a suitable transition position can be adaptively determined based on the background music of the video to be clipped, thereby reducing the labor and burden of video creators.

[0009] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0010] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:

[0011] Figure 1 is a flowchart of a transition position determination method according to an embodiment of the present disclosure.

[0012] Figure 2 is yet another flowchart of a transition position determination method according to an embodiment of the present disclosure.

[0013] Figure 3 is a flowchart of training a multi-layer perception model according to an embodiment of the present disclosure.

[0014] Figure 4 is a schematic block diagram of a transition position determination apparatus according to an embodiment of the present disclosure.

[0015] Figure 5 shows a structural schematic diagram of an electronic device suitable for implementing embodiments of the present disclosure. DETAILED DESCRIPTION

[0016] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It is understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of the present disclosure.

[0017] It is understood that each of the steps recited in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.

[0018] The term "comprising" and variations thereof as used herein are open-ended, and mean "including but not limited to". The term "based on" means "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related terms are defined in the description that follows.

[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.

[0020] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0021] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0022] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0026] At the same time, it is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0027] Before describing the embodiments of the present disclosure in detail, the meanings of the terms involved in the present disclosure are first explained.

[0028] Background music refers to the music information in the video, which can be a song or dubbing, etc.

[0029] A time unit refers to a basic time slice, for example, 0.1 second. An audio of 10 seconds long can be divided into 100 continuous music slices according to the time unit of 0.1 second. There is no overlapping part between the music slices.

[0030] Sampling frequency: sound is a continuous signal, which needs to be quantized before storage. Quantization includes two aspects, one is to quantize the amplitude of the sound signal, and the other is to quantize the time of the sound signal. The sampling frequency here refers to the quantization of the time of the sound signal, that is, how many sample points are collected per second. For example, the sampling frequency is 44100, which means that the audio signal of one second is equally spaced to collect 44100 sample points.

[0031] Transition refers to a transition effect used to connect two video materials.

[0032] Timestamp refers to the time point of the time occurrence.

[0033] Local context: for a time window K, the local context refers to a segment of length (2K+1) composed of [current music slice, K music slices before the current music slice, K music slices after the current music slice].

[0034] Local feature refers to a feature formed after considering the local context, and the receptive field range is (2K+1) time units.

[0035] Global feature refers to a feature formed after considering the entire background music, and the receptive field range is equal to the length of the entire background music.

[0036] Transition position refers to a suitable transition position determined according to the input background music. These transition positions usually have a strong correlation with the melody of the background music, for example, the transition position will appear at the time point of the rhythm change of the background music, forming a card point video.

[0037] Figure 1 is a flowchart of a transition position determination method according to an embodiment of the disclosure. As shown in Figure 1 The transition position determination method includes the following steps S11-S13.

[0038] In step S11, the background music of the video to be edited is segmented to obtain music slices.

[0039] Video to be edited refers to a video to be edited.

[0040] The background music of the video to be edited refers to the song, voiceover, etc. selected by the user and expected to be used as the background music of the video to be edited.

[0041] In some embodiments, the background music of the video to be edited can be segmented using a first time unit as a unit to obtain music slices. The music slices obtained by segmentation are continuous and have no overlapping parts. The value of the first time unit can be set according to actual conditions, for example, it can be 0.1 seconds or other values. For example, if the background music of the video to be edited lasts for 10 seconds and the size of the first time unit is 0.1 seconds, 100 continuous music slices can be obtained through segmentation, and these music slices have no overlapping parts. Each music slice obtained by segmentation is essentially a shorter audio signal.

[0042] In addition, if the time length of the last music slice in the music slice obtained by segmentation is less than the first time unit, the padding information can be used to pad the time length of the last music slice to the first time unit. For example, the duration of the background music of the video to be edited is 10.05 seconds and the size of the first time unit is 0.1 seconds. The duration of the last music slice obtained by segmentation is 0.05 seconds, which is less than the size of the first time unit, that is, less than 0.1 seconds. In this case, the padding information can be used to pad the duration of the last music slice and pad its duration to 0.1 seconds (that is, the size of the first time unit). The padding information can be set as needed, for example, it can be 0.

[0043] In step S12, the sound features of the music slice are extracted.

[0044] In some embodiments, a music slice can be used as an object to extract cepstral coefficient features (such as Mel frequency cepstral coefficient (MFCC) features) and energy features from the music slice, and the sound signal of the music slice can be mapped to extract the deep learning features of the music slice, for example, by using a linear mapping layer to map the sound signal of the music slice to extract the adaptive deep learning-based features of the music slice. That is, for each music slice, the above three features will be extracted respectively. The cepstral coefficient feature reflects the frequency characteristics of the sound in the music slice, and the energy feature reflects the energy characteristics of the sound in the music slice. These two features contain some representative information of the sound. The deep learning feature reflects other features of the sound in the music slice in addition to the cepstral coefficient feature and the energy feature. These other features are obtained through autonomous learning methods such as deep learning.

[0045] In addition, the cepstrum coefficient features, the energy features and the deep learning features of each music slice can be spliced respectively to obtain the sound features of each music slice. The splicing operation can be to splice the cepstrum coefficient features, the energy features and the deep learning features end to end. For example, the cepstrum coefficient features, the energy features and the deep learning features of the first music slice are spliced to obtain the sound features of the first music slice; the cepstrum coefficient features, the energy features and the deep learning features of the second music slice are spliced to obtain the sound features of the second music slice; and so on, to obtain the sound features of all music slices.

[0046] In step S13, a cut position of the video to be cut is determined based on the sound features.

[0047] In some embodiments, the cut position of the video to be cut can be determined based on the sound features by a multilayer perceptron (MLP).

[0048] In some embodiments, determining the cut position of the video to be cut based on the sound features can be implemented in the following manner.

[0049] First, the sound features of the music slices are processed based on the context of the music slices to obtain local features of the music slices. For example, the similarity between the sound features to be processed and the sound features of the music slices falling within the first time window before and after the music slice corresponding to the sound features to be processed can be determined first, and then the sound features to be processed are processed based on the similarity to obtain the local features of the music slices. In this way, each sound feature is adjusted on the basis of considering the context of each sound feature, and the adjustment speed is very fast. This adjustment method can also be referred to as local attention adjustment.

[0050] For example, the length of the background music of the video to be cut is 10 seconds, the size of the first time unit is 0.1 second, 100 music clips are obtained by segmentation, and 100 sound features are obtained by performing sound feature extraction on the 100 music clips, denoted as {f1, f2, f3,..., f100}. Assuming that the width of the first time window is 5, for the sound feature f6 of the 6th music clip, the context thereof has a total of 11 sound features, i.e., {f1, f2, f3, f4, f5, f6, f7, f8, f9, f10, f11}, that is, 5 sound features before and after the sound feature f6. Then, the local feature of the 6th music clip can be obtained in the following manner: calculating the similarity between the sound feature f6 and each context feature thereof (i.e., the sound features f1, f2, f3, f4, f5, f7, f8, f9, f10, f11), for example, a1 = s(f1, f6), a2 = s(f2, f6),..., a11 = s(f11, f6), where s represents similarity calculation. It is worth noting that a6 = 1, because the similarity of itself and itself is 1; after calculating all the similarities, a new sound feature f6 is obtained by weighted summation, i.e., f6 新 = (a1f1 + a2f2 + a3f3... + a11f11) / (a1 + a2 + a3... + a11), f6 新 which is the local feature of the 6th music clip.

[0051] Then, the local feature of the music clip is subjected to attention processing to obtain the global feature of the music clip. This is actually a global attention, i.e., the range of the receptive field is the entire background music.

[0052] For example, a self-attention mechanism with linear complexity (also known as LinFormer) can be used to perform attention processing on the local feature of the music clip. Through the self-attention mechanism with linear complexity, the computational complexity can be optimized to O(N), i.e., the computation time and spatial complexity are linearly related to the length of the background music.

[0053] Then, the transition position of the video to be cut is determined based on the multi-layer perception information of the global feature of the music clip. For example, the multi-layer perception model can be first corrected (e.g., trained) based on the sample music with transition labels and a preset loss function, then the global feature of the music clip is input into the multi-layer perception model to obtain the multi-layer perception information, and then the transition position of the video to be cut is determined based on the multi-layer perception information.

[0054] Since the receptive field range of the local feature of the music slice is (2*first time window+1) first time units, and the receptive field range of the global feature of the music slice is the entire background music, when determining the transition position of the video to be edited, both the situation near each music slice (which can be considered as looking at the segments before and after the music slice to determine whether it is the most appropriate cut point position here) and the situation of the entire background music (which can be considered as looking at the cut point distribution / style of the entire background music to determine whether a cut point is needed here) are considered, ensuring that a suitable transition position can be determined, so that the final effect and content of the video are more smooth and reasonable, and the situation that the transition position does not match the content or the background music will not occur, for example, in a background music with heavy bass, placing the transition position at the position where the bass switches from high to low or the position where the heavy bass gradually strengthens or the position where the low bass gradually strengthens is more comfortable in terms of sensory experience, and placing it at the position where the high bass is entered or the low bass is entered is very abrupt.

[0055] By adopting the technical solution described above, since the background music of the video to be edited is first segmented to obtain music slices, the sound features of the music slices are then extracted, and then the transition position of the video to be edited is determined based on the sound features, the suitable transition position can be adaptively determined based on the background music of the video to be edited, and the labor and burden of the video creator are reduced. The technical solution according to an embodiment of the present disclosure can also be referred to as an audio beat matching (ABM) scheme.

[0056] Figure 2 is another flowchart of a transition position determination method according to an embodiment of the present disclosure. As shown in Figure 2 , first, the background music x of the video to be edited is segmented in first time units to obtain N music slices x1, x2, x3, … x N . The implementation of segmentation has been described above and will not be described here again. Then, feature extraction is performed on each music slice to extract cepstral coefficient features, energy features, and deep learning features. The implementation of feature extraction has been described above and will not be described here again. Then, the cepstral coefficient features, energy features, and deep learning features of each music slice are spliced. Then, the spliced features are subjected to local context processing. The implementation of local context processing has been described above for the acquisition of local features and will not be described here again. After the local features obtained by local context processing are processed by a self-attention mechanism, they are processed by a multilayer perceptron, and a suitable transition position is output by the multilayer perceptron.

[0057] By adopting the technical solution, since the background music of the video to be edited is first segmented to obtain music clips, then the sound features of the music clips are extracted, then the sound features are locally contextually processed and self-attention mechanism processed, and then a suitable transition position is output by the multilayer perceptron, when determining the transition position of the video to be edited, both the situation of each music clip (which can be considered as looking at the segments before and after the music clip to determine whether it is the most suitable cut point position) and the situation of the whole background music (which can be considered as looking at the cut point distribution / style of the whole background music to determine whether a cut point is needed here) are considered, ensuring that a suitable transition position can be determined, so that the final effect and content transition of the video are more smooth and reasonable, and the situation of mismatching between the transition position and the content or background music does not occur.

[0058] Figure 3 is a flowchart of training a multilayer perception model according to an embodiment of the present disclosure.

[0059] As shown in Figure 3 , first, in step S31, sample music with transition labels is segmented to obtain sample music clips.

[0060] The sample music can be obtained from a sample music dataset. The sample music in the sample music dataset is all labeled with transition positions, that is, each data in the sample music dataset contains an original background music audio file and a corresponding transition position timestamp.

[0061] The implementation of segmentation has been described in the foregoing, and will not be repeated here.

[0062] In step S32, sample sound features of the sample music clips are extracted. The implementation of sound feature extraction has been described in the foregoing, and will not be repeated here.

[0063] In step S33, the transition position is predicted based on the sample sound features.

[0064] For example, the sample sound features can be locally contextually processed and self-attention mechanism processed, and then the transition position is predicted based on the processed features. The local contextual processing and self-attention mechanism processing have been described in detail in the foregoing, and will not be repeated here.

[0065] In step S34, the loss between the predicted transition position and the transition label is calculated based on a preset loss function.

[0066] When constructing the preset loss function, two main problems are considered:

[0067] (1) How to balance positive and negative samples. The number of transitions in a video is very small, which is the number of positive samples. For example, a 10-second video may contain 3 transitions. If the background music of the video is divided into 100 segments according to a time unit of 0.1 seconds, then 3 of the 100 segments contain the start time point of the transition, which are positive samples, and the remaining 97 segments are negative samples. This imbalance can seriously affect the performance of the multi-layer perception model.

[0068] (2) How to measure the importance of each label position. For example, the transition position is labeled with 1, and the non-transition position is labeled with 0. A data label containing 6 time units can be [1, 0, 0, 0, 0, 0]. When the multi-layer perception model predicts the transition position, the output [0, 1, 0, 0, 0, 0] is obviously better than the output [0, 0, 0, 0, 0, 1], but the traditional cross-entropy loss cannot measure this difference. Moreover, the closer the position to the transition label 1, the more important it should be to be predicted correctly, and the farther the position from the transition label 1, the less constraint it should be subjected to.

[0069] To solve the above two problems, the loss between the predicted transition position and the transition label can be calculated based on the preset loss function in the following way.

[0070] First, the influence of the transition label belonging to the positive sample is calculated. The influence range of the transition label belonging to the positive sample can be in the form of a Gaussian distribution. In addition, in order to ensure the certainty of the numerical value of the influence of the transition label, it can be normalized to ensure that the numerical value of the influence of the transition label is limited between 0 and 1.

[0071] Then, based on the influence of the transition label and the value of the transition label, the importance of the position of the transition label is calculated. If the transition label is a positive sample, the value of the transition label is 1, and if the transition label is a negative sample, the value of the transition label is 0, that is, only the position of the label 1 will have an influence.

[0072] Then, based on the predicted transition position, the transition label, and the importance of the position of the transition label, the loss between the predicted transition position and the transition label is calculated. In addition, in order to balance the positive and negative samples, the window width of the Gaussian distribution can be estimated according to the sample music data set, so that the sum of the label weights of the positive samples and the sum of the label weights of the negative samples are approximately equal when calculating the loss.

[0073] In step S35, the multi-layer perception model is trained based on the calculated loss to obtain a trained multi-layer perception model.

[0074] By adopting the above technical solutions, the training of the multi-layer perception model can be realized.

[0075] Figure 4 is a schematic block diagram of a transition position determination apparatus according to an embodiment of the present disclosure. As shown in the figure, the transition position determination apparatus comprises a segmentation module 41 configured to segment background music of a video to be edited to obtain music clips; an extraction module 42 configured to extract sound features of the music clips; and a transition determination module 43 configured to determine a transition position of the video to be edited based on the sound features. Figure 4

[0076] By adopting the above technical solution, since the background music of the video to be edited is first segmented to obtain music clips, then the sound features of the music clips are extracted, and then the transition position of the video to be edited is determined based on the sound features, the appropriate transition position can be adaptively determined based on the background music of the video to be edited, and the labor and burden of the video creator are reduced.

[0077] Optionally, the segmentation module 41 segments the background music of the video to be edited, comprising: segmenting the background music of the video to be edited in a first time unit to obtain the music clips; and if the time length of the last music clip in the music clips is less than the first time unit, the time length of the last music clip is padded to the first time unit by using padding information.

[0078] Optionally, the extraction module 42 extracts the sound features of the music clips, comprising: extracting cepstrum coefficient features and energy features from the music clips, and mapping the sound signal of the music clips to extract deep learning features of the music clips.

[0079] Optionally, the transition determination module 43 determines the transition position of the video to be edited based on the sound features, comprising: processing the sound features of the music clips based on the context of the music clips to obtain local features of the music clips; performing attention processing on the local features to obtain global features of the music clips; and determining the transition position of the video to be edited based on multi-layer perception information of the global features.

[0080] Optionally, the transition determination module 43 processes the sound features of the music clips based on the context of the music clips to obtain local features of the music clips, comprising: determining the similarity between the sound features of the music clips to be processed and the sound features of the music clips falling within a first time window before and after the music clip corresponding to the sound features to be processed; and processing the sound features to be processed based on the similarity to obtain the local features of the music clips.

[0081] ​Optionally, the transition determination module 43 determines the transition position of the video to be edited based on the multi-layer perception information of the global feature, including: correcting a multi-layer perception model based on sample music with a transition label and a preset loss function; inputting the global feature into the multi-layer perception model to obtain the multi-layer perception information; and determining the transition position of the video to be edited based on the multi-layer perception information.

[0082] Optionally, the correcting the multi-layer perception model based on sample music with a transition label and a preset loss function includes: segmenting sample music with a transition label to obtain sample music slices; extracting sample sound features of the sample music slices; predicting a transition position based on the sample sound features; calculating a loss between the predicted transition position and the transition label based on a preset loss function; and correcting the multi-layer perception model based on the calculated loss.

[0083] The embodiments of the present disclosure further provide a computer readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of any of the methods of the present disclosure.

[0084] The embodiments of the present disclosure further provide an electronic device, including: a storage device having a computer program stored thereon; and a processing device configured to execute the computer program in the storage device to implement the steps of any of the methods of the present disclosure.

[0085] Reference is made below to Figure 5 , which shows a structural schematic diagram of an electronic device 600 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure can include, but is not limited to, mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Personal Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), and the like, as well as fixed terminal devices such as digital TVs, desktop computers, and the like. Figure 5 The electronic device shown is merely an example and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.

[0086] As shown in Figure 5 , the electronic device 600 can include a processing device (such as a central processor, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 602 or loaded into a random access memory (RAM) 603 from a storage device 608. In the RAM 603, various programs and data required for the operation of the electronic device 600 are also stored. The processing device 601, the ROM 602, and the RAM 603 are connected to each other through a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0087] In general, the following devices can be connected to the I / O interface 605: input devices 606 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; output devices 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; storage devices 608 including, for example, a magnetic tape, a hard disk, and the like; and communication devices 609. The communication devices 609 can allow the electronic device 600 to communicate wirelessly or wired with other devices to exchange data. Although Figure 5 The electronic device 600 is shown with various devices, but it is understood that all of the shown devices are not required to be implemented or present. More or less devices can alternatively be implemented or present.

[0088] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network through the communication devices 609, or installed from the storage devices 608, or installed from the ROM 602. When the computer program is executed by the processing devices 601, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0089] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.

[0090] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.

[0091] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and is not assembled into the electronic device.

[0092] The computer readable medium described above carries one or more programs, which when executed by the electronic device, cause the electronic device to: segment background music of a video to be edited to obtain music clips; extract sound features of the music clips; and determine a transition position of the video to be edited based on the sound features.

[0093] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages, including an object-oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0094] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks noted in succession can in fact be executed substantially concurrently or in the opposite order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0095] The modules described in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of a module does not constitute a limitation on the module itself, for example, an extracting module can also be described as a "module for extracting sound features of the music clips".

[0096] The functionality described herein above can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, an example type of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0097] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0098] According to one or more embodiments of the present disclosure, example 1 provides a transition position determination method, comprising: segmenting background music of a video to be edited to obtain music clips; extracting sound features of the music clips; and determining a transition position of the video to be edited based on the sound features.

[0099] According to one or more embodiments of the present disclosure, example 2 provides the method of example 1, wherein the segmenting the background music of the video to be edited comprises: segmenting the background music of the video to be edited in a first time unit to obtain the music clips; and if a time length of a last music clip in the music clips is less than the first time unit, padding the time length of the last music clip to the first time unit using padding information.

[0100] According to one or more embodiments of the present disclosure, example 3 provides the method of example 1, wherein the extracting the sound features of the music clips comprises: extracting cepstrum coefficient features and energy features from the music clips, and mapping a sound signal of the music clips to extract deep learning features of the music clips.

[0101] According to one or more embodiments of the present disclosure, example 4 provides the method of any one of examples 1 to 3, wherein the determining the cut position of the video to be cut based on the sound features comprises: processing the sound features of the music slice based on a context of the music slice to obtain local features of the music slice; performing attention processing on the local features to obtain global features of the music slice; determining the cut position of the video to be cut based on multi-layer perception information of the global features.

[0102] According to one or more embodiments of the present disclosure, example 5 provides the method of example 4, wherein the processing the sound features of the music slice based on a context of the music slice to obtain local features of the music slice comprises: determining a similarity between the sound features to be processed and sound features of music slices falling within a first time window before and after the music slice corresponding to the sound features to be processed; processing the sound features to be processed based on the similarity to obtain the local features of the music slice.

[0103] According to one or more embodiments of the present disclosure, example 6 provides the method of example 4, wherein the determining the cut position of the video to be cut based on multi-layer perception information of the global features comprises: correcting a multi-layer perception model based on sample music with a cut label and a preset loss function; inputting the global features into the multi-layer perception model to obtain the multi-layer perception information; determining the cut position of the video to be cut based on the multi-layer perception information.

[0104] According to one or more embodiments of the present disclosure, example 7 provides the method of example 6, wherein the correcting the multi-layer perception model based on sample music with a cut label and a preset loss function comprises: segmenting the sample music with the cut label to obtain sample music slices; extracting sample sound features of the sample music slices; predicting a cut position based on the sample sound features; calculating a loss between the predicted cut position and the cut label based on the preset loss function; correcting the multi-layer perception model based on the calculated loss.

[0105] According to one or more embodiments of the present disclosure, example 8 provides the method of example 7, wherein the calculating a loss between the predicted cut position and the cut label based on the preset loss function comprises: calculating an influence of the cut label belonging to a positive sample; calculating a degree of importance of a position of the cut label based on the influence and a value of the cut label; calculating a loss between the predicted cut position and the cut label based on the predicted cut position, the cut label, and the degree of importance of the position of the cut label.

[0106] According to one or more embodiments of the present disclosure, example 9 provides a transition position determination apparatus, comprising: a segmentation module configured to segment background music of a video to be edited to obtain music clips; an extraction module configured to extract sound features of the music clips; and a transition determination module configured to determine a transition position of the video to be edited based on the sound features.

[0107] According to one or more embodiments of the present disclosure, example 10 provides a computer readable medium having stored thereon a computer program which, when executed by a processing apparatus, implements the steps of the method of any one of examples 1-8.

[0108] According to one or more embodiments of the present disclosure, example 11 provides an electronic device, comprising: a storage apparatus having stored thereon a computer program; and a processing apparatus configured to execute the computer program in the storage apparatus to implement the steps of the method of any one of examples 1-8.

[0109] The above description merely illustrates the preferred embodiments of the present disclosure and a principle of applied technologies. It should be understood by those skilled in the art that the disclosure range involved in the present disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by any combinations of the technical features described above or their equivalent features without departing from the above disclosed concept. For example, the technical solutions formed by replacing the above described features with the technical features disclosed in the present disclosure (but not limited to) having similar functions.

[0110] In addition, although each operation is depicted in a particular order, this should not be understood as requiring the operations to be performed in the particular order shown or in sequential order. Under certain circumstances, multitasking and parallel processing can be advantageous. Similarly, although several implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments can also be combined in a single embodiment. Conversely, various features described in the context of a single embodiment can also be separated and implemented in multiple embodiments.

[0111] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims. With respect to the devices in the above-described embodiments, the specific manner in which the various modules perform operations has been described in detail in the embodiments related to the method, and will not be described here in detail.

Claims

1. A method for determining a transition position, characterized in that: include: Segment the background music of the video to be edited to obtain music slices; Extracting sound features of the music slice; Determining a transition position of the video to be edited based on the sound features; The step of determining the transition position of the video to be edited based on the sound feature includes: processing the sound features of the music slice based on the context of the music slice to obtain local features of the music slice; Performing attention processing on the local features to obtain global features of the music slice; The transition position of the video to be edited is determined based on the multi-layer perception information of the global features.

2. The method according to claim 1, characterized in that The segmentation of the background music of the video to be edited includes: Segmenting the background music of the video to be edited based on the first time unit to obtain the music slices; If the time length of the last music slice in the music slices is less than the first time unit, the time length of the last music slice is padded to the first time unit using padding information.

3. The method according to claim 1, characterized in that The extracting the sound features of the music slice comprises: extracting cepstral coefficient features and energy features from the music slice; The sound signal of the music slice is mapped to extract the deep learning features of the music slice.

4. The method according to claim 1, wherein The processing of the sound features of the music slice based on the context of the music slice to obtain the local features of the music slice includes: Determining similarities between the sound feature to be processed and sound features of music slices falling within a first time window before and after the music slice corresponding to the sound feature to be processed; Based on the similarity, the sound features to be processed are processed to obtain local features of the music slice.

5. The method according to claim 1, wherein The determining the transition position of the video to be edited based on the multi-layer perception information of the global feature includes: Calibrate the multi-layer perception model based on sample music with transition labels and a preset loss function; Inputting the global features into the multi-layer perception model to obtain the multi-layer perception information; A transition position of the video to be edited is determined based on the multi-layer perception information.

6. The method according to claim 5, characterized in that The method of correcting the multi-layer perception model based on the sample music with transition labels and the preset loss function includes: Segmenting the sample music with the transition label to obtain sample music slices; Extracting sample sound features of the sample music slice; predicting a transition position based on the sample sound features; Calculating the loss between the predicted transition position and the transition label based on the preset loss function; The multi-layer perception model is calibrated based on the calculated loss.

7. The method according to claim 6, characterized in that The calculating the loss between the predicted transition position and the transition label based on the preset loss function includes: Calculate the influence of the transition label belonging to the positive sample; Calculating the importance of the position of the transition label based on the influence and the value of the transition label; Based on the predicted transition position, the transition label, and the importance of the position of the transition label, a loss between the predicted transition position and the transition label is calculated.

8. A device for determining a transition position, characterized in that: include: A segmentation module is used to segment the background music of the video to be edited into music slices; An extraction module, configured to extract the sound features of the music slice; A transition determination module, configured to determine a transition position of the video to be edited based on the sound features; a transition determination module, further configured to process the sound features of the music slice based on the context of the music slice to obtain local features of the music slice; Performing attention processing on the local features to obtain global features of the music slice; The transition position of the video to be edited is determined based on the multi-layer perception information of the global features.

9. A computer-readable medium having a computer program stored thereon, characterized in that: When the program is executed by a processing device, the steps of the method according to any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: include: a storage device having a computer program stored thereon; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Video scene change point determination and video clip generation method and system

    CN114222159A