Video generation method, device, electronic device and readable storage medium
By extracting the target video paragraphs and background music from the user's video materials and matching the feature information of video and audio, the problem of video in the existing video generation method is solved, and the rhythm matching between video and audio is achieved, enhancing the user experience and memory effect.
Patent Information
- Application Number
- JP2024539569
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-31
- Filing Date
- 2022-12-29
- Publication Date
- 2025-05-08
- Estimated Expiration
- 2042-12-29
AI Technical Summary
The existing video generation methods have limitations when selecting video paragraphs, which leads to the inability to generate videos, poor user experience, and it is difficult to effectively recall deep memories of the past.
By extracting the target video paragraphs and background music from the user's video material, and matching the characteristic information of the video and audio, a target video with a specific rhythm and content matching is generated.
The rhythm matching of video and audio is achieved, which enhances the vividness and user experience of video, allowing users to better recall the deep memories of the past.
Smart Images

Figure 0007673331000001 
Figure 0007673331000002 
Figure 0007673331000003
Abstract
Description
[Technical field]
[0001] [CROSS-REFERENCE TO RELATED APPLICATIONS] This application claims priority to a Chinese patent application filed on December 31, 2021, bearing application number 202111672927.6 and entitled "Video Generation Method, Apparatus, Electronic Device, and Readable Storage Medium," the entire contents of which are incorporated herein by reference.
[0002] [Technical field] The present disclosure relates to the field of video processing technology, and in particular to a video generating method, device, electronic device and readable storage medium. [Background technology]
[0003] A user can generate one video by integrating videos posted within a certain time period in the past. The user may be reminded of some deep memories while watching the generated video. Therefore, videos with review themes are deeply liked by users. Summary of the Invention
[0004] The present disclosure provides a video generation method, apparatus, electronic device, and readable storage medium. According to a first aspect, the present disclosure provides a method for producing a method for detecting a pulmonary circulation disorder, comprising: obtaining a plurality of video footage from a set of original video footage that includes video relating to a user; obtaining a target audio material to be used as background music; for each of the video materials, performing image feature extraction from each video frame of the video material respectively, and performing a segmentation process based on the image feature information respectively corresponding to each of the video frames of each of the video materials to obtain a target video segment corresponding to the video material; and integrating target video segments and the target audio material corresponding to each of the video materials to generate a target video including a plurality of video segments each obtained based on the plurality of target video segments, wherein the plurality of video segments in the target video are played in chronological order of posting, and durations of the plurality of video segments match durations of corresponding musical phrases in the target audio material.
[0005] In one possible embodiment, the step of obtaining target audio material to be used as background music comprises: The method includes determining a target audio material from a set of predetermined audio materials based on audio features of each audio material in the set of predetermined audio materials, beat information of the audio materials, and a duration of each musical phrase of a particular audio segment in the audio material, the particular audio segment in the target audio material being background music of the target video.
[0006] In one possible embodiment, the step of determining target audio material from a predetermined set of audio material based on audio features of each audio material in the set, on beat information of the audio material, and on durations of each musical phrase of a particular audio segment in the audio material, comprises: - filtering out a number of music materials included in the set of predefined music materials based on the predefined set of music features to obtain a first set of candidate audio material; - eliminating each audio material included in the first set of candidate audio material based on a predefined audio beat to obtain a second set of candidate audio material; determining the target audio material based on audio material from the second set of candidate audio material, the duration of each musical phrase contained in a particular audio segment satisfying a predetermined duration condition.
[0007] In one possible embodiment, if there is no audio material in the second set of candidate audio materials, the duration of the musical phrases contained in the particular audio segment satisfies a predefined duration condition, the method further comprises: The method further includes a step of matching the audio features corresponding to each audio material of a pre-specified audio material set based on the user's preferences, and if the matching is successful, determining the target audio material based on the audio material with which the matching is successful.
[0008] In one possible embodiment, the method comprises: performing a weighting calculation for each of said video footage based on image feature information respectively corresponding to each of said video frames of said video footage to obtain an evaluation result respectively corresponding to each of said video frames; extracting a target video frame from each of the video frames of the video material to include a plurality of target video frames extracted from the plurality of video materials based on the evaluation results respectively corresponding to each of the video frames, to obtain a set of video frames for generating an opening and / or an ending of the target video; a step of generating a target video by integrating target video segments and the target audio material respectively corresponding to each of the video materials, the step comprising: The step of aggregating target video segments respectively corresponding to each of the video materials, the set of video frames and the target audio material to generate the target video including an opening and / or an ending generated from the set of video frames.
[0009] In one possible embodiment, the step of performing a segmentation process based on image feature information corresponding to each video frame in each of the video materials to obtain a target video segment includes: image feature information corresponding to each video frame of the video footage, the duration of a corresponding musical phrase in the target audio segment, and the original audio of the video footage Sentence Based on the segmentation result, a segmentation process is performed on the video material to obtain the target video segment.
[0010] In one possible embodiment, the target audio segment is one or more complete subsequences of the corresponding original audio. Sentence Includes.
[0011] According to a second aspect, the present disclosure provides a method for producing a method for detecting a pulmonary circulation, comprising: a video processing module for obtaining a plurality of video footage from a set of original video footage including video relating to a user; an audio processing module for obtaining a target audio material to be used as background music, the video processing module further includes an audio processing module, which is used for performing, for each of the video materials, image feature extraction from each video frame of the video materials respectively, and performing a segmentation process based on image feature information respectively corresponding to each of the video frames of each of the video materials to obtain a target video segment corresponding to the video materials; a video integration module for integrating the target video segments and the target audio material to generate a target video including a plurality of video segments each derived based on the plurality of target video segments, wherein the video segments in the target video are played in chronological order, and durations of the video segments match durations of corresponding musical phrases in the target audio material; A video generating device is provided.
[0012] According to a third aspect, the present disclosure provides an electronic device including a memory and a processor, the memory is configured to store computer program instructions; The processor is configured to execute the computer program instructions to cause the electronic device to implement a video generation method according to any one of the first aspects.
[0013] According to a fourth aspect, the present disclosure provides a readable storage medium comprising computer program instructions which, when executed by at least one processor of an electronic device, cause the electronic device to implement a video generation method according to any one of the first aspects.
[0014] According to a fifth aspect, the present disclosure provides a computer program product which, when executed by a computer, causes the computer to implement the video generation method according to any one of the first aspects.
[0015] According to a sixth aspect, the present disclosure provides a computer program which, when executed by a computer, causes the computer to implement the video generation method according to any one of the first aspects. [Brief description of the drawings]
[0016] The drawings within this specification are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure, and together with the specification serve to explain the principles of the present disclosure.
[0017] In order to more clearly describe the technical solutions in the embodiments or prior art of the present disclosure, the following description will briefly describe the drawings that need to be used in the description of the embodiments or prior art. Obviously, a person skilled in the art can obtain other drawings based on these drawings without creative efforts.
[0018] [Figure 1] FIG. 2 is a diagram showing an application scene of a video generation method according to an embodiment of the present disclosure. [Diagram 2]FIG. 2 illustrates a flowchart of a video generation method according to one embodiment of the present disclosure. [Diagram 3] FIG. 2 illustrates a flowchart of a video generation method according to another embodiment of the present disclosure. [Figure 4] FIG. 2 illustrates a flowchart of a video generation method according to another embodiment of the present disclosure. [Diagram 5] FIG. 2 illustrates a flowchart of a video generation method according to another embodiment of the present disclosure. [Figure 6] FIG. 1 is a schematic diagram of a video generating device according to an embodiment of the present disclosure. [Figure 7] FIG. 1 is a schematic diagram illustrating a configuration of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0019] In order to make the above objects, features and advantages of the present disclosure more clearly understandable, the following further describes the aspects of the present disclosure. What should be described is that, unless contradictory, the embodiments and features in the embodiments of the present disclosure can be combined with each other.
[0020] In order to fully understand the present disclosure, many specific details are described in the following description, but the present disclosure may be implemented in other forms different from those described herein. Apparently, the embodiments in the specification are only some embodiments of the present disclosure, but not all embodiments.
[0021] Currently, a video analysis algorithm is generally used to select some video segments from videos posted by users within a certain time period, and these selected video segments are integrated to generate a review-themed video. However, since the above process is limited to the selection method of how to select the video segments, the generated video as a whole is not vivid enough, and the feeling given to the user falls short of expectations, so how to generate a more vivid and excellent review-themed video is an urgent problem to be solved.
[0022] The present disclosure provides a video generating method, device, electronic device, and readable storage medium, which includes selecting a plurality of video materials from video materials posted by a user within a certain time period, extracting several target video segments from the selected plurality of video materials, and integrating the extracted target video segments with a determined target audio segment to generate a target video corresponding to the certain time period, thereby allowing the user to look back on deep memories within the certain time period through the target video. In addition, the duration of each video segment in the target video matches the duration of a musical phrase in the adopted target audio segment, so that the content rhythm and audio rhythm of the target video match, giving the user a unique experience.
[0023] 1 is a diagram showing a scene of a video generating method according to the present disclosure. The scene 100 shown in FIG. 1 includes a terminal device 101 and a server-side device 102, in which a client is installed, and the client can communicate with the server-side device 102 via the terminal device 101.
[0024] Here, the client may present an entry for acquiring a target video to a user through the terminal device 101, and generate an operation command based on a trigger operation on the entry by the user. Then, the client may generate a video acquisition request according to the operation command. The client may transmit a video acquisition request to the server-side device through the terminal device 101. In response to the video acquisition request transmitted by the client through the terminal device 101, the server-side device 102 may deliver a target video corresponding to a predetermined time period to the client. Upon receiving the target video, the client may load the target video, display a video editing page through the terminal device 101, and play and present the target video on the video editing page. In addition, the user may make some adjustments to the video through the video editing page, for example, replace the soundtrack of the target video.
[0025] The server-side device 102 may select multiple video materials from multiple videos related to the user, extract several target video segments from the selected multiple video materials, determine target audio materials to be used as background music, and integrate the extracted target video segments and the target audio materials to generate a target video. Since the server-side device 102 can pre-generate the target video offline, when the server-side device 102 receives a video acquisition request sent by a client, the server-side device 102 can quickly respond to the video acquisition request to ensure the user experience.
[0026] What should be described is that, in addition to the scene shown in FIG. 1 above, the video generating method according to the present disclosure may be executed locally on the terminal device. For example, the client may use the resources of the terminal device to pre-select multiple video materials from a video related to a user, extract some target video segments from the selected multiple video materials, and cache the extracted target video segments and the determined target audio segment locally on the terminal device. When the client receives a video acquisition request input by a user, the client integrates the multiple target video segments and the target audio material to generate a target video. In this way, the client can process the video material and the audio material by using the resources of the terminal device during the idle time period. Here, the target audio material may be obtained from a server-side device in advance by the terminal device, or may be obtained from the server-side device when the terminal device receives a video acquisition request input by a user.
[0027] Of course, the video generating method according to the present disclosure may be realized in other forms, and the details will not be described here.
[0028] Exemplarily, the video generation method according to the present disclosure may be performed by a video generation device according to the present disclosure, which may be realized in any software and / or hardware form. Exemplarily, the video generation device may be a mobile phone, a tablet computer, a notebook computer, a palmtop computer, an in-vehicle terminal, a wearable device, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), a personal computer (PC), a server, a cloud server, a server cluster, etc. The embodiments of the present disclosure are not specifically limited to a specific type of the video generation device.
[0029] In the following embodiment, a detailed description will be given of an example in which an electronic device executes a video generating method.
[0030] 2 is a flowchart of a video generating method according to an embodiment of the present disclosure. As shown in FIG. 2, the method of the present embodiment includes the following steps:
[0031] S201: A plurality of video materials are obtained from an original video material set.
[0032] Here, the original video material set may include videos related to the user, and the multiple video materials selected from the original video material set are video materials for generating a target video corresponding to the user. Here, the videos related to the user may include, but are not limited to, videos posted by the user, videos edited but not posted by the user (e.g., videos stored in drafts), videos exported by the user, etc. Or, the videos related to the user may be further related to time. That is, the original video material set may include videos related to the user that meet a specific time condition. For example, the original video material set may include videos posted by the user in the past year, videos shot in the past year, videos clipped in the past year, etc.
[0033] In the following embodiment, an example will be described in which the original video material set includes videos posted within a specific time period.
[0034] Assuming that the electronic device is a server-side device, the server-side device may obtain the video material posted by the user during a predetermined time period from a database for storing the video material posted by the user. Assuming that the electronic device is a terminal device, in general, the client may cache the video material posted by the user, the unposted video material in the draft, the video material downloaded by the user, etc. in a local memory space of the terminal device. Thus, the electronic device may further obtain the video material posted by the user during a predetermined time period from the local memory space of the terminal device.
[0035] The present disclosure does not limit the specific time length of the predetermined time period. For example, the predetermined time period may be one year, half a year, one quarter, etc.
[0036] Here, the plurality of video materials determined by the electronic device are video materials that satisfy a certain condition among the plurality of video materials posted by the user during a certain time period, and the selected video materials may include meaningful memories of the user during the certain time period. Here, the certain condition for selecting the video materials may be associated with information in a first dimension of the video materials. The first dimension may include duration, plays, likes, comments, favorites, forwards, downloads, sharing, etc.
[0037] As a possible embodiment, the electronic device may first analyze the duration of all video materials posted by users within a specified time period, determine video materials whose total duration is greater than a first specified duration as candidate video materials, obtain a set of candidate video materials, and then perform a weighting calculation on information in one or more first dimensions, such as plays, likes, comments, favorites, forwards, downloads, and sharing, for each candidate video material in the set of candidate video materials to obtain an overall score corresponding to each candidate video material, sort the candidate video materials based on the overall score corresponding to each candidate video material, and determine multiple candidate video materials whose overall scores meet a specified condition as video materials for generating a target video corresponding to the specified time period.
[0038] For example, for each video material posted in a certain time period, obtain the number of likes X1, the number of comments X2, the number of favorites X3, and the number of forwards X4 of the video material. Here, the weight corresponding to the like dimension is s1, the weight corresponding to the comment dimension is s2, the weight corresponding to the favorite dimension is s3, and the weight corresponding to the forward dimension is s4. A weighting calculation is performed based on the information and weight corresponding to each dimension to obtain a weighting calculation result corresponding to the video material. The above weighting calculation may be expressed as P=s1*X1+s2*X2+s3*X3+s4*X4 by the formula, where P represents the weighting calculation result for the video material.
[0039] Here, a higher weighting calculation indicates a higher likelihood that the posted video meets the specific criteria for selecting video material, and a lower weighting calculation indicates a lower likelihood that the posted video meets the specific criteria for selecting video material.
[0040] It should be noted that the purpose of selecting the video material posted by the user in a certain time period based on the first predetermined duration before performing the weighting calculation is to ensure that the video material with strong content is obtained. If the duration of the video material is too short and the video content is relatively small, it may not be possible to divide the effective video segment for generating the target video, and it is also not possible to match the duration of the music phrase of the soundtrack.
[0041] The first predetermined length of time may be determined based on a minimum length of time corresponding to the target video segment. For example, if the minimum length of time of the target video segment is 2 seconds, the first predetermined length of time may be greater than 2 seconds. For example, the first predetermined length of time may be 3 seconds, 4 seconds, etc.
[0042] It should be noted that the present disclosure does not limit the number of video materials for generating a target video corresponding to a predetermined time period. For example, the number of video materials for generating a target video corresponding to a predetermined time period may be a fixed numerical value. Or, the number of video materials for generating a target video corresponding to a predetermined time period may be any one numerical value within a predetermined numerical range, and may be determined based on the total number of video materials posted by a user during a predetermined time period. Exemplarily, the number of video materials for generating a target video corresponding to a predetermined time period may be set to 5, 10, etc. in the electronic device. Or, the number of video materials for generating a target video corresponding to a predetermined time period may be set to any number within [3, 10].
[0043] In addition, when selecting multiple video materials from the original video material set, the electronic device may first exclude other people's videos transferred by the user, videos set to privacy status by the user, videos deleted by the user after posting, advertisement sales, and other types of videos, and then filter them based on the time length of the video materials and information on the first dimension of the video materials, thereby protecting the user's privacy and ensuring that the content of the target video corresponding to the predetermined time period generated is closely related to the user's memories during the predetermined time period.
[0044] S202: Obtain a target audio material to be used as background music.
[0045] In some cases, if the duration of the target audio material, audio characteristics in dimensions other than the duration of the target audio material, etc., meet the requirements of the present application for background music of the target video, no special processing is required for the target audio material. In other cases, if the duration of the target audio material, audio characteristics of some audio segments in the target audio material, etc., do not meet the requirements of the present application for background music of the target video, target audio segments that meet the conditions may be extracted from the target audio material and used as background music of the target video.
[0046] When the target audio segment in the target audio material is the background music of the target video, the target audio segment may be one specific audio segment in the target audio material determined by the electronic device through a comprehensive analysis of the audio features in the second dimension of each audio material in a predetermined set of audio materials.
[0047] The predetermined set of audio materials may include audio materials added to favorites by the user, audio materials used when posting videos by the user, audio materials in the audio library, etc. The present disclosure does not limit parameters such as the number of audio materials included in the predetermined set of audio materials, the duration of the audio materials, the storage format of the audio materials, etc.
[0048] The second dimension may include one or more dimensions such as rhythm, audio style, audio mood, audio scene, duration of musical phrases, number of videos posted with audio, etc. Of course, the third dimension may further include other dimensional features of the audio material. The present disclosure is not limited to which dimensions the audio features specifically include.
[0049] The target audio material and the target audio segment may also be audio determined by the electronic device in other ways. The following will provide a detailed description of how to determine the target audio material and the target audio segment in the embodiment shown in FIG.
[0050] S203: For each of the video materials, perform image feature extraction on each video frame of the video material respectively, and perform a segmentation process based on the image feature information respectively corresponding to each of the video frames of each of the video materials to obtain a target video segment corresponding to the video material.
[0051] The video footage in this step is the video footage obtained from the original video footage set in step S201.
[0052] The electronic device performs feature extraction in a third dimension for each video frame of a selected plurality of video materials using a pre-trained video processing model, obtains image feature information for each video frame in the video materials output by the video processing model, obtains an evaluation result corresponding to the video frame based on the image feature information corresponding to each video frame in the video materials, and calculates the evaluation results respectively corresponding to each video frame and the original audio of the video materials. Sentence Based on the segmentation result and the duration of the corresponding musical phrase in the target audio segment, a position in the video material of the target video segment to be segmented may be determined and further segmented to obtain the target video segment.
[0053] Here, the above third dimension may include one or more dimensions of art style, image scene, image theme, image mood, image person relationship, image saliency feature, etc. It should be noted that the present disclosure does not limit the type of the video processing model. For example, the video processing model may be a neural network model, a convolution model, etc.
[0054] Here, the evaluation results corresponding to each video frame, the original audio of the video material, Sentence Obtaining the target video segment based on the segmentation result and the time length of the corresponding musical phrase in the target audio segment may adopt any of the following embodiments, but is not limited thereto.
[0055] In one possible embodiment, assuming that the evaluation result is a numerical value, the electronic device analyzes a video frame range in which video frames having a higher numerical value in the evaluation result are relatively concentrated, and the video frame range may be set to be equal to or greater than a maximum time length of each video segment in the target video to be generated, and a relative time to the original audio of the video material is calculated. SentenceAccording to the segmentation result, a cut is performed within the determined video frame range to obtain a target video segment, and one or more complete audio segments corresponding to the target video segment are included in the original audio segment. Sentence It may be so arranged.
[0056] In another possible embodiment, assuming that the evaluation result is a numerical value, the electronic device further performs a search before and after the position where the video frame with the highest evaluation result is located based on the position where the video frame with the highest evaluation result is located to obtain a target video segment, and detects whether one or more complete audio segments are included in the original audio segment corresponding to the target video segment. Sentence It may be so arranged.
[0057] It should be noted that the duration of the target video segment determined by any of the above methods may be less than the duration of the corresponding musical phrase in the target audio segment. For illustrative purposes, assume that the above first embodiment is currently performed on video material 2. The video frame range in which the evaluation result value determined based on the evaluation result corresponding to the video frames is relatively high is from frame 15 to frame 26 of the video material, which is played at a rate of 3 frames / second and has a playback time length of 4 seconds. In the original audio corresponding to video material 2, the original audio segment corresponding to frame 15 to frame 26 is two complete frames. Sentence The first Sentence and the second Sentence The duration of each of the first and second musical phrases is 2 seconds, and the playback duration of the second video sequence is 3 seconds, which is the duration of the second musical phrase in the target audio segment, and a maximum of 9 frames may be played back at a rate of 3 frames per second. Therefore, it is necessary to extract 9 or fewer consecutive video frames from the 15th frame to the 26th frame. Sentence Based on the overall distribution of the numerical values of the evaluation results of the multiple video frames corresponding to each of the video frames or the video frame with the highest numerical value of the evaluation result, Sentencemay be determined as the video frames included in the target audio segment to be extracted.
[0058] For example, assume that the second embodiment described above is currently being performed on video material 3. The video frame with the highest evaluation result value determined based on the evaluation results corresponding to the video frames is the 20th frame, and the original audio corresponding to the video material 3 has the highest evaluation result value corresponding to the 20th frame. Sentence The time length of Sentence is 2 seconds, and the later Sentence The playback time length of the third musical phrase in the target audio segment is 2 seconds, and the playback time length of the video material 3 is 8 seconds. Therefore, the search is performed before and after the 20th frame, and the original audio corresponding to the 20th frame is Sentence , previous Sentence , later Sentence Determine one of them. Sentence may be determined as the video frames included in the target audio segment to be extracted.
[0059] The electronics can perform the process for each of the video materials to extract corresponding target video segments from the selected video materials.
[0060] It should be noted that the electronic device may extract the target video segment from the video material in other ways.
[0061] S204: Integrate the multiple target video segments and the target audio material to generate a target video including multiple video segments each derived based on the multiple target video segments, wherein the multiple video segments in the target video are played in chronological order, and durations of the multiple video segments match durations of corresponding musical phrases in the target audio material.
[0062] The plurality of target video segments includes a target video segment corresponding to each of the video materials.
[0063] In one possible embodiment, the electronic device may generate a target video by filling a clip template with multiple target video segments in a posting chronological order, and integrating the target video with target audio material (or target audio segments of the target audio material), where the clip template is a template for instructing clip manipulation manners corresponding to one or more clip types, such as one or more of transitions, filters, effects, etc.
[0064] During integration, the duration of a target video segment may be greater or less than the duration of the corresponding musical phrase, so the playback speed of the target video segment may be sped up or slowed down so that the speed-adjusted target video segment matches the duration of the corresponding musical phrase, and thus the video segments contained in the target video are said speed-adjusted target video segments.
[0065] In addition, in the target video, information about the target video segment may be displayed in each video frame of each target video segment, such as the posting time of the target video segment, the posting location, the title when the user posted the target video segment, whether the video is a duet type, etc. The information about the target video segment may be displayed at any position of the video frame, such as the upper left, upper right, lower left, lower right, etc. In order to ensure the user's visual experience, the screen of the target video frame may be as unobstructed as possible.
[0066] The method according to the present embodiment selects a plurality of video materials from a video related to a user, extracts some target video segments from the selected plurality of video materials, and integrates the extracted target video segments with the determined target audio material to generate a target video. The target video allows the user to look back on deep memories from the past. In addition, the duration of each video segment in the target video matches the duration of a musical phrase in the adopted target audio material, so that the content rhythm and audio rhythm of the target video match, giving the user a unique experience.
[0067] 3 is a flowchart of a video generating method according to another embodiment of the present disclosure. The embodiment shown in FIG. 3 is mainly for illustrating an embodiment of determining target audio material and target audio segments. As shown in FIG. 3, the method of this embodiment includes the following steps:
[0068] S301: Obtain a first candidate audio set by filtering out a plurality of audio materials included in a predetermined audio material set based on a predetermined audio feature set.
[0069] Step S301 corresponds to filtering the audio in dimensions such as audio style tag, audio mood tag, audio language tag, and audio scene tag.
[0070] Here, the predetermined audio material set may include a plurality of audio materials in an audio library, audio materials added to favorites by a user, audio materials used in videos posted by a user, etc. The present disclosure does not limit the implementation form of determining the predetermined audio material set. The electronic device may obtain an identification of each audio material in the predetermined audio material set, audio feature information of the audio material, lyric information of the audio material, etc. Here, the identification of the audio material may be used to uniquely identify the audio material, so that a search can be easily performed in the predetermined audio material set based on the identification of the target audio material that is determined later. The identification of the audio material may be, for example, an audio material ID, a numeric number, etc.
[0071] The audio features included in the predefined audio feature set are audio features in various dimensions of the audio material that is not recommended as a soundtrack for the target video. For example, the predefined audio feature set may include audio features in one or more dimensions, such as audio style features, audio mood features, audio language features, audio scene features, etc. Of course, the predefined audio feature set may include audio features in other dimensions, and the present disclosure is not limited thereto. Exemplarily, the predefined audio feature set includes audio mood features, such as tension, anger, fatigue, etc.
[0072] In the process of filtering out audio materials from the set of predetermined audio materials based on the set of predetermined audio features, the electronic device may obtain a first set of candidate audio materials by matching music features corresponding to each audio material with each audio feature in the set of predetermined audio features, and determining the audio material as non-candidate audio if the matching is successful, and determining the audio material as candidate audio material if the matching is not successful.
[0073] The first set of candidate audio material may include a plurality of candidate audio material.
[0074] S302: Obtain a second set of candidate audio materials by excluding each audio material included in the first set of candidate audio materials based on a predetermined audio beat.
[0075] Step S302 corresponds to performing selection of the audio material along the audio rhythm dimension.
[0076] Here, the audio beat may be expressed in beats per minute, or BPM (Beat Per Minute), which may be used to determine the rhythmic speed of audio or speech, with a higher BPM value indicating a faster rhythm.
[0077] The electronic device obtains a BPM corresponding to each audio material in the first set of candidate audio material, and compares the BPM corresponding to the audio material with a predetermined audio beat, and audio material that meets the predetermined audio beat is audio material in the second set of candidate audio material, and audio material that does not meet the predetermined audio beat is excluded audio material, which is not a potential soundtrack for the target video.
[0078] The predetermined audio beat may be a BPM range, for example BPM=[80,130].
[0079] It should be noted that the execution order of the above S301 and S302 does not have to be in order, and S302 may be executed first, and then S301.
[0080] Next, the electronic device may determine the target audio material from among the second set of candidate audio materials, where the duration of each musical phrase contained in the particular audio segment satisfies a predefined duration condition. Exemplarily, the method may include steps S303 to S305.
[0081] S303: Perform speech recognition on each of the audio materials in the second set of candidate audio materials to obtain specific audio segments corresponding to each of the audio materials in the second set of candidate audio materials.
[0082] S304: For each audio material in the second candidate audio set, perform musical phrase segmentation on a specific audio segment of the audio material to obtain each musical phrase contained in the specific audio segment, and determine whether the audio material is audio material in a third candidate audio material set according to whether a duration of each musical phrase contained in the target audio segment satisfies a predetermined duration condition.
[0083] Steps S303 and S304 correspond to performing a selection of the audio material in the musical phrase duration dimension and based on the durations of musical phrases contained within particular audio segments of the audio material.
[0084] Among them, the specific audio segment may be understood as a climax audio segment in the audio material. The present disclosure does not limit the implementation of obtaining the corresponding specific audio segment from the audio material. For example, the specific audio segment may be an audio segment obtained by segmenting the target audio by analyzing a specific attribute of the audio material with an audio processing model and combining a number of video materials to generate a target video corresponding to a certain time period. The present disclosure does not limit the type of the audio processing model. For example, the audio processing model may be a neural network model, a convolution model, etc.
[0085] After performing audio processing on each audio material in the second set of candidate audio materials and obtaining specific audio segments corresponding to each audio material in the second set of candidate audio materials, a Voice Activity Detection (VAD) technique may be adopted to analyze the lyrics of the specific audio segments, and determine the start and end positions of the lyrics of each musical phrase in the specific audio segments, i.e., determine the lyric boundaries corresponding to each musical phrase, and map the start and end positions of the lyrics of each musical phrase to the specific audio segment to obtain the duration of each musical phrase.
[0086] If the durations of the musical phrases of the target audio segment all satisfy a second predetermined condition (e.g. a preset duration requirement), the audio material to which the particular audio segment belongs is determined as audio material in the third set of candidate audio material. If the durations of one or more musical phrases of the particular audio segment do not satisfy the second predetermined condition, the audio material to which the particular audio segment belongs is determined as audio material to be excluded.
[0087] It should be noted that the second predefined condition (e.g., a preset duration requirement) may be determined based on the minimum and maximum duration of a single video segment allowed in the target video. For example, if each video segment in the target video is at least 2 seconds long and at most 5 seconds long, the second predefined condition may be no less than 3 seconds and no more than 6 seconds long. Leaving a time margin of 1 second can be used for transition and also avoids errors in merging the video segments later.
[0088] In one possible case, one or more audio materials included in the second set of candidate audio materials may satisfy the condition, so that if the third set of candidate audio materials includes one or more audio materials, the target audio material may be determined from the audio materials included in the third set of candidate audio materials.
[0089] In another possible case, it is possible that none of the audio materials in the second set of candidate audio materials satisfies the second predetermined condition, and thus the target audio material cannot be determined from the third set of candidate audio materials if there is no audio material in the third set of candidate audio materials.
[0090] S305: Determine the target audio material from audio material included in a third set of candidate audio material, the target audio segment being a particular audio segment corresponding to the target audio material.
[0091] In one possible embodiment, the electronic device may perform a weighting calculation for the audio materials in the third set of candidate audio materials based on information such as the total number of videos posted using the audio materials, the number of favorites for the audio materials, and the number of videos posted by the target user using the audio materials, and sort each audio material in the third set of candidate audio materials in ascending order based on the weighting calculation result, and select the audio material ranked first based on the sorting result as the target audio material.
[0092] In another possible embodiment, the electronic device may randomly select one audio material from the third set of candidate audio materials as the target audio material.
[0093] Of course, the electronic device may select the target audio material from the third set of candidate audio material in other ways, and this disclosure is not intended to limit the implementation of how the target audio material is selected from the third set of candidate audio material.
[0094] It should be noted that sorting the candidate audio materials based on information such as the total number of videos posted using the audio material, the number of favorites of the audio, the number of videos posted by the target user using the audio material, etc. may be performed after steps S301 and S302 and before step S303. Thus, in the process of performing S303 to S305, the candidate audio materials may be selected in order of their earliest selection.
[0095] S306: Matching is performed between the user's preferences and audio features corresponding to each audio material in a pre-specified audio material set, and if the matching is successful, a target audio material is determined based on the audio material with which the matching is successful.
[0096] When a user's preferences are successfully matched with the audio characteristics of one or more audio materials in a pre-specified audio material set, one of the one or more audio materials with successful matching may be determined as the target audio material.
[0097] For example, a weighting calculation may be performed based on information such as the total number of videos posted using the audio material, the number of favorites for the audio material, and the number of videos posted by the target user using the audio material, and based on the results of the weighting calculation, the audio materials with successful matching may be sorted in descending order, and the audio material ranked first may be selected as the target audio material, or one of the audio materials with successful matching may be randomly selected as the target audio material.
[0098] If the user's preference does not match the audio characteristics of the predetermined audio, a random audio material may be selected as the target audio material from the set of predetermined audio materials or a set of pre-specified audio materials.
[0099] By selecting a target audio material using the method of this embodiment, the selected target audio material is matched with multiple target video segments extracted from multiple video materials posted by a user within a specified time period, and the selected target audio material is used as background music for the target video, so that the target video can meet the user's expectations as much as possible and improve the user experience.
[0100] The opening and ending of a video are important parts of the video and can enhance the quality of the video, so we will explain in detail how to generate the opening and ending of a target video by the embodiment shown in Figure 4.
[0101] 4 is a flowchart of a video generating method according to another embodiment of the present disclosure. As shown in FIG. 4, the method according to this embodiment includes the following steps:
[0102] S401: A plurality of video materials are obtained from an original video material set.
[0103] Here, the original video material set includes video materials posted by users during a specified time period, and a plurality of video materials selected from the original video material set are used to generate a target video corresponding to the specified time period.
[0104] S402: Obtain a target audio material to be used as background music.
[0105] Steps S401 and S402 in this embodiment are similar to steps S201 and S202 in the embodiment shown in FIG. 2, respectively, and reference may be made to the detailed description of the embodiment shown in FIG. 2, and for the sake of brevity, no further description will be given here.
[0106] It should be noted that since the target video includes an opening and an ending, when determining the target audio material, not only the duration of the target video segment needs to be considered, but also the duration of the opening and the ending needs to correspond to the duration of one musical phrase respectively, so as to ensure that the structure of the whole target video is consistent, i.e., ensure that each video segment in the target video matches the duration of one complete musical phrase.
[0107] S403: For each video material, perform image feature extraction on each video frame of the video material respectively, and obtain target video segments and target video frames corresponding to the video material according to the image feature information respectively corresponding to each video frame of the video material, thereby obtaining a plurality of target video segments and video frame sets.
[0108] A video frame set includes target video frames extracted from each of the video materials, and a number of target video frames included in the video frame set are used to generate an opening and / or an ending of a target video.
[0109] Here, for the implementation of extracting video frame image features from multiple video materials respectively, performing segmentation processing based on the image feature information corresponding to each video frame, and obtaining a target video segment, reference can be made to the detailed description of the embodiment shown in FIG. 2, and for the sake of simplicity, no further description will be given here.
[0110] Here, the target video frames included in the video frame set are image materials for generating the opening and / or ending of the target video, and this section mainly describes how to obtain each target video frame included in the video frame set.
[0111] As one possible embodiment, the target video frame may be a video frame extracted from a target video segment corresponding to the video material, and the image feature information of the target video frame satisfies a condition.
[0112] For example, as described above, when determining the target video segment, the electronic device may obtain an evaluation result corresponding to each video frame of the video material based on the image feature information of each video frame in the video material. Assuming that the evaluation result is a numerical value, the electronic device may select a video frame having the highest numerical value of the evaluation result in the target video segment as the image material for generating the opening and / or ending. Alternatively, the electronic device may select one video frame from a plurality of video frames having the highest numerical value of the evaluation result in the target video segment as the image material for generating the opening and / or ending.
[0113] By performing the above process for each video material, video frames that satisfy the conditions can be obtained from each video material, and used as image material to generate the opening and / or ending of the target video, i.e., a set of video frames can be obtained.
[0114] S404: Generate an opening and / or an ending from the set of video frames, and generate a target video by integrating the opening and / or ending, and the multiple target video segments and the target audio material.
[0115] The target video includes an opening, a plurality of video segments each derived based on a plurality of target video segments, and an ending, the plurality of video segments in the target video are played in chronological order, and the durations of the plurality of video segments match the durations of corresponding musical phrases in the target audio material.
[0116] Assuming that the background music of the target video is the target audio segment of the target audio material, the duration of the opening matches the duration of the first musical phrase of the target audio segment, the duration of the ending matches the duration of the last musical phrase of the target audio segment, the duration of each target video segment matches the duration of the corresponding musical phrase, the duration of the first target video segment matches the duration of the second musical phrase in the target audio segment, the duration of the second target video segment matches the duration of the third musical phrase in the target audio segment, and the duration of the third target audio segment matches the duration of the fourth musical phrase in the target audio segment.
[0117] As one possible embodiment, assuming that an opening and an ending need to be generated, the electronic device may edit (clip) multiple target video frames included in a video frame set with different clip templates to obtain the opening and the ending, respectively. The clip template is used to indicate a clipping scheme for clipping the target video frames included in the video frame set to a video segment. For example, the clipping scheme may include, but is not limited to, one or more of a transition scheme, an effect, a filter, and the like. The present disclosure is not limited to the clip template.
[0118] Since the duration of the opening needs to match the duration of the first musical phrase of the target audio segment, and the duration of the ending needs to match the duration of the last musical phrase of the target audio segment, a clip template for clipping the opening may be selected based on the duration of the first musical phrase of the target audio segment, and a clip template for clipping the ending may be selected based on the duration of the last musical phrase of the target audio segment. Of course, when selecting a clip template, other factors may also be taken into consideration, such as the style of the video segment to be generated based on the clip template, requirements for image material of the clip template, etc.
[0119] It should be noted that during the merging process, the duration of a target video segment may be greater or less than the duration of the corresponding musical phrase, and thus the playback speed of the target video segment may be sped up or slowed down so that the speed-adjusted target video segment matches the duration of the corresponding musical phrase.
[0120] In this embodiment, video frames are extracted from multiple target video segments respectively to generate the opening and ending of the target video, which is advantageous in making the target video more exciting, ensuring that the target video meets the user's expectations as much as possible, improving the user's experience, and increasing the user's interest in posting the video.
[0121] In one specific embodiment, it is assumed that the video generating method is executed by a server-side device. The server-side device specifically includes a business server side, an audio algorithm server side, an image algorithm server side, and a cloud service for executing an integration process. And the server-side device may generate a target video corresponding to a certain time period (such as the past year) offline in advance, and when the business server side receives a video acquisition request sent by a client, send the target video corresponding to the certain time period (such as the past year) to the client.
[0122] The following describes in detail how the server side pre-generates the target video through the embodiment shown in Figure 5. The embodiment shown in Figure 5 shows a specific flow of the video generation method in the above scene.
[0123] As shown in FIG. 5, this embodiment includes the following steps.
[0124] S501: The business server side transmits a video material acquisition request to the database.
[0125] Here, the video material acquisition request is an acquisition request for requesting a video posted by a target user during a specific time period. The video material acquisition request may include an indicator of the target user and instruction information for instructing to request the video material posted by the target user during a specific time period.
[0126] S502: The database transmits a set of original video materials corresponding to the target user to the business server side.
[0127] S503: The business server side determines a plurality of video materials from the original video material set corresponding to the target user.
[0128] S504: The business server side excludes the predetermined audio material set based on the predetermined audio feature set and the predetermined audio beat to obtain a second candidate audio material set.
[0129] S505: The business server side transmits an indication of each audio material included in the second set of candidate audio materials to the audio algorithm server side.
[0130] S506: The audio algorithm server side downloads the audio material according to the identification of the received audio material, identifies a specific audio segment for the audio material, and performs musical phrase segmentation and selection to determine the target audio material.
[0131] S507: The audio algorithm server side transmits the target audio material's identifier and audio time stamp information to the business server side.
[0132] Here, the audio time stamp information includes time stamp information of each musical phrase in a target audio segment in the target audio material, and the audio time stamp information can be used to determine the start time and end time of each musical phrase in each target audio segment, and further, the duration information of each musical phrase can be obtained.
[0133] S508: The business server side sends the identifiers of the multiple video materials, the identifiers of the target audio materials, and the audio time stamp information to the image algorithm server side.
[0134] S509: The image algorithm server side obtains a plurality of video materials and a target audio material from the indicators of the plurality of video materials and the indicator of the target audio material.
[0135] S510: The image algorithm server side obtains time stamp information of each of a plurality of target video segments and a video integration logic between the plurality of target video segments and the target audio segments according to a plurality of video materials, a target audio material and audio time stamp information.
[0136] S511: The image algorithm server side sends time stamp information of each of a plurality of target video segments and a video integration logic between the plurality of target video segments and the target audio segment to the business server side.
[0137] S512: The business server side sends to the cloud service indicators of the multiple video materials, timestamp information of each of the multiple target video segments, indicators of the target audio materials, timestamp information of the target audio segments, and video integration logic between the multiple target video segments and the target audio segments.
[0138] S513: The cloud service performs video integration based on the identification of the multiple video materials, the timestamp information of each of the multiple target video segments, the identification of the target audio material, the timestamp information of the target audio segments, and video integration logic between the multiple target video segments and the target audio segments to generate a target video.
[0139] The generated target video may be stored in a cloud service.
[0140] S514: The client receives a video acquisition request sent by the user.
[0141] S515: The client transmits a video acquisition request to the business server side.
[0142] S516: The business server side transmits a video acquisition request to the cloud service.
[0143] S517: The cloud service transmits the target video to the business server side.
[0144] S518: The business server side transmits the target video to the client.
[0145] S519: The client loads the target video into the video editing page and plays it.
[0146] The client may then clip and post the target video based on the user's actions.
[0147] In the embodiment shown in FIG. 5, the implementation of the audio algorithm server side and the image algorithm server side can be referred to the detailed description of the above method embodiment, and will not be further described in this embodiment for the sake of conciseness.
[0148] The method according to the present embodiment selects a plurality of video materials from the video materials posted by the user in a certain time period, extracts some target video segments from the selected plurality of video materials, and integrates the extracted target video segments with the determined target audio segment to generate a target video corresponding to the certain time period, so that the user can look back on deep memories in the certain time period through the target video. In addition, the duration of each video segment in the target video matches the duration of the musical phrase in the adopted target audio segment, so that the content rhythm and audio rhythm of the target video match, and a unique experience feeling can be given to the user. And this aspect is completed in advance offline by the client. When the server-side device receives the video acquisition request sent by the client, it can quickly respond to the video acquisition request, ensuring the user experience.
[0149] Illustratively, the present disclosure further provides a video generation device.
[0150] 6 is a schematic diagram of a video generating device according to an embodiment of the present disclosure. Here, the video generating device according to this embodiment may be a video generating system. As shown in FIG. 6, the video generating device 600 may include a video processing module 601, an audio processing module 602, and a video integration module 603.
[0151] In the video generation device 600, a video processing module 601 is used to obtain a number of video footage from a set of original video footage that includes videos relating to a user.
[0152] The audio processing module 602 is used to obtain target audio material to be used as background music.
[0153] The video processing module 601 is further used to perform image feature extraction for each video frame of the video material, and perform a segmentation process based on the image feature information corresponding to each of the video frames of each of the video materials to obtain a target video segment corresponding to the video material.
[0154] The integration module 603 is used to integrate the target video segments and the target audio material respectively corresponding to each of the video materials to generate a target video including a plurality of video segments each obtained based on a plurality of target video segments, wherein the plurality of video segments in the target video are played in chronological order of posting, and the durations of the plurality of video segments match the durations of corresponding musical phrases in the target audio material.
[0155] The video generating device according to this embodiment may be used to implement the technical solutions shown in any one of the method embodiments described above, and the implementation principles and technical effects are similar, and for clarity, reference may be made to the detailed description of the above method embodiment.
[0156] In one possible embodiment, the audio processing module 602 is specifically used to determine target audio material from a predefined set of audio materials based on the audio features of each audio material in the predefined set of audio materials, the beat information of the audio materials, and the duration of each musical phrase of a particular audio segment in the audio material, where the particular audio segment in the target audio material is background music of the target video.
[0157] In one possible embodiment, the audio processing module 602 is specifically used for: excluding musical materials included in a predetermined set of musical materials based on a predefined set of musical features to obtain a first set of candidate audio materials; excluding audio materials included in the first set of candidate audio materials based on a predefined audio beat to obtain a second set of candidate audio materials; and determining the target audio material based on audio materials in the second set of candidate audio materials, the durations of musical phrases included in an audio segment of which satisfy a predefined duration condition.
[0158] In one possible embodiment, if there is no audio material in the second candidate audio material set in which the duration of a musical phrase in a specific audio segment satisfies a predetermined duration condition, the audio processing module 602 is further used to match the audio features corresponding to each audio material in the pre-specified audio material set based on the user's preferences, and if the matching is successful, determine the target audio material based on the audio material with which the matching is successful.
[0159] In one possible embodiment, the video processing module 601 further performs a weighting calculation for each of the video materials based on image feature information respectively corresponding to each of the video frames of the video materials to obtain an evaluation result respectively corresponding to each of the video frames; and extracting a target video frame from each of the video frames of the video material to include a plurality of target video frames extracted from the plurality of video materials based on the evaluation results respectively corresponding to each of the video frames, thereby obtaining a set of video frames for generating an opening and / or an ending of the target video.
[0160] Accordingly, the integration module 603 is specifically used for integrating the target video segments, the video frame sets, and the target audio segments corresponding to each of the video materials, respectively, to generate the target video including an opening and / or an ending generated from the video frame sets.
[0161] In one possible embodiment, the video processing module 601 is specifically used to perform a segmentation process on the video material to obtain the target video segment based on image feature information corresponding to each video frame of the video material, the duration of the corresponding musical phrase in the target audio segment, and the sentence segmentation result of the original audio in the video material.
[0162] In one possible embodiment, the target audio segment is one or more complete subsequences of the corresponding original audio. Sentence Includes.
[0163] As one possible embodiment, the video generation device 600 may further include a storage module (not shown in FIG. 6) for storing the generated target video.
[0164] As one possible embodiment, when the video generation device 600 is a server-side device, the video generation device may further include a communication module (not shown in FIG. 6) for receiving a video acquisition request sent from a client and, in response to the video acquisition request, transmitting a target video corresponding to a corresponding predetermined time period to the client.
[0165] What should be explained is that details that are not described in detail in the apparatus embodiment shown in FIG. 6 can be referred to the description of the method embodiment described above, and for the sake of simplicity, they are not described one by one in the apparatus embodiment.
[0166] Illustratively, the present disclosure further provides an electronic device.
[0167] 7 is a schematic diagram of an electronic device according to an embodiment of the present disclosure. As shown in FIG. 7, an electronic device 700 according to this embodiment includes a memory 701 and a processor 702.
[0168] Here, the memory 701 may be an independent physical unit, and may be connected to the processor 702 via a bus 703. The memory 701 and the processor 702 may be integrated together and realized by hardware or the like.
[0169] The memory 701 is used for storing program instructions, and the processor 702 calls the program instructions to execute the technical solutions of any one of the above method embodiments.
[0170] Alternatively, when a part or all of the methods in the above embodiments are realized by software, the above electronic device 700 may include only the processor 702. The memory 701 for storing the programs is outside the electronic device 700, and the processor 702 is connected to the memory through a circuit / wiring and is used to read and execute the programs stored in the memory.
[0171] The processor 702 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and a NP.
[0172] The processor 702 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0173] The memory 701 may include volatile memory, such as random-access memory (RAM). The memory may also include non-volatile memory, such as flash memory, a hard disk drive (HDD) or a solid-state drive (SSD). The memory may also include a combination of the above types of memory.
[0174] The present disclosure further provides a readable storage medium comprising computer program instructions which, when executed by at least one processor of an electronic device, cause the video generation method as set forth in any one of the method embodiments above to be implemented.
[0175] The present disclosure further provides a computer program product which, when executed by a computer, causes the computer to implement the video generation method as set forth in any one of the method embodiments above.
[0176] It should be explained that, in this specification, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not require or imply that such an actual relationship or sequence exists between the entities or operations. Moreover, the terms "comprise", "include", or any other variation thereof, indicate a non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements not expressly stated or elements inherent in such process, method, article, or device. Absent further limitations, an element defined by the musical phrase "comprises a..." does not exclude the inclusion of other identical elements in the process, method, article, or device that includes said elements.
[0177] The above are merely specific embodiments of the present disclosure, and are used to enable those skilled in the art to understand or realize the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Thus, the present disclosure is not limited to these embodiments herein, but has the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. 1. A method of video generation, the method comprising: obtaining a plurality of video footage from a set of original video footage that includes video relating to a user; obtaining a target audio material to be used as background music; For each of the video materials, performing image feature extraction for each video frame of the video material respectively, and performing a segmentation process based on the image feature information respectively corresponding to each of the video frames of each of the video materials to obtain a target video segment corresponding to the video material; aggregating the target video segments and the target audio material corresponding to each of the video materials to generate a target video including a plurality of video segments each derived based on a plurality of target video segments, wherein the plurality of video segments in the target video are played in chronological order, and durations of the plurality of video segments match durations of corresponding musical phrases in the target audio material; The video generation method includes: performing a weighting calculation for each of said video footage based on image feature information respectively corresponding to each of said video frames of said video footage to obtain an evaluation result respectively corresponding to each of said video frames; and extracting a target video frame from each of the video frames of the video footage based on the evaluation results corresponding to each of the video frames, respectively, to obtain a video frame set, the video frame set including a plurality of target video frames extracted from the plurality of video footage, the video frame set being used to generate an opening and / or an ending of the target video. method.
2. The step of obtaining a target audio material to be used as background music includes:
2. The method of claim 1, further comprising determining target audio material from a set of predetermined audio material based on audio features of each audio material in the set, beat information of the audio material, and a duration of each musical phrase of a particular audio segment in the audio material, wherein the particular audio segment in the target audio material is background music of the target video.
3. determining target audio material from the set of predetermined audio material based on audio features of each audio material in the set of predetermined audio material, beat information of the audio material, and durations of each musical phrase of a particular audio segment in the audio material, comprising: - excluding a plurality of music materials included in the predetermined set of music materials based on the predetermined set of music features to obtain a first set of candidate audio material; - excluding each audio material included in the first set of candidate audio material based on a predefined audio beat to obtain a second set of candidate audio material; and determining the target audio material based on audio material in the second set of candidate audio material, the duration of each musical phrase contained in a particular audio segment satisfying a predetermined duration condition.
4. If there is no audio material in the second set of candidate audio materials, the duration of a musical phrase contained in a particular audio segment satisfies a predetermined duration condition, the method further comprises:
4. The method of claim 3, further comprising: matching audio features corresponding to each audio material of a pre-specified set of audio materials based on the user's preferences; and, if matching is successful, determining the target audio material based on the successfully matched audio material.
5. a step of generating a target video by integrating target video segments and the target audio material respectively corresponding to each of the video materials, the step comprising: The method of claim 1, further comprising a step of integrating a target video segment, the set of video frames and the target audio material, each of which corresponds to one of the video materials, to generate the target video including an opening and / or an ending generated from the set of video frames.
6. performing a segmentation process based on image feature information corresponding to each video frame of the video material to obtain a target video segment; The method of claim 1, further comprising: performing a segmentation process on the video material to obtain the target video segment based on image feature information corresponding to each video frame of the video material, the duration of the corresponding musical phrase in the target audio segment, and the sentence segmentation result of the original audio in the video material.
7. The method of claim 6 , wherein the target audio segment comprises one or more complete sentences of the corresponding original audio.
8. 1. A video generation device, the video generation device comprising: a video processing module for obtaining a plurality of video footage from a set of original video footage including video relating to a user; an audio processing module for obtaining a target audio material to be used as background music, The video processing module further includes an audio processing module, which is used for performing image feature extraction for each video frame of each of the video materials, respectively, and performing a segmentation process based on the image feature information corresponding to each of the video frames of each of the video materials, respectively, to obtain a target video segment corresponding to the video materials; a video integration module for integrating the target video segments and the target audio material, each of which corresponds to one of the video materials, to generate a target video including a plurality of video segments each derived based on a plurality of the target video segments, wherein the plurality of video segments in the target video are played in chronological order, and durations of the plurality of video segments match durations of corresponding musical phrases in the target audio material; the video processing module performs a weighting calculation for each of the video materials based on image feature information respectively corresponding to each of the video frames of the video materials, obtains an evaluation result respectively corresponding to each of the video frames, and extracts a target video frame from each of the video frames of the video materials based on the evaluation result respectively corresponding to each of the video frames to obtain a video frame set, the video frame set including a plurality of target video frames extracted from the plurality of video materials, the video frame set being used to generate an opening and / or an ending of the target video; Video generation equipment.
9. An electronic device including a memory and a processor, the memory is configured to store computer program instructions; An electronic device, wherein the processor is configured to execute the computer program instructions to cause the electronic device to perform the method of any one of claims 1 to 7.
10. A readable storage medium comprising computer program instructions which, when executed by at least one processor of an electronic device, cause said electronic device to implement the method according to any one of claims 1 to 7.
11. A computer program which, when executed by a computer, causes said computer to implement the method according to any one of claims 1 to 7.
12. 1. A video production device configured to store a computer program, comprising: A video production device, wherein the computer program, when executed by a processor of the video production device, causes the video production device to implement the method of any one of claims 1 to 7.
Citation Information
Patent Citations
Video processing method and device
CN111866585A
Video processing method and device, electronic equipment and storage medium
CN112235631A
Information processing apparatus, information processing method, and information processing program
JP2016039547A