Audio generation method, device and equipment and computer readable storage medium
By obtaining video features, duration features and audio-visual synchronization features, and combining audio description text, audio is automatically generated, which solves the problems of low matching and low efficiency of audio generation in the existing technology, and achieves efficient and accurate audio generation.
Patent Information
- Application Number
- CN202510853769.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
In the prior art, the user manually operates the audio and video to generate low matching degree, poor generation effect and low efficiency, and the audio generation method is complex.
By obtaining video features, video duration features and audio-visual synchronization features, audio and video are automatically generated to ensure the duration and synchronization of audio and video. Combining audio description text features, multimodal feature fusion and processing model are used to generate audio.
It improves the matching degree and generation efficiency between audio and video, simplifies the generation process, and the generated audio and video are highly consistent and meet user requirements.
Smart Images

Figure CN120358378A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to an audio generation method, apparatus, device, and computer-readable storage medium. Background Art
[0002] With the continuous development of artificial intelligence technology, the generation methods of videos have become more and more diverse. For example, videos can be automatically generated through artificial intelligence models. The videos generated through artificial intelligence models are silent videos, and it is necessary to dub the silent videos to obtain the corresponding audio of the silent videos.
[0003] In the related art, users dub silent videos through video editing software to obtain the corresponding audio of the silent videos.
[0004] However, this method requires manual operation by users and depends on the subjective consciousness of users, resulting in a low matching degree between the generated audio and the video, poor audio generation effect, and a relatively complex audio generation method, thereby leading to low audio generation efficiency. Summary of the Invention
[0005] The present disclosure provides an audio generation method, apparatus, device, and computer-readable storage medium. This method improves the matching degree between the generated audio and the video, has a better audio generation effect, and simplifies the audio generation method, resulting in high audio generation efficiency. The technical solutions of the present disclosure are as follows.
[0006] According to one aspect of the embodiments of the present disclosure, an audio generation method is provided. The method includes: obtaining a video for which audio is to be generated; determining video features, video duration features, and audio-visual synchronization features corresponding to the video, where the video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance between the video's picture and sound; generating the audio corresponding to the video according to the video features, the video duration features, and the audio-visual synchronization features, where the duration of the audio is the same as the duration of the video.
[0007] According to another aspect of the embodiments of the present disclosure, an audio generation apparatus is provided. The apparatus includes: an obtaining unit configured to obtain a video for which audio is to be generated.
[0008] A determining unit configured to determine video features, video duration features, and audio-visual synchronization features corresponding to the video, where the video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance between the video's picture and sound.
[0009] A generating unit, configured to generate the audio corresponding to the video according to the video feature, the video duration feature, and the audio-visual synchronization feature, where the duration of the audio is the same as the duration of the video.
[0010] In some embodiments, the obtaining unit is further configured to obtain an audio description text, where the audio description text is a description text of the audio corresponding to the video.
[0011] The determining unit is further configured to determine the text feature corresponding to the audio description text.
[0012] The generating unit is configured to generate the audio corresponding to the video according to the video feature, the video duration feature, the audio-visual synchronization feature, and the text feature.
[0013] In some embodiments, the determining unit is configured to obtain a first feature corresponding to the audio description text, where the first feature is used to characterize the audio description text; extract keywords from the audio description text to obtain at least one keyword; obtain a second feature corresponding to each keyword, where the second feature corresponding to each keyword is used to characterize the corresponding keyword; and determine the text feature corresponding to the audio description text according to the first feature and the second features corresponding to the respective keywords.
[0014] In some embodiments, the obtaining unit is configured to process the audio description text through a first text feature obtaining model to obtain the first feature corresponding to the audio description text, where the first text feature obtaining model is used to obtain text features.
[0015] In some embodiments, for any one of the keywords, the obtaining unit is configured to process the any one keyword through a second text feature obtaining model to obtain the second feature corresponding to the any one keyword, where the second text feature obtaining model is used to obtain text features, and the number of tokens supported by the second text feature obtaining model is less than the number of tokens supported by the first text feature obtaining model.
[0016] In some embodiments, the generating unit is configured to perform fusing the video feature, the text feature, and the audio noise feature to obtain a fused feature, where the audio noise feature and the audio feature of the audio corresponding to the video have the same dimension; splitting the fused feature to obtain a split video feature, a split text feature, and a split audio noise feature, where the split video feature, the split text feature, and the split audio noise feature are all fused with the video feature, the text feature, and the audio noise feature; generating an alignment feature according to the audio-visual synchronization feature, where the dimension of the alignment feature is the same as that of the audio noise feature; generating a global feature according to the text feature, the video feature, the video duration feature, and the alignment feature, where the global feature is a feature after considering the text feature, the video feature, the video duration feature, and the alignment feature; and generating the audio corresponding to the video according to the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature.
[0017] In some embodiments, the generating unit is configured to perform upsampling the audio-visual synchronization feature to the same dimension as the audio noise feature to obtain the alignment feature.
[0018] In some embodiments, the generating unit is configured to perform pooling on the text feature, the video feature, and the video duration feature to obtain a pooled feature, where the pooled feature is used to characterize the consistency between the audio description text and the video; and generating the global feature according to the pooled feature and the alignment feature.
[0019] In some embodiments, the generating unit is configured to perform adding the pooled feature and the alignment feature to obtain the global feature.
[0020] In some embodiments, the generating unit is configured to perform determining the audio feature of the audio corresponding to the video according to the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature; and generating the audio corresponding to the video according to the audio feature.
[0021] In some embodiments, the generating unit is configured to perform processing the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature through an audio feature generation model to obtain the audio feature of the audio corresponding to the video, where the audio feature generation model is used to generate audio features.
[0022] In some embodiments, the determining unit is configured to perform frame extraction on the video to obtain multiple pictures included in the video; determine the picture features of each picture, where the picture features of each picture are used to characterize the corresponding picture; and determine the video features corresponding to the video according to the picture features of each picture.
[0023] In some embodiments, the determining unit is configured to perform, for any one of the pictures, processing the any one picture through a picture feature acquisition model to obtain the picture features of the any one picture, where the picture feature acquisition model is used to acquire picture features.
[0024] In some embodiments, the determining unit is configured to obtain the number of frames and the frame rate of the video; determine the duration of the video according to the number of frames and the frame rate of the video; and determine the video duration features corresponding to the video according to the duration of the video.
[0025] In some embodiments, the determining unit is configured to determine the quotient between the number of frames and the frame rate of the video; in the case where the quotient is an integer, determine the quotient as the duration of the video; and in the case where the quotient is a non-integer, round up the quotient to obtain the duration of the video.
[0026] In some embodiments, the determining unit is configured to map the duration of the video to an array; process the array through a duration learning model to obtain the video duration features corresponding to the video, where the duration learning model is used to determine video duration features.
[0027] In some embodiments, the duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer, and a third convolutional layer; the determining unit is configured to input the array into the fully connected layer to obtain the output result of the fully connected layer; input the output result of the fully connected layer into the first convolutional layer and the second convolutional layer respectively to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; map the output result of the first convolutional layer through an activation function to obtain an attention weight; multiply the attention weight and the output result of the second convolutional layer to obtain a multiplication result; and input the multiplication result into the third convolutional layer to obtain the video duration features of the video.
[0028] In some embodiments, the determining unit is configured to process the video through an audio-visual synchronization encoder to obtain the audio-visual synchronization features corresponding to the video, where the audio-visual synchronization encoder is used to acquire audio-visual synchronization features.
[0029] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, which includes: one or more processors; a memory for storing program code executable by the processor; wherein, the processor is configured to execute the program code to implement the above audio generation method.
[0030] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the program code in the computer-readable storage medium is executed by a processor of a terminal, enabling the electronic device to execute the above audio generation method.
[0031] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, which implements the above audio generation method when executed by a processor.
[0032] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure.
[0033] The embodiments of the present disclosure provide an audio generation method, which generates audio corresponding to a video according to the video, without manual operation by the user. Therefore, the audio generation method is simplified, and the audio generation efficiency is relatively high. Moreover, it does not rely on the subjective awareness of the user, so that the matching degree between the generated audio and the video is relatively high, and the audio generation efficiency is relatively good. In addition, since not only video features are obtained, but also video duration features and audio-visual synchronization features are obtained, when generating audio, not only the video features of the video itself are considered, but also the video duration features and audio-visual synchronization features are considered, so that the duration of the generated audio is the same as that of the video, and the generated audio is synchronized with the video in terms of sound and picture, further improving the matching degree between the generated audio and the video and the audio generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] The drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure, and do not constitute an improper limitation to the present disclosure.
[0035] Figure 1 is a schematic diagram of an implementation environment shown according to an exemplary embodiment.
[0036] Figure 2 is a flowchart of an audio generation method shown according to an exemplary embodiment.
[0037] Figure 3 is a flowchart of another audio generation method shown according to an exemplary embodiment.
[0038] Figure 4 is a schematic diagram of the determination process of the video duration feature of a video shown according to an exemplary embodiment.
[0039] Figure 5 is a flowchart of yet another audio generation method shown according to an exemplary embodiment.
[0040] Figure 6 is a flowchart of still another audio generation method shown according to an exemplary embodiment.
[0041] Figure 7 is a block diagram of an audio generation device shown according to an exemplary embodiment.
[0042] Figure 8 is a block diagram of a terminal shown according to an exemplary embodiment.
[0043] Figure 9 is a block diagram of a server shown according to an exemplary embodiment. Detailed implementation manners
[0044] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0045] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.
[0046] It should be noted that the information (including but not limited to user equipment information, user personal information, etc.), data (including but not limited to data for analysis, stored data, displayed data, etc.) and signals involved in the present disclosure are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the videos and texts involved in the present disclosure are all obtained under full authorization.
[0047] The audio generation method provided by the embodiments of the present disclosure can be executed by an electronic device. Figure 1 is a schematic diagram of an implementation environment provided by the embodiments of the present disclosure. See Figure 1, the implementation environment includes: an electronic device 101. The electronic device 101 can be a terminal or a server, and the embodiments of the present disclosure do not limit this. A target application is installed in the electronic device 101, and the target application is used to generate audio. In the embodiments of the present disclosure, the electronic device 101 obtains a video for which audio is to be generated, and determines the video features, video duration features, and audio-visual synchronization features corresponding to the video. The video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance between the video's picture and sound; according to the video features, video duration features, and audio-visual synchronization features, audio corresponding to the video is generated, and the duration of the audio is the same as the duration of the video. The target application can be any application with an audio generation function. For example, the target application can be a short video application, a social application, etc., and no specific limitation is made here.
[0048] Optionally, the electronic device 101 is a terminal, and the terminal can be at least one of devices such as a smart phone, a smart watch, a desktop computer, a laptop computer, a virtual reality terminal, an augmented reality terminal, a wireless terminal, and a laptop portable computer. The terminal has a communication function and can access a wired network or a wireless network. The terminal can generally refer to one of multiple terminals, and those skilled in the art can know that the number of the above terminals can be more or less. The electronic device 101 is a server, and the server can be an independent physical server, or a server cluster or a distributed file system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server and the terminal are directly or indirectly connected through a wired or wireless communication method, and the embodiments of the present disclosure do not limit this. Optionally, the number of the above servers can be more or less, and the embodiments of the present disclosure do not limit this. Of course, the server can also include other functional servers to provide more comprehensive and diverse services. Among them, the server undertakes the main computing work, and the terminal undertakes the secondary computing work; or, the server undertakes the secondary computing work, and the terminal undertakes the main computing work; or, the server or the terminal can each independently undertake the computing work, and the embodiments of the present disclosure do not limit this.
[0049] Figure 2 is a flowchart of an audio generation method shown according to an exemplary embodiment. As Figure 2 shown, this method is executed by an electronic device, and this method includes the following steps.
[0050] In step S201, the electronic device obtains a video for which audio is to be generated.
[0051] In the embodiments of the present disclosure, the video for which the audio is to be generated can be any video, and the embodiments of the present disclosure do not limit this. A target application is installed in the electronic device, and the target application is used to generate audio. The target application can be any application capable of generating audio, and the embodiments of the present disclosure do not limit the type of the target application.
[0052] In some embodiments, the electronic device displays relevant information of the target application, and the relevant information includes at least one of the name of the target application and the icon of the target application. In response to a trigger operation on the relevant information of the target application, a video upload interface is displayed. A video upload control is displayed in the video upload interface. In response to a trigger operation on the video upload control, a video selection interface is displayed. At least one selectable video is displayed in the video selection interface. In response to a trigger operation on any one of the at least one selectable videos, the triggered selectable video is used as the video for which the audio is to be generated.
[0053] Among them, the at least one selectable video displayed in the video selection interface is a video stored in the electronic device. The trigger operation on the relevant information of the target application can be a click operation on the relevant information of the target application, a double-click operation on the relevant information of the target application, or other operations on the relevant information of the target application. The embodiments of the present disclosure do not limit this. The process of the trigger operation on other content is similar to the process of the trigger operation on the relevant information of the target application, and the embodiments of the present disclosure do not limit this either.
[0054] In step S202, the electronic device determines the video feature, video duration feature, and audio-visual synchronization feature corresponding to the video. The video feature is used to characterize the video, the video duration feature is used to characterize the duration of the video, and the audio-visual synchronization feature is used to characterize the time consistency and semantic relevance of the video image and sound.
[0055] In the embodiments of the present disclosure, the process of determining the video feature corresponding to the video includes: performing frame division processing on the video to obtain multiple pictures included in the video; determining the picture feature of each picture, and the picture feature of each picture is used to characterize the corresponding picture; and determining the video feature corresponding to the video according to the picture features of each picture.
[0056] In this embodiment, by decomposing the video into multiple pictures, extracting the picture features and integrating them to obtain the video feature, the processing complexity can be reduced while fusing the information in the spatial and temporal dimensions, improving the comprehensiveness, robustness, and adaptability of the video feature expression to different videos. Furthermore, the matching degree between the generated video feature and the video is higher, and the accuracy of the video feature is higher.
[0057] In the embodiments of the present disclosure, the process of determining the image features of each image includes: for any one of the images, processing the image through an image feature acquisition model to obtain the image features of the image, where the image feature acquisition model is used to acquire image features.
[0058] In this embodiment, by using an image feature acquisition model to obtain image features, the image information can be efficiently and accurately converted into a structured representation that can be understood by an electronic device, making the generated image features more accurate. Since the image features are used to generate video features, the accuracy of the generated video features can also be improved.
[0059] In some embodiments, the process of determining the video duration feature corresponding to a video includes: obtaining the number of frames and the frame rate of the video; determining the duration of the video according to the number of frames and the frame rate of the video; and determining the video duration feature corresponding to the video according to the duration of the video.
[0060] In this embodiment, by obtaining the number of frames and the frame rate of the video to determine the duration of the video and extract the duration feature, the time dimension information of the video can be converted into structured data in a quantitative manner, making the generated video duration feature more accurate.
[0061] In some embodiments, the process of determining the duration of a video according to the number of frames and the frame rate of the video includes: determining the quotient between the number of frames and the frame rate of the video; in the case where the quotient is an integer, determining the quotient as the duration of the video; and in the case where the quotient is a non-integer, rounding up the quotient to obtain the duration of the video.
[0062] In this embodiment, by calculating the quotient between the number of frames and the frame rate and determining the duration of the video according to the quotient, the duration of the video can be determined in a standardized calculation method, ensuring the standardization and adaptability of the duration calculation.
[0063] In some embodiments, the process of determining the video duration feature corresponding to a video according to the duration of the video includes: mapping the duration of the video to an array; and processing the array through a duration learning model to obtain the video duration feature corresponding to the video, where the duration learning model is used to determine the video duration feature.
[0064] In this embodiment, mapping the duration of the video to an array and processing it through a duration learning model can convert the duration information into a structured feature of the model, so as to adaptively extract the deep semantic representation containing temporal rules to adapt to subsequent analysis tasks.
[0065] In some embodiments, the duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer, and a third convolutional layer. The process of processing an array through the duration learning model to obtain the video duration feature corresponding to the video includes: inputting the array into the fully connected layer to obtain the output result of the fully connected layer; respectively inputting the output result of the fully connected layer into the first convolutional layer and the second convolutional layer to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; mapping the output result of the first convolutional layer through an activation function to obtain the attention weight; multiplying the attention weight and the output result of the second convolutional layer to obtain the multiplication result; inputting the multiplication result into the third convolutional layer to obtain the video duration feature of the video.
[0066] In this embodiment, through the hierarchical processing of the fully connected layer and multiple convolutional layers, and by weighting the output results of the convolutional layers with the attention weights generated by the activation function, the key features related to the video duration in the array can be efficiently captured, and the information unrelated to the video duration in the array can be suppressed, so as to accurately extract the video duration feature.
[0067] In some embodiments, the process of determining the audio-visual synchronization feature corresponding to the video includes: processing the video through an audio-visual synchronization encoder to obtain the audio-visual synchronization feature corresponding to the video, and the audio-visual synchronization encoder is used to obtain the audio-visual synchronization feature.
[0068] In this embodiment, by jointly encoding the sound and the picture of the video through the audio-visual synchronization encoder, the temporal consistency and semantic relevance between the two can be effectively captured, so as to accurately extract the audio-visual synchronization feature reflecting the audio-visual coordination relationship.
[0069] In step S203, the electronic device generates the audio corresponding to the video according to the video feature, the video duration feature, and the audio-visual synchronization feature, and the duration of the audio is the same as the duration of the video.
[0070] The embodiments of the present disclosure provide an audio generation method. This method generates the audio corresponding to the video according to the video, without manual operation by the user. Therefore, the audio generation method is simplified, and the audio generation efficiency is relatively high. Moreover, it does not rely on the subjective consciousness of the user, so that the matching degree between the generated audio and the video is relatively high, and the audio generation efficiency is relatively good. In addition, since not only the video feature is obtained, but also the video duration feature and the audio-visual synchronization feature are obtained, when generating the audio, not only the video feature of the video itself is considered, but also the video duration feature and the audio-visual synchronization feature are considered, so that the duration of the generated audio is the same as the duration of the video, and the generated audio is synchronized with the video in terms of sound and picture, further improving the matching degree between the generated audio and the video and the audio generation efficiency.
[0071] In some embodiments, the method further includes: obtaining an audio description text, where the audio description text is a description text of the audio corresponding to the video; determining the text features corresponding to the audio description text; and generating the audio corresponding to the video according to the video features, the video duration features, and the audio-visual synchronization features, including: generating the audio corresponding to the video according to the video features, the video duration features, the audio-visual synchronization features, and the text features.
[0072] In this embodiment, it is also possible to obtain the audio description text, and when generating the audio, the text features corresponding to the audio description text are also considered, so that the generated audio is not only an audio that matches the video, but also an audio that meets the user's requirements.
[0073] In some embodiments, the process of determining the text features corresponding to the audio description text includes: obtaining a first feature corresponding to the audio description text, where the first feature is used to characterize the audio description text; extracting at least one keyword from the audio description text; obtaining a second feature corresponding to each keyword, where the second feature corresponding to each keyword is used to characterize the corresponding keyword; and determining the text features corresponding to the audio description text according to the first feature and the second features corresponding to each keyword.
[0074] In this embodiment, by combining the overall semantic features of the audio description text (i.e., the first feature) with the key semantic features extracted from the audio description text (i.e., the second feature), it is possible to more comprehensively and prominently capture the audio-related information in the audio description text, thereby generating more accurate and effective text features.
[0075] In some embodiments, the process of obtaining the first feature corresponding to the audio description text includes: processing the audio description text through a first text feature acquisition model to obtain the first feature corresponding to the audio description text, where the first text feature acquisition model is used to obtain text features.
[0076] In this embodiment, by processing the audio description text through the first text feature acquisition model, the first feature that can accurately represent the overall semantic information of the audio description text can be effectively extracted, providing a basis for subsequent acquisition of text features.
[0077] In some embodiments, the process of obtaining the second feature corresponding to each keyword includes: for any one of the keywords, processing the any one keyword through a second text feature acquisition model to obtain the second feature corresponding to the any one keyword, where the second text feature acquisition model is used to obtain text features, and the number of tokens supported by the second text feature acquisition model is less than the number of tokens supported by the first text feature acquisition model.
[0078] In this embodiment, the keyword is processed by the second text feature acquisition model, and the second feature that can accurately represent the keyword can be effectively extracted, providing a basis for subsequent text feature acquisition.
[0079] In some embodiments, the process of generating the audio corresponding to the video according to the video feature, video duration feature, audio-visual synchronization feature, and text feature includes: fusing the video feature, text feature, and audio noise feature to obtain a fused feature, where the audio noise feature and the audio feature of the audio corresponding to the video have the same dimension; splitting the fused feature to obtain the split video feature, split text feature, and split audio noise feature, and the split video feature, split text feature, and split audio noise feature are all fused with the video feature, text feature, and audio noise feature; generating an alignment feature according to the audio-visual synchronization feature, where the dimension of the alignment feature is the same as that of the audio noise feature; generating a global feature according to the text feature, video feature, video duration feature, and alignment feature, where the global feature is the feature after considering the text feature, video feature, video duration feature, and alignment feature; generating the audio corresponding to the video according to the alignment feature, global feature, split video feature, split text feature, and split audio noise feature.
[0080] In this embodiment, first fusing and splitting the video feature, text feature, and audio noise feature can enable full interaction of various modal features, break the modal barrier, enhance the correlation and complementarity between features, generate an alignment feature based on the audio-visual synchronization feature to ensure the coordination of audio and video in time and semantics, generate a global feature based on multiple features to comprehensively integrate multi-dimensional information, fully present the core content of the video and text and the audio noise situation, and finally generate the audio corresponding to the video by combining various features. The rich features obtained in the previous processing process can be fully utilized, considering from multiple aspects, accurately generating an audio that fits the video content, is synchronized with the video picture, and conforms to the semantics of the audio description text, effectively improving the accuracy and fitting degree of the generated audio and optimizing the quality of the generated audio.
[0081] In some embodiments, the process of generating the alignment feature according to the audio-visual synchronization feature includes: upsampling the audio-visual synchronization feature to the same dimension as the audio noise feature to obtain the alignment feature.
[0082] In this embodiment, since the frame rates between the audio-visual synchronization feature and the audio noise feature are different, in order to align the timing relationship between the visual content and the audio, by upsampling the audio-visual synchronization feature to the same dimension as the audio noise feature to obtain the alignment feature, the scale unification between features of different dimensions can be realized, providing a basis for subsequent audio generation.
[0083] In some embodiments, the process of generating global features based on text features, video features, video duration features, and alignment features includes: pooling the text features, video features, and video duration features to obtain pooled features, where the pooled features are used to characterize the consistency between the audio description text and the video; generating global features based on the pooled features and the alignment features.
[0084] In this embodiment, by pooling the text features, video features, and video duration features to extract key information on multimodal consistency, and then combining the alignment features, global features that integrate the consistency between modalities and the adaptability of audio-visual synchronization dimensions can be generated, making the generated global features more accurate and providing comprehensive feature support for subsequent audio generation.
[0085] In some embodiments, the process of generating global features based on the pooled features and the alignment features includes: adding the pooled features and the alignment features to obtain global features.
[0086] In this embodiment, adding the pooled features and the alignment features to generate global features can incorporate information on the audio-visual synchronization dimension while retaining the multimodal consistency representation, achieving complementary fusion at the feature level to enhance the comprehensiveness of the global representation.
[0087] In some embodiments, the process of generating the audio corresponding to the video based on the alignment features, global features, split video features, split text features, and split audio noise features includes: determining the audio features of the audio corresponding to the video based on the alignment features, global features, split video features, split text features, and split audio noise features; generating the audio corresponding to the video based on the audio features.
[0088] In this embodiment, by comprehensively using the alignment features, global features, and split multimodal features to determine the audio features, the synchronization relationship between the video and the audio, the overall association of multimodal content, and the local detail information of each modality can be fully integrated. The alignment features ensure the coordination between the audio and the video in terms of time and semantics; the global features fuse the multimodal consistency and the adaptability of audio-visual synchronization. The split features retain the original information and interaction details of different modalities. Generating audio features based on this can accurately capture the key audio elements required by the video content, ensuring that the generated audio highly matches the video in all aspects, making the generated audio not only accurately match the video but also combine with the audio description text, resulting in a better audio effect.
[0089] In some embodiments, the process of determining the audio features of the audio corresponding to a video based on alignment features, global features, split video features, split text features, and split audio noise features includes: processing the alignment features, global features, split video features, split text features, and split audio noise features through an audio feature generation model to obtain the audio features of the audio corresponding to the video, where the audio feature generation model is used to generate audio features.
[0090] In this embodiment, by processing multi-dimensional features through an audio feature generation model to generate audio features, the audio-visual synchronization, multi-modal global association, and detailed information of each modality can be efficiently fused to accurately match the video content and generate suitable audio features, making the generated audio features more accurate and thus making the subsequent generated audio have a better effect.
[0091] The above Figure 2 The embodiments briefly introduce the process of audio generation. The following further introduces the process of audio generation based on Figure 3 the embodiments. Refer to Figure 3 Figure 3 FIG.
[0092] In step S301, the electronic device obtains the video for which the audio is to be generated.
[0093] In the embodiments of the present disclosure, the video for which the audio is to be generated can be any video, and the present disclosure does not limit this. A target application is installed in the electronic device, and the target application is used to generate audio. The target application can be any application capable of generating audio, and the present disclosure also does not limit the type of the target application. Exemplarily, the target application is a video editing application.
[0094] In some embodiments, relevant information of the target application is displayed on the electronic device, and the relevant information of the target application includes at least one of the name and icon of the target application. In response to a trigger operation on the relevant information of the target application, a video upload interface is displayed, and a video upload control for uploading a video is displayed in the video upload interface; in response to a trigger operation on the video upload control, a video selection interface is displayed, and at least one selectable video is displayed in the video selection interface; in response to a trigger operation on any one of the at least one selectable videos, the triggered selectable video is used as the video for which the audio is to be generated.
[0095] Among them, at least one selectable video displayed in the video selection interface is a video stored in the electronic device. The triggering operation for the relevant information of the target application may be a click operation on the relevant information of the target application, a double-click operation on the relevant information of the target application, or other operations on the relevant information of the target application. The embodiments of the present disclosure do not limit this. The process of the triggering operation for other content is similar to the process of the triggering operation for the relevant information of the target application, and the embodiments of the present disclosure do not limit this either.
[0096] In step S302, the electronic device determines the video feature, video duration feature, and audio-visual synchronization feature corresponding to the video. The video feature is used to characterize the video, the video duration feature is used to characterize the duration of the video, and the audio-visual synchronization feature is used to characterize the temporal consistency and semantic relevance between the video's picture and sound.
[0097] In some embodiments, after obtaining the video in the above steps, the video feature corresponding to the video may be determined first, then the video duration feature corresponding to the video, and finally the audio-visual synchronization feature corresponding to the video; or the video duration feature corresponding to the video may be determined first, then the video feature corresponding to the video, and finally the audio-visual synchronization feature corresponding to the video; or the audio-visual synchronization feature corresponding to the video may be determined first, then the video feature corresponding to the video, and finally the video duration feature corresponding to the video. The embodiments of the present disclosure do not limit the determination order of the video feature, video duration feature, and audio-visual synchronization feature corresponding to the video.
[0098] In some embodiments, the process for the electronic device to determine the video feature corresponding to the video includes: performing frame extraction on the video to obtain multiple pictures included in the video; determining the picture feature of each picture, where the picture feature of each picture is used to characterize the corresponding picture; and determining the video feature corresponding to the video according to the picture features of each picture.
[0099] In this embodiment, by decomposing the video into multiple pictures, extracting and integrating the picture features to obtain the video feature, it is possible to reduce the processing complexity while fusing the information in the spatial and temporal dimensions, improving the comprehensiveness, robustness, and adaptability of the video feature expression to different videos. Furthermore, the generated video feature has a higher matching degree with the video, and the accuracy of the video feature is higher.
[0100] Among them, the process of extracting frames from a video to obtain multiple pictures included in the video includes: extracting a target number of pictures from the pictures included in each second of the video to obtain multiple pictures included in the video. The target number is set based on experience or flexibly adjusted according to the implementation environment, and the embodiments of the present disclosure do not limit this. Exemplarily, if the target number is 8 and the duration of the video is 10 seconds, then 8 frames of pictures are extracted per second, and a total of 80 frames of pictures are extracted. These 80 frames of pictures are the multiple pictures included in the video.
[0101] In some embodiments, after extracting frames from a video to obtain multiple pictures included in the video, the process of determining the picture features of each picture includes: for any one of the pictures, processing the any one of the pictures through a picture feature acquisition model to obtain the picture features of the any one of the pictures, and the picture feature acquisition model is used to acquire picture features.
[0102] Among them, the picture feature acquisition model can be the visual encoder of Metaclip (Metadata-Curated Language-Image Pre–training, a language-image pre-training model) with a parameter scale of 2.5B (Billion), trained using the contrastive learning algorithm, or a visual VAE (Variational Autoencoder) model, or a visual representation model. The embodiments of the present disclosure do not limit this.
[0103] Optionally, input any one of the pictures into the picture feature acquisition model, and the output result of the picture feature acquisition model is the picture features of the any one of the pictures.
[0104] In this embodiment, by using the picture feature acquisition model to acquire picture features, the image information can be efficiently and accurately converted into a structured representation that can be understood by the electronic device, making the generated picture features more accurate. Since the picture features are used to generate video features, the accuracy of the generated video features can also be improved.
[0105] In some embodiments, the process of determining the video features corresponding to the video according to the picture features of each picture includes: splicing the picture features of each picture to obtain the video features corresponding to the video.
[0106] In some embodiments, the process of determining the video duration features corresponding to the video includes: obtaining the number of frames and the frame rate of the video; determining the duration of the video according to the number of frames and the frame rate of the video; and determining the video duration features corresponding to the video according to the duration of the video.
[0107] In this embodiment, by obtaining the number of frames and the frame rate of a video to determine the duration of the video and extract duration features, the time dimension information of the video can be converted into structured data in a quantitative manner, making the generated video duration features more accurate.
[0108] In some embodiments, the process of obtaining the number of frames and the frame rate of a video includes obtaining the number of frames and the frame rate of the video through the OpenCV tool (Open Source Computer Vision Library, an open-source computer vision and machine learning software library).
[0109] In some embodiments, the process of determining the duration of a video according to the number of frames and the frame rate of the video includes: determining the quotient between the number of frames and the frame rate of the video; in the case where the quotient is an integer, determining the quotient as the duration of the video; in the case where the quotient is a non-integer, rounding up the quotient to obtain the duration of the video.
[0110] In this embodiment, by calculating the quotient between the number of frames and the frame rate and determining the duration of the video according to the quotient, the duration of the video can be determined in a standardized calculation manner, ensuring the standardization and adaptability of the duration calculation.
[0111] Exemplarily, if the number of frames of a video is 3600 frames and the frame rate is 25 frames per second (fps), then the quotient between the number of frames and the frame rate of the video is 3600 / 25 = 144. Since the quotient between the number of frames and the frame rate of the video is an integer, therefore, the quotient between the number of frames and the frame rate of the video is used as the duration of the video, that is, the duration of the video is 144 seconds.
[0112] Another example, if the number of frames of a video is 3600 frames and the frame rate is 35 frames per second (fps), then the quotient between the number of frames and the frame rate of the video is 3600 / 35 ≈ 102.9. Since the quotient between the number of frames and the frame rate of the video is a non-integer, therefore, the quotient between the number of frames and the frame rate of the video is rounded up to get 103, that is, the duration of the video is 103 seconds.
[0113] In some embodiments, the process of determining the video duration feature corresponding to a video according to the duration of the video includes: mapping the duration of the video to an array; processing the array through a duration learning model to obtain the video duration feature corresponding to the video, and the duration learning model is used to determine the video duration feature.
[0114] In this embodiment, mapping the duration of the video to an array and processing it through a duration learning model can convert the duration information into structured features of the model science system, so as to adaptively extract deep semantic representations containing temporal laws to adapt to subsequent analysis tasks.
[0115] In some embodiments, the process of mapping the duration of a video to an array includes: mapping the duration of the video to an array of [512, 768].
[0116] In some embodiments, the duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer, and a third convolutional layer. Among them, the first convolutional layer, the second convolutional layer, and the third convolutional layer are all one-dimensional convolutional layers. The process of obtaining the video duration feature corresponding to the video by processing the array through the duration learning model includes: inputting the array into the fully connected layer to obtain the output result of the fully connected layer; respectively inputting the output result of the fully connected layer into the first convolutional layer and the second convolutional layer to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; mapping the output result of the first convolutional layer through an activation function to obtain an attention weight; multiplying the attention weight and the output result of the second convolutional layer to obtain a multiplication result; inputting the multiplication result into the third convolutional layer to obtain the video duration feature of the video.
[0117] Among them, the activation function can be the SiLu (Sigmoid Linear Unit, a commonly used activation function in deep learning) activation function.
[0118] In this embodiment, through the hierarchical processing of the fully connected layer and multiple convolutional layers, and by using the attention weight generated by the activation function to weight the output result of the convolutional layer, the key features related to the video duration in the array are efficiently captured, and the information unrelated to the video duration in the array is suppressed, so as to accurately extract the video duration feature.
[0119] Such as Figure 4 is a schematic diagram of a process for determining the video duration feature of a video provided by an embodiment of the present application. Input the array into the fully connected layer 401 to obtain the output result of the fully connected layer; input the output result of the fully connected layer into the first convolutional layer 402 and the second convolutional layer 403 to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; map the output result of the first convolutional layer through an activation function to obtain an attention weight; multiply the output result of the second convolutional layer and the attention weight to obtain a multiplication result; input the multiplication result into the third convolutional layer 404 to obtain the video duration feature of the video.
[0120] In some embodiments, the process of determining the audio-visual synchronization feature corresponding to the video includes: processing the video through an audio-visual synchronization encoder to obtain the audio-visual synchronization feature corresponding to the video, and the audio-visual synchronization encoder is used to obtain the audio-visual synchronization feature.
[0121] The audio-visual synchronization encoder decouples feature extraction and synchronization modeling by performing multi-modal segment-level contrast pre-training on the video, and realizes efficient audio-visual synchronization under sparse synchronization clues.
[0122] In some embodiments, the video is input into the audio-visual synchronization encoder to obtain the audio-visual synchronization features corresponding to the video.
[0123] In this embodiment, the audio-visual synchronization encoder jointly encodes the sound of the video and the picture of the video, which can effectively capture the temporal consistency and semantic relevance between the two, so as to accurately extract the audio-visual synchronization features reflecting the audio-visual coordination relationship.
[0124] In step S303, the electronic device obtains the audio description text, and the audio description text is the description text of the audio corresponding to the video.
[0125] The embodiments of the present disclosure do not limit the manner of obtaining the audio description text. Optionally, a plurality of optional texts are stored in the electronic device, and any one of the plurality of optional texts is used as the audio description text. Optionally, a text input box can also be displayed, and in response to the input operation of the user in the text input box, the input content is used as the audio description text.
[0126] Exemplarily, the audio description text is the sound of a racket hitting a ball.
[0127] The embodiments of the present disclosure do not limit the order of obtaining the video for which the audio is to be generated and obtaining the audio description text. Optionally, first obtain the video for which the audio is to be generated, and then obtain the audio description text, or first obtain the audio description text, and then obtain the video for which the audio is to be generated.
[0128] In step S304, the electronic device determines the text features corresponding to the audio description text.
[0129] In some embodiments, after obtaining the audio description text in the above steps, the process of determining the text features corresponding to the audio description text includes: obtaining the first feature corresponding to the audio description text, where the first feature is used to characterize the audio description text; performing keyword extraction on the audio description text to obtain at least one keyword; obtaining the second feature corresponding to each keyword, where the second feature corresponding to each keyword is used to characterize the corresponding keyword; and determining the text features corresponding to the audio description text according to the first feature and the second features corresponding to each keyword.
[0130] In this embodiment, by combining the overall semantic feature of the audio description text (that is, the first feature) and the key semantic features extracted from the audio description text (that is, the second feature), it is possible to more comprehensively and prominently capture the audio-related information in the audio description text, so as to generate more accurate and effective text features.
[0131] In some embodiments, the process of obtaining the first feature corresponding to the audio description text includes: processing the audio description text through a first text feature acquisition model to obtain the first feature corresponding to the audio description text, where the first text feature acquisition model is used to acquire text features.
[0132] Among them, inputting the audio description text into the first text feature acquisition model to obtain the first feature corresponding to the audio description text.
[0133] In some embodiments, the first text feature acquisition model may be a T5 (Text-To-Text Transfer Transformer) large text model or a GLM model (General Language Model). The embodiments of the present disclosure do not limit this. The T5 large text model supports a maximum of 512 text tokens and can be used for the understanding of long texts.
[0134] In this embodiment, processing the audio description text through the first text feature acquisition model can effectively extract the first feature that can accurately represent the overall semantic information of the audio description text, providing a basis for subsequent acquisition of text features.
[0135] In some embodiments, the process of extracting at least one keyword from the audio description text includes: performing word segmentation on the audio description text to obtain a plurality of words, and using the words that exist in the keyword library among the plurality of words as keywords.
[0136] In some embodiments, the process of obtaining the second feature corresponding to each keyword includes: for any one of the keywords, processing the any one keyword through a second text feature acquisition model to obtain the second feature corresponding to the any one keyword, where the second text feature acquisition model is used to acquire text features, and the number of tokens supported by the second text feature acquisition model is less than the number of tokens supported by the first text feature acquisition model.
[0137] Among them, inputting the any one keyword into the second text feature acquisition model to obtain the second feature corresponding to the any one keyword.
[0138] In some embodiments, the second text feature acquisition model may be the text encoder of MetaClip. The text encoder of MetaClip is a text encoder trained by using short texts and keywords in a contrastive learning manner, which can improve the retrieval and understanding ability of keywords. The text encoder of MetaClip supports 77 tokens and does not support the understanding of texts with a larger length.
[0139] In this embodiment, the keywords are processed by the second text feature acquisition model, so that the second features that can accurately characterize the keywords can be effectively extracted, providing a basis for the subsequent acquisition of text features.
[0140] In some embodiments, the process of determining the text feature corresponding to the audio description text based on the first feature and the second feature corresponding to each keyword includes: concatenating the first feature and the second feature corresponding to each keyword to obtain the text feature corresponding to the audio description text.
[0141] In step S305, the electronic device generates audio corresponding to the video according to the video features, video duration features, audio-visual synchronization features and text features, and the duration of the audio is the same as the duration of the video.
[0142] In some embodiments, the process of generating the audio corresponding to the video according to the video features, video duration features, audio-visual synchronization features and text features includes: fusing the video features, text features and audio noise features to obtain the fused features, the audio noise features and the audio features of the audio corresponding to the video having the same dimension; splitting the fused features to obtain split video features, split text features and split audio noise features, wherein the split video features, split text features and split audio noise features all contain the video features, text features and audio noise features; generating an alignment feature according to the audio-visual synchronization features, wherein the dimension of the alignment feature is the same as the dimension of the audio noise feature; generating a global feature according to the text features, video features, video duration features and alignment features, wherein the global feature is the feature after considering the text features, video features, video duration features and alignment features; generating the audio corresponding to the video according to the alignment features, the global features, the split video features, the split text features and the split audio noise features.
[0143] In this embodiment, the video features, text features and audio noise features are first merged and split, so that the features of each modality can fully interact, break down the modal barriers, enhance the correlation and complementarity between features, generate alignment features based on the audio-visual synchronization features, ensure the temporal and semantic coordination of audio and video, generate global features based on multiple features, comprehensively integrate multi-dimensional information, and fully present the core content of the video and text and the audio noise conditions. Finally, various features are combined to generate audio corresponding to the video, which can make full use of the rich features obtained in the previous processing process, comprehensively consider multiple aspects, and accurately generate audio that fits the video content, is synchronized with the video picture, and conforms to the semantics of the audio description text, effectively improve the accuracy and fit of the generated audio, and optimize the quality of the generated audio.
[0144] In some embodiments, the process of fusing video features, text features, and audio noise features to obtain fused features includes: fusing the video features, text features, and audio noise features through a cross-attention mechanism to obtain fused features. The video features, text features, and audio noise features are respectively the query, key, and value in the cross-attention mechanism. The fused features are the features after fusing the video features, text features, and audio noise features.
[0145] Among them, the dimension of the audio noise feature is set based on experience, or flexibly adjusted according to the implementation environment, and the embodiments of the present disclosure do not limit this.
[0146] In some embodiments, the process of splitting the fused features to obtain the split video features, split text features, and split audio noise features includes: splitting the fused features according to the fusion order when the video features, text features, and audio noise features are fused, to obtain the split video features, split text features, and split audio noise features. The dimension of the split video features is the same as that of the video features, the dimension of the split text features is the same as that of the text features; the dimension of the split audio noise features is the same as that of the audio noise features.
[0147] In some embodiments, the process of generating alignment features according to the audio-visual synchronization features includes: upsampling the audio-visual synchronization features to the same dimension as the audio noise features to obtain alignment features.
[0148] In this embodiment, since the frame rates between the audio-visual synchronization features and the audio noise features are different, in order to align the temporal relationship between the visual content and the audio, by upsampling the audio-visual synchronization features to the same dimension as the audio noise features to obtain alignment features, the scale unification between features of different dimensions can be achieved, providing a basis for subsequent audio generation.
[0149] In some embodiments, the process of generating global features according to the text features, video features, video duration features, and alignment features includes: performing pooling on the text features, video features, and video duration features to obtain pooled features, and the pooled features are used to characterize the consistency between the audio description text and the video; generating global features according to the pooled features and the alignment features.
[0150] In this implementation, by performing pooling on the text features, video features, and video duration features to extract key information of multimodal consistency, and then combining the alignment features, global features that integrate the consistency association between modalities and the adaptability of the audio-visual synchronization dimension can be generated, making the generated global features more accurate and providing comprehensive feature support for subsequent audio generation.
[0151] In some embodiments, the process of generating global features based on pooled features and aligned features includes: adding the pooled features and the aligned features to obtain global features.
[0152] In this embodiment, adding the pooled features and the aligned features to generate global features can incorporate the audio-visual synchronization dimension information while retaining the multi-modal consistency representation, achieving complementary fusion at the feature level to enhance the comprehensiveness of the global representation.
[0153] In some embodiments, the process of generating the audio corresponding to the video based on the aligned features, global features, split video features, split text features, and split audio noise features includes: determining the audio features of the audio corresponding to the video according to the aligned features, global features, split audio features, split text features, and split audio noise features; generating the audio corresponding to the video according to the audio features.
[0154] In this embodiment, by comprehensively using the aligned features, global features, and split multi-modal features to determine the audio features, it is possible to comprehensively integrate the synchronization relationship between the video and the audio, the overall correlation of the multi-modal content, and the local detailed information of each modality. The aligned features ensure the coordination of the audio and the video in terms of time and semantics; the global features fuse the multi-modal consistency and the audio-visual synchronization adaptability, and the split features retain the original information and interaction details of different modalities. On this basis, generating the audio features can accurately capture the key audio elements required by the video content, ensuring that the generated audio highly fits the video in all aspects, so that the generated audio can not only accurately match the video, but also combine with the audio description text, making the effect of the generated audio better.
[0155] In some embodiments, the process of determining the audio features of the audio corresponding to the video according to the aligned features, global features, split video features, split text features, and split audio noise features includes: processing the aligned features, global features, split video features, split text features, and split audio noise features through an audio feature generation model to obtain the audio features of the audio corresponding to the video, where the audio feature generation model is used to generate audio features.
[0156] In some embodiments, the aligned features, global features, split video features, split text features, and split audio noise features are input into the audio feature generation model through AdaLN (Adaptive Layer Normalization) and cross-attention mechanism to obtain the audio features of the audio corresponding to the video.
[0157] Among them, the audio feature generation model is a denoising-based multi-modal DIT (Denoising Diffusion Implicit Models) generation model.
[0158] In this embodiment, a multi-dimensional feature is processed by an audio feature generation model to generate an audio feature, which can efficiently fuse the audio-visual synchronization, multi-modal global association, and the detailed information of each modality, so as to accurately match the video content to generate a suitable audio feature, making the generated audio feature more accurate, and further making the subsequent generated audio have a better effect.
[0159] In some embodiments, the process of generating the audio corresponding to the video according to the audio feature includes: processing the audio feature through an audio encoder to obtain a Mel spectrogram corresponding to the audio feature; processing the Mel spectrogram corresponding to the audio feature through a vocoder to obtain the audio corresponding to the video.
[0160] Among them, the vocoder is any vocoder that supports music and sound effect generation, and the type of the vocoder is not limited in the embodiments of the present disclosure.
[0161] In some embodiments, the audio feature is input into an audio encoder to obtain a Mel spectrogram corresponding to the audio feature; the Mel spectrogram corresponding to the audio feature is input into a vocoder to obtain the audio corresponding to the video.
[0162] The audio generation method provided by the embodiments of the present disclosure can be used to dub videos of AIGC (Artificial Intelligence Generated Content) or UGC (User Generated Content), and at the same time process multi-source superposition scenarios (such as traffic sounds, human voices, construction sounds, etc. in urban streets). By decoupling multi-modal features, corresponding sound effects are generated respectively, replacing the cumbersome process of manual superposition one by one. Moreover, it supports the dual-condition generation of "audio description text + video", such as controlling the style of the footsteps of the characters in the video through the audio description text "suspense movie footsteps", providing more creative possibilities for film and television dubbing, and at the same time adapting to emerging scenarios such as game real-time sound effect generation and virtual anchor dubbing.
[0163] In some embodiments, the audio generation method provided by the embodiments of the present disclosure is a unified end-to-end multi-modal audio generation technology. The input is a video or a video and text, and the output is an audio. That is to say, the embodiments of the present disclosure can implement audio generation tasks such as video-audio and text joint video-audio. Moreover, the embodiments of the present disclosure can also be extended to support the generation of character dialogue voices. The embodiments of the present disclosure use an audio generation feature model to generate audio features, effectively realizing the fusion of information between modalities and improving the content relevance of audio generation. At the same time, the ability to align the video and audio at the beat is also improved. In addition, by adding a duration learning model, the generation of audio of various durations is effectively supported, increasing the flexibility of video dubbing.
[0164] Embodiments of the present disclosure provide an audio generation method. This method generates the audio corresponding to a video based on the video without manual operation by the user. Therefore, the audio generation method is simplified, and the audio generation efficiency is relatively high. Moreover, it does not rely on the subjective consciousness of the user, so that the generated audio has a relatively high matching degree with the video, and the audio generation efficiency is relatively good. In addition, since not only video features are obtained, but also video duration features and audio-visual synchronization features are obtained, when generating the audio, not only the video features of the video itself are considered, but also the video duration features and audio-visual synchronization features are considered, so that the duration of the generated audio is the same as the duration of the video, and the generated audio is synchronized with the video in terms of sound and picture, further improving the matching degree between the generated audio and the video and the audio generation efficiency.
[0165] In addition, an audio description text can be obtained. When generating the audio corresponding to the video, not only video features, video duration features, and audio-visual synchronization features are considered, but also text features corresponding to the audio description text are considered, so that the generated audio is not only an audio that matches the video, but also an audio that has a high matching degree with the audio description text, making the generated audio more accurate.
[0166] Figure 5 It is a flowchart of an audio generation method shown according to an exemplary embodiment. As Figure 5 shown, the method includes the following steps.
[0167] Step 501, the electronic device obtains the video and the audio description text of the audio to be generated.
[0168] Wherein, the audio description text is the description text of the audio corresponding to the video.
[0169] In some embodiments, the process of obtaining the video of the audio to be generated has been described in step S301 above, and the process of obtaining the audio description text has been described in step S303 above. Embodiments of the present disclosure will not repeat them here.
[0170] Step 502, the electronic device determines the video features, video duration features, and audio-visual synchronization features of the video.
[0171] Wherein, the video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the time consistency and semantic relevance between the picture and the sound of the video.
[0172] In some embodiments, the process of determining the video features, video duration features, and audio-visual synchronization features of the video has been described in step S302 above. Embodiments of the present disclosure will not repeat it here.
[0173] Step 503, the electronic device determines the text features corresponding to the audio description text.
[0174] In some embodiments, the process of determining the text features corresponding to the audio description text has been described in step S304 above, and will not be elaborated in the embodiments of the present disclosure herein.
[0175] Step 504: The electronic device fuses the video features, text features, and audio noise features to obtain fused features.
[0176] Among them, the audio noise features and the audio features of the audio corresponding to the video have the same dimension. The fused features are features that fuse the video features, text features, and audio noise features.
[0177] In some embodiments, the process of fusing the video features, text features, and audio noise features to obtain fused features has been described in step S305 above, and will not be elaborated in the embodiments of the present disclosure herein.
[0178] Step 505: The electronic device splits the fused features to obtain the split video features, split text features, and split audio noise features.
[0179] Among them, the split video features, split text features, and split audio noise features all fuse the video features, text features, and audio noise features.
[0180] In some embodiments, the process of splitting the fused features to obtain the split video features, split text features, and split audio noise features has been described in step S305 above, and will not be elaborated in the embodiments of the present disclosure herein.
[0181] Step 506: The electronic device generates alignment features according to the audio-visual synchronization features.
[0182] Among them, the dimension of the alignment features is the same as that of the audio noise features.
[0183] In some embodiments, the process of generating alignment features according to the audio-visual synchronization features has been described in step S305 above, and will not be elaborated in the embodiments of the present disclosure herein.
[0184] Step 507: The electronic device generates global features according to the text features, video features, video duration features, and alignment features.
[0185] Among them, the global features are features that consider the text features, video features, video duration features, and alignment features.
[0186] In some embodiments, the process of generating global features according to the text features, video features, video duration features, and alignment features has been described in step S305 above, and will not be elaborated in the embodiments of the present disclosure herein.
[0187] Step 508: The electronic device processes the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature through an audio feature generation model to obtain the audio feature of the audio corresponding to the video.
[0188] Among them, the audio feature generation model is used to generate audio features.
[0189] In some embodiments, the process of processing the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature through the audio feature generation model to obtain the audio feature of the audio corresponding to the video has been described in the above step S305, and the embodiments of the present disclosure do not limit this.
[0190] Step 509: The electronic device processes the audio feature through an audio encoder to obtain the Mel spectrogram corresponding to the audio feature.
[0191] In some embodiments, the process of processing the audio feature through the audio encoder to obtain the Mel spectrogram corresponding to the audio feature has been described in the above step S305, and the embodiments of the present disclosure do not limit this.
[0192] Step 510: The electronic device processes the Mel spectrogram corresponding to the audio feature through a vocoder to obtain the audio corresponding to the video.
[0193] In some embodiments, the process of processing the Mel spectrogram corresponding to the audio feature through the vocoder to obtain the audio corresponding to the video has been described in the above step S305, and the embodiments of the present disclosure do not limit this.
[0194] Figure 6 is a flowchart of an audio generation method shown according to an exemplary embodiment. As Figure 6 shown, the method includes: performing frame extraction on the video of the audio to be generated to obtain multiple pictures, processing each picture through a picture feature acquisition model to obtain the picture feature of each picture, and determining the video feature of the video according to the picture features of each picture. Processing the video through a duration learning model to obtain the video duration feature of the video. Processing the video through an audio-visual synchronization encoder to obtain the audio-visual synchronization feature of the video.
[0195] Processing the audio description text through a first text feature acquisition model to obtain the first feature corresponding to the audio description text. Extracting keywords from the audio description text to obtain multiple keywords; processing each keyword through a second text feature acquisition model to obtain the second feature corresponding to each keyword. Obtaining the text feature of the audio description text according to the first feature and the second features corresponding to each keyword.
[0196] Fuse and split the video features, text features, and audio noise features to obtain the split video features, split text features, and split audio noise features.
[0197] Determine the alignment features and global features based on the video features, text features, video duration features, and audio-visual synchronization features.
[0198] Process the alignment features, global features, split video features, split text features, and split audio noise features through an audio feature generation model to obtain the audio features of the audio corresponding to the video.
[0199] Process the audio features through an audio encoder to obtain the Mel spectrogram corresponding to the audio features.
[0200] Process the Mel spectrogram through a vocoder to obtain the audio corresponding to the video.
[0201] Figure 7 It is a block diagram of an audio generation device shown according to an exemplary embodiment. Refer to Figure 7 , the device includes the following.
[0202] An acquisition unit 701, configured to acquire the video for which the audio is to be generated.
[0203] A determination unit 702, configured to determine the video features, video duration features, and audio-visual synchronization features corresponding to the video, where the video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance between the video image and sound.
[0204] A generation unit 703, configured to generate the audio corresponding to the video according to the video features, video duration features, and audio-visual synchronization features, and the duration of the audio is the same as the duration of the video.
[0205] In some embodiments, the acquisition unit 701 is further configured to acquire an audio description text, where the audio description text is the description text of the audio corresponding to the video.
[0206] The determination unit 702 is further configured to determine the text features corresponding to the audio description text.
[0207] The generation unit 703 is configured to generate the audio corresponding to the video according to the video features, video duration features, audio-visual synchronization features, and text features.
[0208] In some embodiments, the determining unit 702 is configured to execute: obtaining a first feature corresponding to the audio description text, where the first feature is used to characterize the audio description text; performing keyword extraction on the audio description text to obtain at least one keyword; obtaining a second feature corresponding to each keyword, where the second feature corresponding to each keyword is used to characterize the corresponding keyword; and determining a text feature corresponding to the audio description text according to the first feature and the second features corresponding to the respective keywords.
[0209] In some embodiments, the obtaining unit 701 is configured to execute: processing the audio description text through a first text feature obtaining model to obtain a first feature corresponding to the audio description text, where the first text feature obtaining model is used to obtain text features.
[0210] In some embodiments, for any one of the respective keywords, the obtaining unit 701 is configured to execute: processing any one keyword through a second text feature obtaining model to obtain a second feature corresponding to any one keyword, where the second text feature obtaining model is used to obtain text features, and the number of tokens supported by the second text feature obtaining model is less than the number of tokens supported by the first text feature obtaining model.
[0211] In some embodiments, the generating unit 703 is configured to execute: fusing the video feature, the text feature, and the audio noise feature to obtain a fused feature, where the audio noise feature and the audio feature of the audio corresponding to the video have the same dimension; splitting the fused feature to obtain a split video feature, a split text feature, and a split audio noise feature, where the split video feature, the split text feature, and the split audio noise feature are all fused with the video feature, the text feature, and the audio noise feature; generating an alignment feature according to the audio-visual synchronization feature, where the dimension of the alignment feature is the same as the dimension of the audio noise feature; generating a global feature according to the text feature, the video feature, the video duration feature, and the alignment feature, where the global feature is a feature after considering the text feature, the video feature, the video duration feature, and the alignment feature; and generating the audio corresponding to the video according to the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature.
[0212] In some embodiments, the generating unit 703 is configured to execute: upsampling the audio-visual synchronization feature to the same dimension as the audio noise feature to obtain an alignment feature.
[0213] In some embodiments, the generating unit 703 is configured to execute: performing pooling on the text feature, the video feature, and the video duration feature to obtain a pooled feature, where the pooled feature is used to characterize the consistency between the audio description text and the video; and generating a global feature according to the pooled feature and the alignment feature.
[0214] In some embodiments, the generating unit 703 is configured to add the pooled feature and the aligned feature to obtain a global feature.
[0215] In some embodiments, the generating unit 703 is configured to determine an audio feature of the audio corresponding to the video according to the aligned feature, the global feature, the split video feature, the split text feature, and the split audio noise feature; and generate the audio corresponding to the video according to the audio feature.
[0216] In some embodiments, the generating unit 703 is configured to process the aligned feature, the global feature, the split video feature, the split text feature, and the split audio noise feature through an audio feature generation model to obtain an audio feature of the audio corresponding to the video, and the audio feature generation model is used to generate an audio feature.
[0217] In some embodiments, the determining unit 702 is configured to perform frame extraction on the video to obtain multiple pictures included in the video; determine a picture feature of each picture, and the picture feature of each picture is used to characterize the corresponding picture; and determine a video feature corresponding to the video according to the picture features of the pictures.
[0218] In some embodiments, for any one of the pictures, the determining unit 702 is configured to process the any one of the pictures through a picture feature acquisition model to obtain a picture feature of the any one of the pictures, and the picture feature acquisition model is used to acquire a picture feature.
[0219] In some embodiments, the determining unit 702 is configured to obtain the number of frames and the frame rate of the video; determine the duration of the video according to the number of frames and the frame rate of the video; and determine a video duration feature corresponding to the video according to the duration of the video.
[0220] In some embodiments, the determining unit 702 is configured to determine the quotient between the number of frames and the frame rate of the video; in the case where the quotient is an integer, determine that the quotient is the duration of the video; and in the case where the quotient is a non-integer, round up the quotient to obtain the duration of the video.
[0221] In some embodiments, the determining unit 702 is configured to map the duration of the video to an array; process the array through a duration learning model to obtain a video duration feature corresponding to the video, and the duration learning model is used to determine a video duration feature.
[0222] In some embodiments, the duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer, and a third convolutional layer; a determination unit 702 configured to perform operations of inputting an array into the fully connected layer to obtain an output result of the fully connected layer; inputting the output result of the fully connected layer into the first convolutional layer and the second convolutional layer respectively to obtain an output result of the first convolutional layer and an output result of the second convolutional layer; mapping the output result of the first convolutional layer through an activation function to obtain attention weights; multiplying the attention weights and the output result of the second convolutional layer to obtain a multiplication result; and inputting the multiplication result into the third convolutional layer to obtain video duration features of the video.
[0223] In some embodiments, the determination unit 702 is configured to perform operations of processing the video through an audio-visual synchronization encoder to obtain audio-visual synchronization features corresponding to the video, where the audio-visual synchronization encoder is used to acquire the audio-visual synchronization features.
[0224] The embodiments of the present disclosure provide an audio generation device. The device generates audio corresponding to a video according to the video, without manual operation by the user. Therefore, the audio generation method is simplified, and the audio generation efficiency is relatively high. Moreover, it does not depend on the subjective consciousness of the user, so that the generated audio has a relatively high matching degree with the video, and the audio generation efficiency is relatively good. In addition, since not only video features are acquired, but also video duration features and audio-visual synchronization features are acquired, when generating audio, not only the video features of the video itself are considered, but also the video duration features and audio-visual synchronization features are considered, so that the duration of the generated audio is the same as the duration of the video, and the generated audio is synchronized with the video in terms of sound and picture, further improving the matching degree between the generated audio and the video and the audio generation efficiency.
[0225] Regarding the device in the above embodiments, the specific manners in which each unit performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0226] Figure 8 The block diagram of a terminal 800 provided by an exemplary embodiment of the present disclosure is shown. The terminal 800 may be: a smart phone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 (Moving Picture Experts Group Audio Layer IV) player, a laptop computer, or a desktop computer. The terminal 800 may also be referred to by other names such as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, etc.
[0227] Generally, the terminal 800 includes a processor 801 and a memory 802.
[0228] The processor 801 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 801 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), or PLA (Programmable Logic Array). The processor 801 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 801 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 801 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0229] The memory 802 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 802 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 802 is used to store at least one program code, and the at least one program code is used to be executed by the processor 801 to implement the audio generation method provided in the method embodiments of the present disclosure.
[0230] In some embodiments, the terminal 800 may further optionally include: a peripheral device interface 803 and at least one peripheral device. The processor 801, the memory 802, and the peripheral device interface 803 may be connected by a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 803 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.
[0231] The peripheral device interface 803 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on separate chips or circuit boards, and this embodiment does not limit this.
[0232] The radio frequency circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 804 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts the received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 804 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and so on. The radio frequency circuit 804 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 804 may also include circuits related to NFC (Near Field Communication), and this disclosure does not limit this.
[0233] The display screen 805 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 805 is a touch display screen, the display screen 805 also has the ability to collect touch signals on or above the surface of the display screen 805. The touch signals can be input as control signals to the processor 801 for processing. At this time, the display screen 805 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments, there can be one display screen 805, which is provided on the front panel of the terminal 800; in other embodiments, there can be at least two display screens 805, which are respectively provided on different surfaces of the terminal 800 or are in a folding design; in still other embodiments, the display screen 805 can be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 800. Even further, the display screen 805 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 805 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).
[0234] The camera module 806 is used to capture images or videos. Optionally, the camera module 806 includes a front camera and a rear camera. Generally, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera respectively, to implement functions such as the combination of the main camera and the depth-of-field camera to achieve the background blurring function, the combination of the main camera and the wide-angle camera to achieve panoramic shooting and VR (Virtual Reality) shooting functions or other combined shooting functions. In some embodiments, the camera module 806 can also include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0235] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 801 for processing, or input to the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 800. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 807 may also include a headphone jack.
[0236] The power supply 808 is used to supply power to each component in the terminal 800. The power supply 808 may be alternating current, direct current, a primary battery or a rechargeable battery. When the power supply 808 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.
[0237] Those skilled in the art can understand that Figure 8 the structure shown in
[0238] Figure 9 FIG. shows a structural block diagram of a server 900 provided by an exemplary embodiment of the present disclosure. The server 900 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 901 and one or more memories 902. Among them, at least one program code is stored in the one or more memories 902, and the at least one program code is loaded and executed by the one or more processors 901 to implement the audio generation method provided by each of the above method embodiments. Of course, the server 900 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server 900 may also include other components for implementing device functions, which will not be elaborated here.
[0239] In an exemplary embodiment, there is also provided a computer-readable storage medium including instructions, such as a memory including instructions. The above instructions may be executed by a processor of a terminal to complete the above audio generation method. Optionally, the computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0240] In an exemplary embodiment, a computer program product is further provided, including a computer program which, when executed by a processor, implements the above-mentioned audio generation method. In some embodiments, the computer program product involved in the embodiments of the present disclosure may be deployed to be executed on a terminal, or on multiple terminals located at one place, or on multiple terminals distributed at multiple places and interconnected through a communication network. The multiple terminals distributed at multiple places and interconnected through a communication network may form a blockchain system.
[0241] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims. All the above optional technical solutions can be combined arbitrarily to form optional embodiments of the present application, which will not be elaborated herein one by one.
[0242] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An audio generation method, characterized in that, The method includes: Obtaining a video for which an audio is to be generated; Determining a video feature, a video duration feature, and an audio-visual synchronization feature corresponding to the video, where the video feature is used to characterize the video, the video duration feature is used to characterize the duration of the video, and the audio-visual synchronization feature is used to characterize the temporal consistency and semantic relevance between the video's picture and sound; Generating an audio corresponding to the video according to the video feature, the video duration feature, and the audio-visual synchronization feature, where the duration of the audio is the same as the duration of the video.
2. The method according to claim 1, wherein The method further includes: Obtaining an audio description text, where the audio description text is a description text for the audio corresponding to the video; Determining a text feature corresponding to the audio description text; The step of generating an audio corresponding to the video according to the video feature, the video duration feature, and the audio-visual synchronization feature includes: Generating an audio corresponding to the video according to the video feature, the video duration feature, the audio-visual synchronization feature, and the text feature.
3. The method according to claim 2, characterized in that, The step of determining a text feature corresponding to the audio description text includes: Obtaining a first feature corresponding to the audio description text, where the first feature is used to characterize the audio description text; Performing keyword extraction on the audio description text to obtain at least one keyword; Obtaining a second feature corresponding to each keyword, where the second feature corresponding to each keyword is used to characterize the corresponding keyword; Determining a text feature corresponding to the audio description text according to the first feature and the second features corresponding to the respective keywords.
4. The method according to claim 3, characterized in that, The step of obtaining a first feature corresponding to the audio description text includes: Processing the audio description text through a first text feature acquisition model to obtain a first feature corresponding to the audio description text, where the first text feature acquisition model is used to acquire text features.
5. The method according to claim 3, characterized in that The step of obtaining a second feature corresponding to each keyword includes: For any one keyword among the respective keywords, processing the any one keyword through a second text feature acquisition model to obtain a second feature corresponding to the any one keyword, where the second text feature acquisition model is used to acquire text features, and the number of tokens supported by the second text feature acquisition model is less than the number of tokens supported by the first text feature acquisition model.
6. The method according to claim 2, wherein The step of generating an audio corresponding to the video according to the video feature, the video duration feature, the audio-visual synchronization feature, and the text feature includes: Fusing the video feature, the text feature, and an audio noise feature to obtain a fused feature, where the audio noise feature has the same dimension as the audio feature of the audio corresponding to the video; Splitting the fused feature to obtain a split video feature, a split text feature, and a split audio noise feature, where the split video feature, the split text feature, and the split audio noise feature are all fused with the video feature, the text feature, and the audio noise feature; Generating an alignment feature according to the audio-visual synchronization feature, where the dimension of the alignment feature is the same as the dimension of the audio noise feature; Generate a global feature based on the text feature, the video feature, the video duration feature, and the alignment feature, where the global feature is the feature after considering the text feature, the video feature, the video duration feature, and the alignment feature; Generate the audio corresponding to the video based on the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature.
7. The method according to claim 6, wherein The generating of the alignment feature according to the audiovisual synchronization feature includes: Upsample the audiovisual synchronization feature to the same dimension as the audio noise feature to obtain the alignment feature.
8. The method according to claim 6, wherein The generating of the global feature according to the text feature, the video feature, the video duration feature, and the alignment feature includes: Perform pooling on the text feature, the video feature, and the video duration feature to obtain a pooled feature, where the pooled feature is used to characterize the consistency between the audio description text and the video; Generate the global feature according to the pooled feature and the alignment feature.
9. The method according to claim 8, characterized in that, The generating of the global feature according to the pooled feature and the alignment feature includes: Add the pooled feature and the alignment feature to obtain the global feature.
10. The method according to claim 6, wherein The generating of the audio corresponding to the video according to the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature includes: Determine the audio feature of the audio corresponding to the video according to the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature; Generate the audio corresponding to the video according to the audio feature.
11. The method according to claim 10, characterized in that, The determining of the audio feature of the audio corresponding to the video according to the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature includes: Process the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature through an audio feature generation model to obtain the audio feature of the audio corresponding to the video, where the audio feature generation model is used to generate audio features.
12. The method according to any one of claims 1 to 11, characterized in that Determine the video feature corresponding to the video, including: Perform frame extraction on the video to obtain multiple pictures included in the video; Determine the picture feature of each picture, where the picture feature of each picture is used to characterize the corresponding picture; Determine the video feature corresponding to the video according to the picture features of the pictures.
13. The method according to claim 12, characterized in that, The determining of the picture feature of each picture includes: For any one of the pictures, process the any one picture through a picture feature acquisition model to obtain the picture feature of the any one picture, where the picture feature acquisition model is used to acquire picture features.
14. The method according to any one of claims 1 to 11, characterized in that, Determine the video duration feature corresponding to the video, including: Obtain the number of frames and the frame rate of the video; Determine the duration of the video according to the number of frames and the frame rate of the video; Determine the video duration feature corresponding to the video according to the duration of the video.
15. The method according to claim 14, wherein Determining the duration of the video according to the number of frames and the frame rate of the video includes: Determining the quotient between the number of frames and the frame rate of the video; When the quotient is an integer, determining that the quotient is the duration of the video; When the quotient is a non-integer, rounding up the quotient to obtain the duration of the video.
16. The method according to claim 14, wherein Determining the video duration feature corresponding to the video according to the duration of the video includes: Mapping the duration of the video into an array; Processing the array through a duration learning model to obtain the video duration feature corresponding to the video, where the duration learning model is used to determine video duration features.
17. The method according to claim 16, characterized in that, The duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer, and a third convolutional layer; Processing the array through the duration learning model to obtain the video duration feature corresponding to the video includes: Inputting the array into the fully connected layer to obtain the output result of the fully connected layer; Inputting the output result of the fully connected layer into the first convolutional layer and the second convolutional layer respectively to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; Mapping the output result of the first convolutional layer through an activation function to obtain an attention weight; Multiplying the attention weight and the output result of the second convolutional layer to obtain a multiplication result; Inputting the multiplication result into the third convolutional layer to obtain the video duration feature of the video.
18. The method according to any one of claims 1 to 11, characterized in that, Determining the audio-visual synchronization feature corresponding to the video includes: Processing the video through an audio-visual synchronization encoder to obtain the audio-visual synchronization feature corresponding to the video, where the audio-visual synchronization encoder is used to obtain audio-visual synchronization features.
19. An audio generation device, characterized in that, The device includes: An acquisition unit configured to acquire a video of the audio to be generated; A determination unit configured to determine the video feature, video duration feature, and audio-visual synchronization feature corresponding to the video, where the video feature is used to characterize the video, the video duration feature is used to characterize the duration of the video, and the audio-visual synchronization feature is used to characterize the temporal consistency and semantic relevance between the video's picture and sound; A generation unit configured to generate the audio corresponding to the video according to the video feature, the video duration feature, and the audio-visual synchronization feature, where the duration of the audio is the same as the duration of the video.
20. An electronic device, characterized in that, Includes: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the audio generation method according to any one of claims 1 to 18.
21. A computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the audio generation method according to any one of claims 1 to 18.
22. A computer program product, characterized in that, The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the audio generation method according to any one of claims 1 to 18.
Citation Information
Patent Citations
Silent video mimicking method, electronic equipment and storage medium
CN118828050A
Audio generation method and device, electronic equipment and storage medium
CN119255028A
Audio and video generation method, electronic equipment and computer readable storage medium
CN119316678A
Audio generation method
CN119364129A
Synthetic audio output method and apparatus, storage medium, and electronic device
US12051400B1