Audio generation method, device, equipment and computer-readable storage medium

By acquiring video features, duration features and audio-visual synchronization features, automatically generate audio with the same duration as the video, solving the problems of low audio generation efficiency and poor matching in the prior art, and achieving efficient and high-quality audio generation.

CN120358378BActive Publication Date: 2025-09-02BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510853769.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-02
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the prior art, when generating video audio through an artificial intelligence model, users need to operate manually, resulting in low matching between audio and video, poor generation effect and low efficiency.

Method used

By obtaining the video features, video duration features and audio-visual synchronization features of the video, audio and video synchronization features are generated, and audio generation is automatically generated using the audio generation device, combining audio description text features to improve matching.

Benefits of technology

It realizes high matching degree and efficient generation between audio and video, simplifies the generation process, and improves the efficiency and effect of audio generation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120358378B_ABST
    Figure CN120358378B_ABST
Patent Text Reader

Abstract

The present disclosure relates to an audio generation method, device, equipment and computer-readable storage medium, and relates to the field of artificial intelligence technology. The method includes: obtaining a video for which audio is to be generated; determining the video features, video duration features and audio-visual synchronization features corresponding to the video, wherein the video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance of the image and sound of the video; based on the video features, the video duration features and the audio-visual synchronization features, generating audio corresponding to the video, wherein the duration of the audio is the same as the duration of the video. This method ensures a high degree of matching between the generated audio and the video, thereby improving the efficiency of audio generation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an audio generation method, apparatus, device, and computer-readable storage medium. Background Art

[0002] With the continuous development of artificial intelligence technology, video generation methods are becoming increasingly diverse. For example, videos can be automatically generated using AI models. However, videos generated by AI models are silent, requiring dubbing to produce the corresponding audio.

[0003] In the related art, a user dubs a silent video through video editing software to obtain the audio corresponding to the silent video.

[0004] However, this method requires manual operation by the user and relies on the user's subjective consciousness, resulting in a low match between the generated audio and video, poor audio generation effect, and a more complicated audio generation method, which in turn leads to low audio generation efficiency. Summary of the Invention

[0005] The present disclosure provides an audio generation method, apparatus, device, and computer-readable storage medium. The method improves the matching degree between the generated audio and video, improves the audio generation effect, simplifies the audio generation method, and improves the audio generation efficiency. The technical solution of the present disclosure is as follows.

[0006] According to one aspect of an embodiment of the present disclosure, a method for generating audio is provided, the method comprising: obtaining a video for which audio is to be generated; determining video features, video duration features, and audio-visual synchronization features corresponding to the video, the video features being used to characterize the video, the video duration features being used to characterize the duration of the video, and the audio-visual synchronization features being used to characterize the temporal consistency and semantic relevance of the picture and sound of the video; generating audio corresponding to the video based on the video features, the video duration features, and the audio-visual synchronization features, the duration of the audio being the same as the duration of the video.

[0007] According to another aspect of an embodiment of the present disclosure, there is provided an audio generation apparatus, the apparatus including: an acquisition unit configured to acquire a video for generating audio.

[0008] The determination unit is configured to determine the video features, video duration features and audio-visual synchronization features corresponding to the video, wherein the video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance of the picture and sound of the video.

[0009] The generating unit is configured to generate audio corresponding to the video according to the video feature, the video duration feature and the audio-visual synchronization feature, wherein the duration of the audio is the same as the duration of the video.

[0010] In some embodiments, the acquisition unit is further configured to execute acquisition of audio description text, where the audio description text is a description text of the audio corresponding to the video.

[0011] The determining unit is further configured to determine text features corresponding to the audio description text.

[0012] The generating unit is configured to generate audio corresponding to the video according to the video feature, the video duration feature, the audio-visual synchronization feature and the text feature.

[0013] In some embodiments, the determination unit is configured to execute the steps of obtaining a first feature corresponding to the audio description text, where the first feature is used to characterize the audio description text; performing keyword extraction on the audio description text to obtain at least one keyword; obtaining a second feature corresponding to each keyword, where the second feature corresponding to each keyword is used to characterize the corresponding keyword; and determining the text feature corresponding to the audio description text based on the first feature and the second feature corresponding to each keyword.

[0014] In some embodiments, the acquisition unit is configured to process the audio description text through a first text feature acquisition model to obtain a first feature corresponding to the audio description text, where the first text feature acquisition model is used to acquire text features.

[0015] In some embodiments, the acquisition unit is configured to process any one of the keywords through a second text feature acquisition model to obtain a second feature corresponding to the keyword, the second text feature acquisition model is used to acquire text features, and the number of tokens supported by the second text feature acquisition model is less than the number of tokens supported by the first text feature acquisition model.

[0016] In some embodiments, the generation unit is configured to perform a fusion of the video features, the text features and the audio noise features to obtain a fusion feature, wherein the audio noise feature has the same dimension as the audio feature of the audio corresponding to the video; split the fusion feature to obtain split video features, split text features and split audio noise features, wherein the split video features, the split text features and the split audio noise features are all fused with the video features, the text features and the audio noise features; generate an alignment feature based on the audio-visual synchronization feature, wherein the dimension of the alignment feature is the same as the dimension of the audio noise feature; generate a global feature based on the text feature, the video feature, the video duration feature and the alignment feature, wherein the global feature is a feature after considering the text feature, the video feature, the video duration feature and the alignment feature; generate the audio corresponding to the video based on the alignment feature, the global feature, the split video feature, the split text feature and the split audio noise feature.

[0017] In some embodiments, the generating unit is configured to up-sample the audio-visual synchronization feature to the same dimension as the audio noise feature to obtain the alignment feature.

[0018] In some embodiments, the generation unit is configured to perform pooling of the text features, the video features, and the video duration features to obtain pooled features, wherein the pooled features are used to characterize the consistency between the audio description text and the video; and generate the global features based on the pooled features and the alignment features.

[0019] In some embodiments, the generating unit is configured to add the pooling feature and the alignment feature to obtain the global feature.

[0020] In some embodiments, the generation unit is configured to determine the audio features of the audio corresponding to the video based on the alignment features, the global features, the split video features, the split text features and the split audio noise features; and generate the audio corresponding to the video based on the audio features.

[0021] In some embodiments, the generation unit is configured to process the alignment features, the global features, the split video features, the split text features and the split audio noise features through an audio feature generation model to obtain audio features of the audio corresponding to the video, and the audio feature generation model is used to generate audio features.

[0022] In some embodiments, the determination unit is configured to perform frame extraction processing on the video to obtain multiple pictures included in the video; determine the picture features of each picture, and the picture features of each picture are used to represent the corresponding picture; and determine the video features corresponding to the video based on the picture features of each picture.

[0023] In some embodiments, the determination unit is configured to process any one of the pictures through a picture feature acquisition model to obtain picture features of the picture, and the picture feature acquisition model is used to obtain picture features.

[0024] In some embodiments, the determination unit is configured to obtain the number of frames and frame rate of the video; determine the duration of the video based on the number of frames and frame rate of the video; and determine the video duration feature corresponding to the video based on the duration of the video.

[0025] In some embodiments, the determination unit is configured to determine the quotient between the number of frames and the frame rate of the video; when the quotient is an integer, determine the quotient as the duration of the video; when the quotient is a non-integer, round up the quotient to obtain the duration of the video.

[0026] In some embodiments, the determination unit is configured to map the duration of the video into an array; process the array through a duration learning model to obtain a video duration feature corresponding to the video, and the duration learning model is used to determine the video duration feature.

[0027] In some embodiments, the duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer and a third convolutional layer; the determination unit is configured to input the array into the fully connected layer to obtain the output result of the fully connected layer; input the output result of the fully connected layer into the first convolutional layer and the second convolutional layer respectively to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; map the output result of the first convolutional layer through an activation function to obtain an attention weight; dot-multiply the attention weight and the output result of the second convolutional layer to obtain a dot-multiplication result; input the dot-multiplication result into the third convolutional layer to obtain the video duration feature of the video.

[0028] In some embodiments, the determining unit is configured to process the video through an audio-visual synchronization encoder to obtain an audio-visual synchronization feature corresponding to the video, where the audio-visual synchronization encoder is used to obtain the audio-visual synchronization feature.

[0029] According to another aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: one or more processors; a memory for storing program codes executable by the processors; wherein the processors are configured to execute the program codes to implement the above-mentioned audio generation method.

[0030] According to another aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When a program code in the computer-readable storage medium is executed by a processor of a terminal, the electronic device is enabled to perform the above-mentioned audio generation method.

[0031] According to another aspect of an embodiment of the present disclosure, a computer program product is provided, including a computer program, which implements the above-mentioned audio generation method when executed by a processor.

[0032] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure.

[0033] The disclosed embodiments provide an audio generation method, which generates audio corresponding to a video based on the video without manual operation by the user, thereby simplifying the audio generation method and making the audio generation more efficient. Moreover, there is no need to rely on the user's subjective consciousness, so that the matching degree between the generated audio and the video is higher, and the audio generation efficiency is better. In addition, since not only the video features but also the video duration features and the audio-visual synchronization features are obtained, when generating audio, not only the video features of the video itself but also the video duration features and the audio-visual synchronization features are taken into account, so that the duration of the generated audio is the same as the duration of the video, and the generated audio is synchronized with the video audio and video, further improving the matching degree between the generated audio and the video and the audio generation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.

[0035] Figure 1 It is a schematic diagram showing an implementation environment according to an exemplary embodiment.

[0036] Figure 2 The figure is a flowchart of an audio generation method according to an exemplary embodiment.

[0037] Figure 3 The figure is a flowchart of another audio generation method according to an exemplary embodiment.

[0038] Figure 4 The figure is a schematic diagram showing a process of determining a video duration feature of a video according to an exemplary embodiment.

[0039] Figure 5 is a flowchart of another audio generation method according to an exemplary embodiment.

[0040] Figure 6 is a flowchart of yet another audio generation method according to an exemplary embodiment.

[0041] Figure 7 The figure is a block diagram of an audio generating apparatus according to an exemplary embodiment.

[0042] Figure 8 It is a block diagram of a terminal according to an exemplary embodiment.

[0043] Figure 9 The figure is a block diagram of a server according to an exemplary embodiment. DETAILED DESCRIPTION

[0044] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0045] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.

[0046] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, storage, and display, etc.), and signals involved in this disclosure are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the videos and texts involved in this disclosure were obtained with full authorization.

[0047] The audio generation method provided by the embodiment of the present disclosure can be executed by an electronic device. Figure 1 This is a schematic diagram of an implementation environment provided by the embodiment of the present disclosure, see Figure 1, the implementation environment includes: an electronic device 101. The electronic device 101 can be a terminal or a server, and the embodiment of the present disclosure does not limit this. A target application is installed in the electronic device 101, and the target application is used to generate audio. In the embodiment of the present disclosure, the electronic device 101 obtains the video for which the audio is to be generated, determines the video features, video duration features and audio-visual synchronization features corresponding to the video, the video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance of the video's picture and sound; based on the video features, video duration features and audio-visual synchronization features, the audio corresponding to the video is generated, and the duration of the audio is the same as the duration of the video. The target application can be any application with an audio generation function, such as a short video application, a social application, etc., which is not specifically limited here.

[0048] Optionally, electronic device 101 is a terminal, which can be at least one of a smartphone, smartwatch, desktop computer, laptop, virtual reality terminal, augmented reality terminal, wireless terminal, and laptop computer. The terminal has communication capabilities and can access a wired or wireless network. The term "terminal" can generally refer to one of multiple terminals, and those skilled in the art will appreciate that the number of terminals can be greater or lesser. Electronic device 101 is a server, which can be a standalone physical server, a server cluster consisting of multiple physical servers, or a distributed file system. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. In some embodiments, the server and the terminal are directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment. Optionally, the number of servers can be greater or lesser, which is not limited in this embodiment. Of course, the server can also include other functional servers to provide more comprehensive and diverse services. Among them, the server undertakes the main computing work and the terminal undertakes the secondary computing work; or, the server undertakes the secondary computing work and the terminal undertakes the main computing work; or, the server or the terminal can undertake the computing work independently, which is not limited in the embodiments of the present disclosure.

[0049] Figure 2 is a flow chart of an audio generation method according to an exemplary embodiment. Figure 2 As shown, the method is executed by an electronic device, and the method includes the following steps.

[0050] In step S201 , the electronic device obtains a video for which audio is to be generated.

[0051] In the embodiments of the present disclosure, the video for which audio is to be generated can be any video, and the embodiments of the present disclosure do not limit this. A target application is installed in the electronic device, and the target application is used to generate audio. The target application can be any application capable of generating audio, and the embodiments of the present disclosure do not limit the type of the target application.

[0052] In some embodiments, the electronic device displays relevant information about the target application, including at least one of the name and icon of the target application. In response to a triggering operation on the relevant information about the target application, a video upload interface is displayed, wherein a video upload control is displayed on the video upload interface. In response to a triggering operation on the video upload control, a video selection interface is displayed, wherein at least one selectable video is displayed on the video selection interface. In response to a triggering operation on any of the at least one selectable video, the triggered selectable video is used as the video for which audio is to be generated.

[0053] Among them, at least one optional video displayed in the video selection interface is a video stored in the electronic device. The triggering operation for the relevant information of the target application can be a click operation for the relevant information of the target application, or a double-click operation for the relevant information of the target application, or other operations for the relevant information of the target application, which is not limited in the embodiments of the present disclosure. The process of triggering operations for other content is similar to the process of triggering operations for the relevant information of the target application, which is not limited in the embodiments of the present disclosure.

[0054] In step S202, the electronic device determines the video features, video duration features, and audio-visual synchronization features corresponding to the video. The video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance of the video surface and sound.

[0055] In an embodiment of the present disclosure, the process of determining the video features corresponding to a video includes: performing frame processing on the video to obtain multiple pictures included in the video; determining the image features of each picture, and the image features of each picture are used to characterize the corresponding picture; and determining the video features corresponding to the video based on the image features of each picture.

[0056] In this embodiment, by decomposing the video into multiple pictures, extracting picture features and integrating them to obtain video features, it is possible to reduce processing complexity while integrating information in spatial and temporal dimensions, thereby improving the comprehensiveness, robustness and adaptability of video feature expression to different videos, thereby making the generated video features more compatible with the video and more accurate.

[0057] In the embodiment of the present disclosure, the process of determining the image features of each picture includes: for any one of the pictures, any one of the pictures is processed by a picture feature acquisition model to obtain the image features of any one of the pictures, and the picture feature acquisition model is used to obtain the picture features.

[0058] In this embodiment, by acquiring image features through an image feature acquisition model, image information can be efficiently and accurately converted into a structured representation that can be understood by electronic devices, so that the accuracy of the generated image features is higher. Since image features are used to generate video features, the accuracy of the generated video features can also be improved.

[0059] In some embodiments, the process of determining the video duration features corresponding to a video includes: obtaining the number of frames and frame rate of the video; determining the duration of the video based on the number of frames and frame rate of the video; and determining the video duration features corresponding to the video based on the duration of the video.

[0060] In this embodiment, by obtaining the number of frames and frame rate of the video to determine the duration of the video and extracting the duration features, the time dimension information of the video can be converted into structured data in a quantitative manner, so that the generated video duration features are more accurate.

[0061] In some embodiments, the process of determining the duration of a video based on the number of frames and frame rate of the video includes: determining the quotient between the number of frames and the frame rate of the video; when the quotient is an integer, determining the quotient as the duration of the video; when the quotient is a non-integer, rounding up the quotient to obtain the duration of the video.

[0062] In this embodiment, by calculating the quotient between the number of frames and the frame rate and determining the duration of the video based on the quotient, the duration of the video can be determined in a standardized calculation method, ensuring the standardization and adaptability of the duration calculation.

[0063] In some embodiments, the process of determining the video duration features corresponding to a video based on the duration of the video includes: mapping the duration of the video into an array; processing the array through a duration learning model to obtain the video duration features corresponding to the video, and the duration learning model is used to determine the video duration features.

[0064] In this embodiment, the duration of the video is mapped into an array and processed through a duration learning model, which can convert the duration information into structured features of the model science system, thereby adaptively extracting deep semantic representations containing temporal regularities to adapt to subsequent analysis tasks.

[0065] In some embodiments, the duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer, and a third convolutional layer. The process of processing the array through the duration learning model to obtain the video duration features corresponding to the video includes: inputting the array into the fully connected layer to obtain the output result of the fully connected layer; inputting the output result of the fully connected layer into the first convolutional layer and the second convolutional layer respectively to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; mapping the output result of the first convolutional layer through an activation function to obtain the attention weight; dot-multiplying the attention weight and the output result of the second convolutional layer to obtain the dot product result; inputting the dot product result into the third convolutional layer to obtain the video duration features of the video.

[0066] In this embodiment, through hierarchical processing of fully connected layers and multiple convolutional layers, and with the help of attention weights generated by activation functions, the output results of the convolutional layers are weighted, so as to efficiently capture the key features related to the video duration in the array, and suppress the information in the array that is not related to the video duration, thereby achieving accurate extraction of video duration features.

[0067] In some embodiments, the process of determining the audio-visual synchronization feature corresponding to the video includes: processing the video through an audio-visual synchronization encoder to obtain the audio-visual synchronization feature corresponding to the video, and the audio-visual synchronization encoder is used to obtain the audio-visual synchronization feature.

[0068] In this embodiment, the audio and video images of the video are jointly encoded by an audio and video synchronization encoder, which can effectively capture the temporal consistency and semantic relevance of the two, thereby accurately extracting audio and video synchronization features that reflect the collaborative relationship between audio and video.

[0069] In step S203, the electronic device generates audio corresponding to the video according to the video feature, the video duration feature and the audio-video synchronization feature, where the duration of the audio is the same as the duration of the video.

[0070] The disclosed embodiments provide an audio generation method, which generates audio corresponding to a video based on the video without manual operation by the user, thereby simplifying the audio generation method and making the audio generation more efficient. Moreover, there is no need to rely on the user's subjective consciousness, so that the matching degree between the generated audio and the video is higher, and the audio generation efficiency is better. In addition, since not only the video features but also the video duration features and the audio-visual synchronization features are obtained, when generating audio, not only the video features of the video itself but also the video duration features and the audio-visual synchronization features are taken into account, so that the duration of the generated audio is the same as the duration of the video, and the generated audio is synchronized with the video audio and video, further improving the matching degree between the generated audio and the video and the audio generation efficiency.

[0071] In some embodiments, the method also includes: obtaining audio description text, which is a description text of the audio corresponding to the video; determining text features corresponding to the audio description text; and generating the audio corresponding to the video based on video features, video duration features, and audio-visual synchronization features. The process includes: generating the audio corresponding to the video based on video features, video duration features, audio-visual synchronization features, and text features.

[0072] In this embodiment, audio description text can also be obtained. When generating audio, text features corresponding to the audio description text are also taken into consideration, so that the generated audio is not only audio that matches the video, but also audio that meets user requirements.

[0073] In some embodiments, the process of determining the text features corresponding to the audio description text includes: obtaining a first feature corresponding to the audio description text, the first feature is used to characterize the audio description text; performing keyword extraction on the audio description text to obtain at least one keyword; obtaining a second feature corresponding to each keyword, the second feature corresponding to each keyword is used to characterize the corresponding keyword; and determining the text features corresponding to the audio description text based on the first feature and the second feature corresponding to each keyword.

[0074] In this embodiment, by combining the overall semantic features of the audio description text (i.e., the first features) with the key semantic features extracted from the audio description text (i.e., the second features), the audio-related information in the audio description text can be captured more comprehensively and with emphasis, thereby generating more accurate and effective text features.

[0075] In some embodiments, the process of obtaining the first feature corresponding to the audio description text includes: processing the audio description text through a first text feature acquisition model to obtain the first feature corresponding to the audio description text, and the first text feature acquisition model is used to obtain text features.

[0076] In this embodiment, the audio description text is processed by the first text feature acquisition model, which can effectively extract the first feature that can accurately characterize the overall semantic information of the audio description text, providing a basis for subsequent acquisition of text features.

[0077] In some embodiments, the process of obtaining the second feature corresponding to each keyword includes: for any keyword among the keywords, processing any keyword through a second text feature acquisition model to obtain the second feature corresponding to any keyword, the second text feature acquisition model is used to obtain text features, and the number of tokens supported by the second text feature acquisition model is less than the number of tokens supported by the first text feature acquisition model.

[0078] In this embodiment, the keywords are processed by the second text feature acquisition model, which can effectively extract the second features that can accurately characterize the keywords, providing a basis for subsequent acquisition of text features.

[0079] In some embodiments, the process of generating the audio corresponding to the video according to the video features, video duration features, audio-visual synchronization features and text features includes: fusing the video features, text features and audio noise features to obtain fused features, the audio noise features and the audio features of the audio corresponding to the video having the same dimension; splitting the fused features to obtain split video features, split text features and split audio noise features, the split video features, split text features and split audio noise features all being fused with video features, text features and audio noise features; generating alignment features according to the audio-visual synchronization features, the dimension of the alignment features being the same as the dimension of the audio noise features; generating global features according to the text features, video features, video duration features and alignment features, the global features being features after considering the text features, video features, video duration features and alignment features; generating the audio corresponding to the video according to the alignment features, global features, split video features, split text features and split audio noise features.

[0080] In this embodiment, the video features, text features and audio noise features are first fused and split, which can enable the full interaction of each modal feature, break the modal barriers, enhance the correlation and complementarity between features, generate alignment features based on the audio and video synchronization features, ensure the temporal and semantic coordination of audio and video, generate global features based on multiple features, comprehensively integrate multi-dimensional information, and fully present the core content of the video and text and the audio noise conditions. Finally, various features are combined to generate audio corresponding to the video. The rich features obtained in the previous processing process can be fully utilized, and comprehensive considerations from multiple aspects can be taken to accurately generate audio that fits the video content, is synchronized with the video picture, and conforms to the semantics of the audio description text, effectively improving the accuracy and fit of the generated audio, and optimizing the quality of the generated audio.

[0081] In some embodiments, the process of generating the alignment feature according to the audio-video synchronization feature includes: up-sampling the audio-video synchronization feature to the same dimension as the audio noise feature to obtain the alignment feature.

[0082] In this embodiment, due to the different frame rates between the audio-visual synchronization features and the audio noise features, in order to align the timing relationship between the visual content and the audio, the audio-visual synchronization features are upsampled to the same dimension as the audio noise features to obtain alignment features. This can achieve scale unification between features of different dimensions, providing a basis for subsequent audio generation.

[0083] In some embodiments, the process of generating global features based on text features, video features, video duration features and alignment features includes: pooling the text features, video features and video duration features to obtain pooled features, and the pooled features are used to characterize the consistency between the audio description text and the video; generating global features based on the pooled features and the alignment features.

[0084] In this embodiment, by pooling text features, video features and video duration features to extract key information of multimodal consistency, and then combining with alignment features, a global feature can be generated that integrates the consistency association between modalities and the adaptability of the audio and video synchronization dimension, so that the generated global feature is more accurate and provides comprehensive feature support for the subsequent generation of audio.

[0085] In some embodiments, the process of generating the global feature according to the pooled feature and the aligned feature includes: adding the pooled feature and the aligned feature to obtain the global feature.

[0086] In this embodiment, the pooled features and the aligned features are added together to generate global features, which can incorporate audio-visual synchronization dimension information while retaining the multimodal consistency representation, achieving complementary fusion at the feature level to enhance the comprehensiveness of the global representation.

[0087] In some embodiments, the process of generating audio corresponding to a video based on alignment features, global features, split video features, split text features, and split audio noise features includes: determining audio features of the audio corresponding to the video based on alignment features, global features, split video features, split text features, and split audio noise features; and generating audio corresponding to the video based on the audio features.

[0088] In this embodiment, audio features are determined by comprehensively utilizing alignment features, global features, and split multimodal features, which can fully integrate the synchronization relationship between video and audio, the overall correlation of multimodal content, and the local detail information of each modality. The alignment features ensure the temporal and semantic coordination of audio and video; the global features integrate the consistency of multimodality and the adaptability of audio and video synchronization, and the split features retain the original information and interaction details of different modalities. On this basis, audio features are generated, which can accurately capture the key audio elements required for the video content, ensure that the generated audio is highly consistent with the video in all aspects, so that the generated audio can not only accurately match the video, but also be combined with the audio description text to make the generated audio more effective.

[0089] In some embodiments, the process of determining the audio features of the audio corresponding to the video based on the alignment features, global features, split video features, split text features and split audio noise features includes: processing the alignment features, global features, split video features, split text features and split audio noise features through an audio feature generation model to obtain the audio features of the audio corresponding to the video, and the audio feature generation model is used to generate audio features.

[0090] In this embodiment, the audio feature generation model is used to process multi-dimensional features to generate audio features, which can efficiently integrate the synchronization of sound and picture, multi-modal global correlation and detailed information of each modality to accurately match the video content to generate suitable audio features, so that the generated audio features are more accurate, and the effect of the subsequently generated audio is better.

[0091] The above is based on Figure 2 The embodiment of the present invention briefly introduces the process of audio generation. Figure 3 The embodiment of the present invention further introduces the process of audio generation. Figure 3 , Figure 3 The flowchart of an audio generation method according to an exemplary embodiment is shown. The method includes the following steps.

[0092] In step S301, the electronic device obtains a video for which audio is to be generated.

[0093] In the embodiments of the present disclosure, the video for which audio is to be generated can be any video, and the embodiments of the present disclosure do not limit this. A target application is installed in the electronic device, and the target application is used to generate audio. The target application can be any application capable of generating audio, and the embodiments of the present disclosure do not limit the type of the target application. For example, the target application is a video editing application.

[0094] In some embodiments, the electronic device displays relevant information about a target application, including at least one of a name and an icon of the target application. In response to a triggering operation on the relevant information about the target application, a video upload interface is displayed, wherein a video upload control is displayed on the video upload interface, and the video upload control is used to upload a video. In response to a triggering operation on the video upload control, a video selection interface is displayed, wherein at least one selectable video is displayed on the video selection interface. In response to a triggering operation on any of the at least one selectable video, the triggered selectable video is used as the video for which audio is to be generated.

[0095] Among them, at least one optional video displayed in the video selection interface is a video stored in the electronic device. The triggering operation for the relevant information of the target application can be a click operation for the relevant information of the target application, or a double-click operation for the relevant information of the target application, or other operations for the relevant information of the target application, which is not limited in the embodiments of the present disclosure. The process of triggering operations for other content is similar to the process of triggering operations for the relevant information of the target application, which is not limited in the embodiments of the present disclosure.

[0096] In step S302, the electronic device determines the video features, video duration features, and audio-visual synchronization features corresponding to the video. The video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance of the video's picture and sound.

[0097] In some embodiments, after acquiring the video in the above steps, the video features corresponding to the video may be determined first, followed by the video duration features corresponding to the video, and finally the audio-visual synchronization features corresponding to the video may be determined; the video duration features corresponding to the video may be determined first, followed by the video features corresponding to the video, and finally the audio-visual synchronization features corresponding to the video may be determined; or the audio-visual synchronization features corresponding to the video may be determined first, followed by the video features corresponding to the video, and finally the video duration features corresponding to the video. The disclosed embodiments do not limit the order in which the video features corresponding to the video, the video duration features corresponding to the video, and the audio-visual synchronization features corresponding to the video are determined.

[0098] In some embodiments, the process of an electronic device determining the video features corresponding to a video includes: performing frame extraction on the video to obtain multiple pictures included in the video; determining the picture features of each picture, and the picture features of each picture are used to represent the corresponding picture; and determining the video features corresponding to the video based on the picture features of each picture.

[0099] In this embodiment, by decomposing the video into multiple pictures, extracting picture features and integrating them to obtain video features, it is possible to reduce processing complexity while integrating information in spatial and temporal dimensions, thereby improving the comprehensiveness, robustness and adaptability of video feature expression to different videos, thereby making the generated video features more compatible with the video and more accurate.

[0100] The process of extracting frames from a video to obtain multiple images included in the video includes: extracting a target number of images from the images included in each second of the video to obtain the multiple images included in the video. The target number is set based on experience or flexibly adjusted according to the implementation environment, and the embodiments of the present disclosure do not limit this. For example, if the target number is 8 and the video duration is 10 seconds, then 8 frames of images are extracted per second, for a total of 80 frames of images. These 80 frames of images are the multiple images included in the video.

[0101] In some embodiments, after the video is subjected to frame extraction processing to obtain multiple pictures included in the video, the process of determining the image features of each picture includes: for any one of the pictures, any one of the pictures is processed through a picture feature acquisition model to obtain the image features of any one of the pictures, and the picture feature acquisition model is used to obtain the picture features.

[0102] Among them, the image feature acquisition model can be a visual encoder (encoder) of Metaclip (Metadata-Curated Language-Image Pre-training, a language-image pre-training model) with a parameter scale of 2.5B (billion) trained using a contrastive learning algorithm, or it can be a visual VAE (Variational Autoencoder) model, or it can be a visual representation model, which is not limited in the embodiments of the present disclosure.

[0103] Optionally, any picture is input into the picture feature acquisition model, and the output result of the picture feature acquisition model is the picture feature of any picture.

[0104] In this embodiment, by acquiring image features through an image feature acquisition model, image information can be efficiently and accurately converted into a structured representation that can be understood by electronic devices, so that the accuracy of the generated image features is higher. Since image features are used to generate video features, the accuracy of the generated video features can also be improved.

[0105] In some embodiments, the process of determining the video features corresponding to the video based on the image features of each image includes: splicing the image features of each image to obtain the video features corresponding to the video.

[0106] In some embodiments, the process of determining the video duration features corresponding to a video includes: obtaining the number of frames and frame rate of the video; determining the duration of the video based on the number of frames and frame rate of the video; and determining the video duration features corresponding to the video based on the duration of the video.

[0107] In this embodiment, by obtaining the number of frames and frame rate of the video to determine the duration of the video and extracting the duration features, the time dimension information of the video can be converted into structured data in a quantitative manner, so that the generated video duration features are more accurate.

[0108] In some embodiments, the process of obtaining the number of frames and the frame rate of the video includes obtaining the number of frames and the frame rate of the video through the OpenCV tool (OpenSource Computer Vision Library, an open source computer vision and machine learning software library).

[0109] In some embodiments, the process of determining the duration of a video based on the number of frames and frame rate of the video includes: determining the quotient between the number of frames and the frame rate of the video; when the quotient is an integer, determining the quotient as the duration of the video; when the quotient is a non-integer, rounding up the quotient to obtain the duration of the video.

[0110] In this embodiment, by calculating the quotient between the number of frames and the frame rate and determining the duration of the video based on the quotient, the duration of the video can be determined in a standardized calculation method, ensuring the standardization and adaptability of the duration calculation.

[0111] For example, the number of frames of a video is 3600 and the frame rate is 25 frames per second (fps). The quotient between the number of frames and the frame rate of the video is 3600 / 25=144. Since the quotient between the number of frames and the frame rate of the video is an integer, the quotient between the number of frames and the frame rate of the video is used as the duration of the video, that is, the duration of the video is 144 seconds.

[0112] For another example, if the number of frames in a video is 3600 and the frame rate is 35 frames per second (fps), then the quotient between the number of frames and the frame rate is 3600 / 35≈102.9. Since the quotient between the number of frames and the frame rate is a non-integer, the quotient between the number of frames and the frame rate is rounded up to 103, which means that the length of the video is 103 seconds.

[0113] In some embodiments, the process of determining the video duration features corresponding to a video based on the duration of the video includes: mapping the duration of the video into an array; processing the array through a duration learning model to obtain the video duration features corresponding to the video, and the duration learning model is used to determine the video duration features.

[0114] In this embodiment, the duration of the video is mapped into an array and processed through a duration learning model, which can convert the duration information into structured features of the model science system, thereby adaptively extracting deep semantic representations containing temporal regularities to adapt to subsequent analysis tasks.

[0115] In some embodiments, the process of mapping the duration of the video into an array includes: mapping the duration of the video into an array of [512, 768].

[0116] In some embodiments, the duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer, and a third convolutional layer. The first convolutional layer, the second convolutional layer, and the third convolutional layer are all one-dimensional convolutional layers. The process of processing an array through the duration learning model to obtain the video duration features corresponding to a video includes: inputting the array into the fully connected layer to obtain the output result of the fully connected layer; inputting the output result of the fully connected layer into the first convolutional layer and the second convolutional layer respectively to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; mapping the output result of the first convolutional layer through an activation function to obtain an attention weight; dot-multiplying the attention weight and the output result of the second convolutional layer to obtain a dot-multiplication result; and inputting the dot-multiplication result into the third convolutional layer to obtain the video duration features of the video.

[0117] The activation function may be a SiLu (Sigmoid Linear Unit, an activation function commonly used in deep learning) activation function.

[0118] In this embodiment, through hierarchical processing of fully connected layers and multiple convolutional layers, and with the help of attention weights generated by activation functions, the output results of the convolutional layers are weighted, so as to efficiently capture the key features related to the video duration in the array, and suppress the information in the array that is not related to the video duration, thereby achieving accurate extraction of video duration features.

[0119] like Figure 4 Schematic diagram of a process for determining the video duration feature of a video provided by an embodiment of the present application. An array is input into a fully connected layer 401 to obtain the output of the fully connected layer; the output of the fully connected layer is input into a first convolutional layer 402 and a second convolutional layer 403 to obtain the output of the first convolutional layer and the output of the second convolutional layer; the output of the first convolutional layer is mapped through an activation function to obtain an attention weight; the output of the second convolutional layer is multiplied by the attention weight to obtain a dot product result; and the dot product result is input into a third convolutional layer 404 to obtain the video duration feature of the video.

[0120] In some embodiments, the process of determining the audio-visual synchronization feature corresponding to the video includes: processing the video through an audio-visual synchronization encoder to obtain the audio-visual synchronization feature corresponding to the video, and the audio-visual synchronization encoder is used to obtain the audio-visual synchronization feature.

[0121] The audio-visual synchronization encoder decouples feature extraction from synchronization modeling by performing multimodal segment-level comparative pre-training on videos, achieving efficient audio-visual synchronization under sparse synchronization cues.

[0122] In some embodiments, the video is input into an audio-visual synchronization encoder to obtain audio-visual synchronization features corresponding to the video.

[0123] In this embodiment, the audio and video images of the video are jointly encoded by an audio and video synchronization encoder, which can effectively capture the temporal consistency and semantic relevance of the two, thereby accurately extracting audio and video synchronization features that reflect the collaborative relationship between audio and video.

[0124] In step S303, the electronic device obtains an audio description text, where the audio description text is a description text of the audio corresponding to the video.

[0125] The disclosed embodiments do not limit the method for obtaining the audio description text. Optionally, the electronic device may store multiple optional texts, and any one of the multiple optional texts may be used as the audio description text. Optionally, a text input box may be displayed, and in response to a user input operation in the text input box, the input content may be used as the audio description text.

[0126] For example, the audio description text is the sound of a racket hitting a ball.

[0127] The disclosed embodiments do not limit the order of obtaining the video of the audio to be generated and obtaining the audio description text. Optionally, the video of the audio to be generated is obtained first, and then the audio description text is obtained, or the audio description text is obtained first, and then the video of the audio to be generated is obtained.

[0128] In step S304, the electronic device determines text features corresponding to the audio description text.

[0129] In some embodiments, after the audio description text is obtained in the above steps, the process of determining the text features corresponding to the audio description text includes: obtaining the first feature corresponding to the audio description text, the first feature is used to characterize the audio description text; performing keyword extraction on the audio description text to obtain at least one keyword; obtaining the second feature corresponding to each keyword, the second feature corresponding to each keyword is used to characterize the corresponding keyword; determining the text features corresponding to the audio description text based on the first feature and the second feature corresponding to the keyword.

[0130] In this embodiment, by combining the overall semantic features of the audio description text (i.e., the first features) with the key semantic features extracted from the audio description text (i.e., the second features), the audio-related information in the audio description text can be captured more comprehensively and with emphasis, thereby generating more accurate and effective text features.

[0131] In some embodiments, the process of obtaining the first feature corresponding to the audio description text includes: processing the audio description text through a first text feature acquisition model to obtain the first feature corresponding to the audio description text, and the first text feature acquisition model is used to obtain text features.

[0132] The audio description text is input into the first text feature acquisition model to obtain the first feature corresponding to the audio description text.

[0133] In some embodiments, the first text feature acquisition model can be a T5 (Text-To-Text Transfer Transformer) large text model or a GLM (General Language Model) model, which is not limited in this disclosure. The T5 large text model supports up to 512 text tokens and can be used for understanding long texts.

[0134] In this embodiment, the audio description text is processed by the first text feature acquisition model, which can effectively extract the first feature that can accurately characterize the overall semantic information of the audio description text, providing a basis for subsequent acquisition of text features.

[0135] In some embodiments, the process of extracting keywords from the audio description text to obtain at least one keyword includes: performing word segmentation on the audio description text to obtain multiple words, and using the words in the keyword library among the multiple words as keywords.

[0136] In some embodiments, the process of obtaining the second feature corresponding to each keyword includes: for any keyword among the keywords, processing any keyword through a second text feature acquisition model to obtain the second feature corresponding to any keyword, the second text feature acquisition model is used to obtain text features, and the number of tokens supported by the second text feature acquisition model is less than the number of tokens supported by the first text feature acquisition model.

[0137] Among them, any keyword is input into the second text feature acquisition model to obtain the second feature corresponding to any keyword.

[0138] In some embodiments, the second text feature acquisition model can be MetaClip's text encoder. MetaClip's text encoder is trained using short text and keywords using contrastive learning, which can improve keyword retrieval and comprehension capabilities. MetaClip's text encoder supports 77 tokens and does not support text comprehension of longer lengths.

[0139] In this embodiment, the keywords are processed by the second text feature acquisition model, which can effectively extract the second features that can accurately characterize the keywords, providing a basis for subsequent acquisition of text features.

[0140] In some embodiments, the process of determining the text feature corresponding to the audio description text based on the first feature and the second feature corresponding to each keyword includes: splicing the first feature and the second feature corresponding to each keyword to obtain the text feature corresponding to the audio description text.

[0141] In step S305, the electronic device generates audio corresponding to the video according to the video features, video duration features, audio-visual synchronization features and text features, and the duration of the audio is the same as the duration of the video.

[0142] In some embodiments, the process of generating the audio corresponding to the video according to the video features, video duration features, audio-visual synchronization features and text features includes: fusing the video features, text features and audio noise features to obtain fused features, the audio noise features and the audio features of the audio corresponding to the video having the same dimension; splitting the fused features to obtain split video features, split text features and split audio noise features, the split video features, split text features and split audio noise features all being fused with video features, text features and audio noise features; generating alignment features according to the audio-visual synchronization features, the dimension of the alignment features being the same as the dimension of the audio noise features; generating global features according to the text features, video features, video duration features and alignment features, the global features being features after considering the text features, video features, video duration features and alignment features; generating the audio corresponding to the video according to the alignment features, global features, split video features, split text features and split audio noise features.

[0143] In this embodiment, the video features, text features and audio noise features are first fused and split, which can enable the full interaction of each modal feature, break the modal barriers, enhance the correlation and complementarity between features, generate alignment features based on the audio and video synchronization features, ensure the temporal and semantic coordination of audio and video, generate global features based on multiple features, comprehensively integrate multi-dimensional information, and fully present the core content of the video and text and the audio noise conditions. Finally, various features are combined to generate audio corresponding to the video. The rich features obtained in the previous processing process can be fully utilized, and comprehensive considerations from multiple aspects can be taken to accurately generate audio that fits the video content, is synchronized with the video picture, and conforms to the semantics of the audio description text, effectively improving the accuracy and fit of the generated audio, and optimizing the quality of the generated audio.

[0144] In some embodiments, fusing video features, text features, and audio noise features to obtain a fused feature includes fusing the video features, text features, and audio noise features through a cross-attention mechanism to obtain the fused feature. The video features, text features, and audio noise features are the query, key, and value in the cross-attention mechanism, respectively. The fused feature is the feature that is the result of fusing the video features, text features, and audio noise features.

[0145] The dimensions of the audio noise features are set based on experience, or flexibly adjusted according to the implementation environment, which is not limited in the embodiments of the present disclosure.

[0146] In some embodiments, the process of splitting the fused features to obtain split video features, split text features, and split audio noise features includes: splitting the fused features in the order in which the video features, text features, and audio noise features are fused to obtain split video features, split text features, and split audio noise features. The dimension of the split video features is the same as the dimension of the video features, the dimension of the split text features is the same as the dimension of the text features, and the dimension of the split audio noise features is the same as the dimension of the audio noise features.

[0147] In some embodiments, the process of generating the alignment feature according to the audio-video synchronization feature includes: up-sampling the audio-video synchronization feature to the same dimension as the audio noise feature to obtain the alignment feature.

[0148] In this embodiment, due to the different frame rates between the audio-visual synchronization features and the audio noise features, in order to align the timing relationship between the visual content and the audio, the audio-visual synchronization features are upsampled to the same dimension as the audio noise features to obtain alignment features. This can achieve scale unification between features of different dimensions, providing a basis for subsequent audio generation.

[0149] In some embodiments, the process of generating global features based on text features, video features, video duration features and alignment features includes: pooling the text features, video features and video duration features to obtain pooled features, and the pooled features are used to characterize the consistency between the audio description text and the video; generating global features based on the pooled features and the alignment features.

[0150] In this implementation, by pooling text features, video features, and video duration features to extract key information of multimodal consistency, and then combining them with alignment features, a global feature can be generated that integrates the consistency association between modalities and the adaptability of the audio-visual synchronization dimension. This makes the generated global feature more accurate and provides comprehensive feature support for the subsequent audio generation.

[0151] In some embodiments, the process of generating the global feature according to the pooled feature and the aligned feature includes: adding the pooled feature and the aligned feature to obtain the global feature.

[0152] In this embodiment, the pooled features and the aligned features are added together to generate global features, which can incorporate audio-visual synchronization dimension information while retaining the multimodal consistency representation, achieving complementary fusion at the feature level to enhance the comprehensiveness of the global representation.

[0153] In some embodiments, the process of generating audio corresponding to a video based on alignment features, global features, split video features, split text features, and split audio noise features includes: determining audio features of the audio corresponding to the video based on alignment features, global features, split audio features, split text features, and split audio noise features; and generating audio corresponding to the video based on the audio features.

[0154] In this embodiment, audio features are determined by comprehensively utilizing alignment features, global features, and split multimodal features, which can fully integrate the synchronization relationship between video and audio, the overall correlation of multimodal content, and the local detail information of each modality. The alignment features ensure the temporal and semantic coordination of audio and video; the global features integrate the consistency of multimodality and the adaptability of audio and video synchronization, and the split features retain the original information and interaction details of different modalities. On this basis, audio features are generated, which can accurately capture the key audio elements required for the video content, ensure that the generated audio is highly consistent with the video in all aspects, so that the generated audio can not only accurately match the video, but also be combined with the audio description text to make the generated audio more effective.

[0155] In some embodiments, the process of determining the audio features of the audio corresponding to the video based on the alignment features, global features, split video features, split text features and split audio noise features includes: processing the alignment features, global features, split video features, split text features and split audio noise features through an audio feature generation model to obtain the audio features of the audio corresponding to the video, and the audio feature generation model is used to generate audio features.

[0156] In some embodiments, the aligned features, global features, split video features, split text features, and split audio noise features are input into an audio feature generation model through AdaLN (adaptive layer normalization) and a cross-attention mechanism to obtain audio features of the audio corresponding to the video.

[0157] Among them, the audio feature generation model is a denoising-based multimodal DIT (Denoising Diffusion Implicit Models) generation model.

[0158] In this embodiment, the audio feature generation model is used to process multi-dimensional features to generate audio features, which can efficiently integrate the synchronization of sound and picture, multi-modal global correlation and detailed information of each modality to accurately match the video content to generate suitable audio features, so that the generated audio features are more accurate, and the effect of the subsequently generated audio is better.

[0159] In some embodiments, the process of generating audio corresponding to a video based on audio features includes: processing the audio features through an audio encoder to obtain a Mel-spectrogram corresponding to the audio features; and processing the Mel-spectrogram corresponding to the audio features through a voice coder to obtain the audio corresponding to the video.

[0160] The vocoder is any vocoder that supports the generation of music sound effects, and the embodiment of the present disclosure does not limit the type of the vocoder.

[0161] In some embodiments, the audio feature is input into an audio encoder to obtain a mel-spectrogram corresponding to the audio feature; the mel-spectrogram corresponding to the audio feature is input into a vocoder to obtain audio corresponding to the video.

[0162] The audio generation method provided in the embodiments of the present disclosure can be used to dub AIGC (Artificial Intelligence Generated Content) or UGC (User Generated Content) videos, and simultaneously process multi-source superposition scenes (such as traffic sounds, human voices, construction sounds, etc. on city streets). By decoupling multimodal features, corresponding sound effects are generated separately, replacing the tedious process of manual superposition one by one. In addition, it supports dual-conditional generation of "audio description text + video", such as controlling the footsteps style of the characters in the video through the audio description text "suspense movie footsteps", providing more creative possibilities for film and television dubbing, while adapting to emerging scenarios such as real-time game sound effect generation and virtual anchor dubbing.

[0163] In some embodiments, the audio generation method provided by the embodiments of the present disclosure is a unified end-to-end multimodal audio generation technology, with video or video and text as input and audio as output. That is to say, the embodiments of the present disclosure can implement audio generation tasks such as video-audio, text combined with video-audio, and moreover, the embodiments of the present disclosure can also be expanded to support the generation of character dialogue voices. The embodiments of the present disclosure use an audio generation feature model to generate audio features, effectively realizing the fusion of information between various modalities and improving the content relevance of audio generation. At the same time, the card point alignment capability of video and audio is also improved. In addition, by adding a duration learning model, it effectively supports the generation of audio of various durations, increasing the flexibility of video dubbing.

[0164] The disclosed embodiments provide an audio generation method, which generates audio corresponding to a video based on the video without manual operation by the user, thereby simplifying the audio generation method and making the audio generation more efficient. Moreover, there is no need to rely on the user's subjective consciousness, so that the matching degree between the generated audio and the video is higher, and the audio generation efficiency is better. In addition, since not only the video features but also the video duration features and the audio-visual synchronization features are obtained, when generating audio, not only the video features of the video itself but also the video duration features and the audio-visual synchronization features are taken into account, so that the duration of the generated audio is the same as the duration of the video, and the generated audio is synchronized with the video audio and video, further improving the matching degree between the generated audio and the video and the audio generation efficiency.

[0165] In addition, audio description text can also be obtained. When generating audio corresponding to the video, not only the video features, video duration features, and audio-visual synchronization features are taken into consideration, but also the text features corresponding to the audio description text are taken into consideration, so that the generated audio is not only audio that matches the video, but also audio that matches the audio description text, making the generated audio more accurate.

[0166] Figure 5 FIG. 1 is a flow chart of an audio generation method according to an exemplary embodiment. Figure 5 As shown, the method includes the following steps.

[0167] Step 501: The electronic device obtains the video and audio description text of the audio to be generated.

[0168] The audio description text is the description text of the audio corresponding to the video.

[0169] In some embodiments, the process of obtaining the video to generate audio has been described in the above step S301, and the process of obtaining the audio description text has been described in the above step S303, and the embodiments of the present disclosure will not be repeated here.

[0170] Step 502: The electronic device determines the video features, video duration features, and audio-video synchronization features of the video.

[0171] Among them, video features are used to characterize videos, video duration features are used to characterize the duration of videos, and audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance of the video's picture and sound.

[0172] In some embodiments, the process of determining the video features, video duration features, and audio-visual synchronization features of the video has been described in the above step S302, and will not be repeated here in the embodiments of the present disclosure.

[0173] Step 503: The electronic device determines text features corresponding to the audio description text.

[0174] In some embodiments, the process of determining the text features corresponding to the audio description text has been described in the above step S304 and will not be repeated here in the embodiment of the present disclosure.

[0175] Step 504: The electronic device fuses the video features, text features, and audio noise features to obtain fused features.

[0176] The audio noise feature has the same dimension as the audio feature of the audio corresponding to the video. The fusion feature is a feature that combines video features, text features, and audio noise features.

[0177] In some embodiments, the process of fusing video features, text features, and audio noise features to obtain fused features has been described in the above step S305 and will not be repeated here in the embodiments of the present disclosure.

[0178] Step 505: The electronic device splits the fused features to obtain split video features, split text features, and split audio noise features.

[0179] Among them, the split video features, the split text features and the split audio noise features all integrate video features, text features and audio noise features.

[0180] In some embodiments, the process of splitting the fused features to obtain split video features, split text features, and split audio noise features has been described in the above step S305 and will not be repeated here in the embodiments of the present disclosure.

[0181] Step 506: The electronic device generates an alignment feature based on the audio and video synchronization feature.

[0182] The dimension of the alignment feature is the same as that of the audio noise feature.

[0183] In some embodiments, the process of generating alignment features based on the audio-visual synchronization features has been described in the above step S305 and will not be repeated here in the embodiments of the present disclosure.

[0184] Step 507: The electronic device generates a global feature based on the text feature, the video feature, the video duration feature, and the alignment feature.

[0185] Among them, the global features are the features after considering text features, video features, video duration features and alignment features.

[0186] In some embodiments, the process of generating global features based on text features, video features, video duration features and alignment features has been described in the above step S305 and will not be repeated here in the embodiments of the present disclosure.

[0187] Step 508: The electronic device processes the alignment features, global features, split video features, split text features, and split audio noise features through an audio feature generation model to obtain audio features of the audio corresponding to the video.

[0188] Among them, the audio feature generation model is used to generate audio features.

[0189] In some embodiments, the process of processing the alignment features, global features, split video features, split text features and split audio noise features through the audio feature generation model to obtain the audio features of the audio corresponding to the video has been described in the above step S305, and the embodiments of the present disclosure are not limited to this.

[0190] Step 509: The electronic device processes the audio feature through an audio encoder to obtain a Mel-spectrogram corresponding to the audio feature.

[0191] In some embodiments, the process of processing the audio features by the audio encoder to obtain the mel-spectrogram corresponding to the audio features has been described in the above step S305, and the embodiments of the present disclosure are not limited to this.

[0192] Step 510: The electronic device processes the mel-spectrogram corresponding to the audio feature through a vocoder to obtain audio corresponding to the video.

[0193] In some embodiments, the process of processing the mel-spectrogram corresponding to the audio feature by a vocoder to obtain the audio corresponding to the video has been described in the above step S305, and the embodiments of the present disclosure are not limited to this.

[0194] Figure 6 FIG. 1 is a flow chart of an audio generation method according to an exemplary embodiment. Figure 6 As shown, the method includes: extracting frames from a video to generate audio to obtain multiple images, processing each image using an image feature acquisition model to obtain image features for each image, and determining video features based on the image features of each image. Processing the video using a duration learning model to obtain video duration features. Processing the video using an audio-video synchronization encoder to obtain audio-video synchronization features.

[0195] The audio description text is processed using a first text feature acquisition model to obtain a first feature corresponding to the audio description text. Keyword extraction is performed on the audio description text to obtain multiple keywords. Each keyword is then processed using a second text feature acquisition model to obtain a second feature corresponding to each keyword. Based on the first feature and the second feature corresponding to each keyword, text features of the audio description text are obtained.

[0196] The video features, text features and audio noise features are fused and split to obtain split video features, split text features and split audio noise features.

[0197] Alignment features and global features are determined through video features, text features, video duration features, and audio-visual synchronization features.

[0198] The audio feature generation model is used to process the alignment features, global features, split video features, split text features and split audio noise features to obtain the audio features of the audio corresponding to the video.

[0199] The audio features are processed by the audio encoder to obtain the Mel spectrum corresponding to the audio features.

[0200] The mel spectrogram is processed by a vocoder to obtain the audio corresponding to the video.

[0201] Figure 7 FIG. 1 is a block diagram of an audio generating device according to an exemplary embodiment. Figure 7 , the device includes the following contents.

[0202] The acquisition unit 701 is configured to acquire a video for generating audio.

[0203] The determination unit 702 is configured to determine the video features, video duration features and audio-visual synchronization features corresponding to the video, where the video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance of the video's picture and sound.

[0204] The generating unit 703 is configured to generate audio corresponding to the video according to the video feature, the video duration feature and the audio-video synchronization feature, where the duration of the audio is the same as the duration of the video.

[0205] In some embodiments, the acquiring unit 701 is further configured to acquire an audio description text, where the audio description text is a description text of the audio corresponding to the video.

[0206] The determining unit 702 is further configured to determine text features corresponding to the audio description text.

[0207] The generating unit 703 is configured to generate audio corresponding to the video according to the video features, video duration features, audio-visual synchronization features and text features.

[0208] In some embodiments, the determination unit 702 is configured to execute the steps of obtaining a first feature corresponding to the audio description text, where the first feature is used to characterize the audio description text; performing keyword extraction on the audio description text to obtain at least one keyword; obtaining a second feature corresponding to each keyword, where the second feature corresponding to each keyword is used to characterize the corresponding keyword; and determining the text feature corresponding to the audio description text based on the first feature and the second feature corresponding to each keyword.

[0209] In some embodiments, the acquisition unit 701 is configured to process the audio description text through a first text feature acquisition model to obtain a first feature corresponding to the audio description text, and the first text feature acquisition model is used to acquire text features.

[0210] In some embodiments, the acquisition unit 701 is configured to execute processing of any keyword among the keywords through a second text feature acquisition model to obtain a second feature corresponding to any keyword, the second text feature acquisition model is used to acquire text features, and the number of tokens supported by the second text feature acquisition model is less than the number of tokens supported by the first text feature acquisition model.

[0211] In some embodiments, the generation unit 703 is configured to perform a fusion of video features, text features and audio noise features to obtain a fused feature, and the audio noise feature has the same dimension as the audio feature of the audio corresponding to the video; split the fused feature to obtain split video features, split text features and split audio noise features, and the split video features, split text features and split audio noise features are all fused with video features, text features and audio noise features; generate an alignment feature based on the audio-visual synchronization feature, and the dimension of the alignment feature is the same as the dimension of the audio noise feature; generate a global feature based on the text feature, video feature, video duration feature and alignment feature, and the global feature is the feature after considering the text feature, video feature, video duration feature and alignment feature; generate the audio corresponding to the video based on the alignment feature, the global feature, the split video feature, the split text feature and the split audio noise feature.

[0212] In some embodiments, the generating unit 703 is configured to upsample the audio-video synchronization feature to the same dimension as the audio noise feature to obtain an alignment feature.

[0213] In some embodiments, the generation unit 703 is configured to perform pooling of text features, video features, and video duration features to obtain pooled features, which are used to characterize the consistency between the audio description text and the video; and generate global features based on the pooled features and the alignment features.

[0214] In some embodiments, the generating unit 703 is configured to perform addition of the pooled features and the aligned features to obtain the global features.

[0215] In some embodiments, the generation unit 703 is configured to determine the audio features of the audio corresponding to the video based on the alignment features, global features, split video features, split text features and split audio noise features; and generate the audio corresponding to the video based on the audio features.

[0216] In some embodiments, the generation unit 703 is configured to process the alignment features, global features, split video features, split text features and split audio noise features through an audio feature generation model to obtain audio features of the audio corresponding to the video, and the audio feature generation model is used to generate audio features.

[0217] In some embodiments, the determination unit 702 is configured to perform frame extraction processing on the video to obtain multiple pictures included in the video; determine the picture features of each picture, and the picture features of each picture are used to represent the corresponding picture; and determine the video features corresponding to the video based on the picture features of each picture.

[0218] In some embodiments, the determination unit 702 is configured to process any one of the pictures through a picture feature acquisition model to obtain picture features of any one of the pictures, where the picture feature acquisition model is used to obtain picture features.

[0219] In some embodiments, the determination unit 702 is configured to obtain the number of frames and frame rate of the video; determine the duration of the video based on the number of frames and frame rate of the video; and determine the video duration feature corresponding to the video based on the duration of the video.

[0220] In some embodiments, the determination unit 702 is configured to determine the quotient between the number of frames and the frame rate of the video; when the quotient is an integer, the quotient is determined to be the duration of the video; when the quotient is a non-integer, the quotient is rounded up to obtain the duration of the video.

[0221] In some embodiments, the determination unit 702 is configured to map the duration of the video into an array; process the array through a duration learning model to obtain a video duration feature corresponding to the video, and the duration learning model is used to determine the video duration feature.

[0222] In some embodiments, the duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer, and a third convolutional layer; the determination unit 702 is configured to execute the input of the array into the fully connected layer to obtain the output result of the fully connected layer; the output result of the fully connected layer is input into the first convolutional layer and the second convolutional layer respectively to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; the output result of the first convolutional layer is mapped through an activation function to obtain an attention weight; the attention weight is dot-multiplied with the output result of the second convolutional layer to obtain a dot-multiplication result; the dot-multiplication result is input into the third convolutional layer to obtain the video duration feature of the video.

[0223] In some embodiments, the determining unit 702 is configured to process the video through an audio-visual synchronization encoder to obtain an audio-visual synchronization feature corresponding to the video, where the audio-visual synchronization encoder is used to obtain the audio-visual synchronization feature.

[0224] The disclosed embodiment provides an audio generation device, which generates audio corresponding to a video based on the video, without the need for manual operation by the user, thereby simplifying the audio generation method and making the audio generation more efficient. Moreover, there is no need to rely on the user's subjective consciousness, so that the matching degree between the generated audio and the video is higher, and the audio generation efficiency is better. In addition, since not only the video features but also the video duration features and the audio-visual synchronization features are obtained, when generating audio, not only the video features of the video itself but also the video duration features and the audio-visual synchronization features are taken into account, so that the duration of the generated audio is the same as the duration of the video, and the generated audio is synchronized with the video audio and video, further improving the matching degree between the generated audio and the video and the audio generation efficiency.

[0225] Regarding the apparatus in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment of the method, and will not be elaborated on here.

[0226] Figure 8 The following is a block diagram illustrating the structure of a terminal 800 provided by an exemplary embodiment of the present disclosure. Terminal 800 may be a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. Terminal 800 may also be referred to as user equipment, portable terminal, laptop terminal, desktop terminal, or other similar names.

[0227] Typically, the terminal 800 includes a processor 801 and a memory 802 .

[0228] Processor 801 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 801 may be implemented in hardware using at least one of the following: a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), or a PLA (Programmable Logic Array). Processor 801 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 801 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content required for display. In some embodiments, processor 801 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0229] The memory 802 may include one or more computer-readable storage media, which may be non-transitory. The memory 802 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 802 is used to store at least one program code, which is executed by the processor 801 to implement the audio generation method provided in the method embodiment of the present disclosure.

[0230] In some embodiments, terminal 800 may optionally include a peripheral device interface 803 and at least one peripheral device. The processor 801, memory 802, and peripheral device interface 803 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 803 via a bus, signal lines, or circuit boards. Specifically, the peripheral device may include at least one of a radio frequency circuit 804, a display screen 805, a camera assembly 806, an audio circuit 807, and a power supply 808.

[0231] The peripheral device interface 803 can be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 801 and the memory 802. In some embodiments, the processor 801, the memory 802, and the peripheral device interface 803 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 801, the memory 802, and the peripheral device interface 803 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0232] The RF circuit 804 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 804 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 804 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the RF circuit 804 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, and the like. The RF circuit 804 can communicate with other terminals via at least one wireless communication protocol. Such wireless communication protocols include, but are not limited to, metropolitan area networks, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 804 may also include circuitry related to Near Field Communication (NFC), although this disclosure does not limit this.

[0233] Display screen 805 is used to display a user interface (UI). This UI may include graphics, text, icons, videos, or any combination thereof. If display screen 805 is a touchscreen display, it is also capable of detecting touch signals on or above the surface of display screen 805. These touch signals can be input as control signals to processor 801 for processing. Display screen 805 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there can be a single display screen 805, located on the front panel of terminal 800. In other embodiments, there can be at least two display screens 805, located on different surfaces of terminal 800 or in a foldable design. In still other embodiments, display screen 805 can be a flexible display, located on a curved or foldable surface of terminal 800. Display screen 805 can also be configured as a non-rectangular, irregular shape, also known as a special-shaped screen. Display screen 805 can be made of materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0234] The camera assembly 806 is used to capture images or videos. Optionally, the camera assembly 806 includes a front camera and a rear camera. Typically, the front camera is set on the front panel of the terminal, and the rear camera is set on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions. In some embodiments, the camera assembly 806 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation at different color temperatures.

[0235] The audio circuit 807 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals that are input into the processor 801 for processing, or input into the radio frequency circuit 804 to achieve voice communication. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, each located in different parts of the terminal 800. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert electrical signals from the processor 801 or the radio frequency circuit 804 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert electrical signals into sound waves audible to humans, but also convert electrical signals into sound waves inaudible to humans for purposes such as ranging. In some embodiments, the audio circuit 807 may also include a headphone jack.

[0236] Power supply 808 is used to power various components in terminal 800. Power supply 808 can be AC ​​power, DC power, a disposable battery, or a rechargeable battery. When power supply 808 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.

[0237] Those skilled in the art will understand that Figure 8 The structure shown in the figure does not constitute a limitation on the terminal 800, and the terminal 800 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0238] Figure 9 The following is a block diagram of a server 900 provided by an exemplary embodiment of the present disclosure. The server 900 may vary significantly due to different configurations or performance, and may include one or more processors (Central Processing Units, CPUs) 901 and one or more memories 902. The one or more memories 902 store at least one program code, which is loaded and executed by the one or more processors 901 to implement the audio generation methods provided by the above-mentioned various method embodiments. Of course, the server 900 may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input and output. The server 900 may also include other components for implementing device functions, which will not be described in detail here.

[0239] In an exemplary embodiment, a computer-readable storage medium including instructions, such as a memory including instructions, is also provided. The instructions are executable by a processor of a terminal to perform the above-described audio generation method. Alternatively, the computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, or the like.

[0240] In an exemplary embodiment, a computer program product is also provided, including a computer program that, when executed by a processor, implements the above-described audio generation method. In some embodiments, the computer program product involved in the embodiments of the present disclosure can be deployed and executed on a single terminal, or on multiple terminals located in a single location, or on multiple terminals distributed in multiple locations and interconnected by a communication network. Multiple terminals distributed in multiple locations and interconnected by a communication network can constitute a blockchain system.

[0241] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The description and examples are to be regarded as exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims. All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of the present application, and will not be described in detail here.

[0242] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.

Claims

1. An audio generation method, characterized in that: The method comprises: Obtaining a video and an audio description text of the audio to be generated, wherein the audio description text is a description text of the audio corresponding to the video; Determining video features, video duration features, and audio-visual synchronization features corresponding to the video, wherein the video features are used to characterize the video, the video duration features are used to characterize the duration of the video, and the audio-visual synchronization features are used to characterize the temporal consistency and semantic relevance of the video's image and sound; Determining text features corresponding to the audio description text; An audio corresponding to the video is generated according to the video feature, the video duration feature, the audio-visual synchronization feature, and the text feature, where the duration of the audio is the same as the duration of the video.

2. The method according to claim 1, characterized in that Determining the text features corresponding to the audio description text includes: Obtaining a first feature corresponding to the audio description text, where the first feature is used to characterize the audio description text; Performing keyword extraction on the audio description text to obtain at least one keyword; Obtaining the second feature corresponding to each keyword, where the second feature corresponding to each keyword is used to characterize the corresponding keyword; The text feature corresponding to the audio description text is determined according to the first feature and the second feature corresponding to each keyword.

3. The method according to claim 2, characterized in that The obtaining of the first feature corresponding to the audio description text includes: The audio description text is processed by a first text feature acquisition model to obtain a first feature corresponding to the audio description text, where the first text feature acquisition model is used to acquire text features.

4. The method according to claim 2, characterized in that The obtaining of the second feature corresponding to each keyword includes: For any keyword among the keywords, the keyword is processed by a second text feature acquisition model to obtain a second feature corresponding to the keyword. The second text feature acquisition model is used to obtain text features, and the number of tokens supported by the second text feature acquisition model is less than the number of tokens supported by the first text feature acquisition model.

5. The method according to claim 1, characterized in that Generating audio corresponding to the video according to the video feature, the video duration feature, the audio-visual synchronization feature, and the text feature includes: fusing the video feature, the text feature, and the audio noise feature to obtain a fused feature, wherein the audio noise feature has the same dimension as the audio feature of the audio corresponding to the video; Splitting the fused features to obtain split video features, split text features, and split audio noise features, wherein the split video features, the split text features, and the split audio noise features are all fused with the video features, the text features, and the audio noise features; generating an alignment feature according to the audio-video synchronization feature, wherein the dimension of the alignment feature is the same as the dimension of the audio noise feature; Generate a global feature based on the text feature, the video feature, the video duration feature, and the alignment feature, wherein the global feature is a feature after considering the text feature, the video feature, the video duration feature, and the alignment feature; The audio corresponding to the video is generated according to the alignment features, the global features, the split video features, the split text features and the split audio noise features.

6. The method according to claim 5, characterized in that Generating an alignment feature according to the audio-video synchronization feature includes: The audio-video synchronization feature is up-sampled to the same dimension as the audio noise feature to obtain the alignment feature.

7. The method according to claim 5, characterized in that Generating a global feature according to the text feature, the video feature, the video duration feature, and the alignment feature includes: Pooling the text features, the video features, and the video duration features to obtain pooled features, wherein the pooled features are used to characterize the consistency between the audio description text and the video; The global feature is generated according to the pooled feature and the aligned feature.

8. The method according to claim 7, characterized in that Generating the global feature according to the pooling feature and the alignment feature includes: The pooled features and the aligned features are added together to obtain the global features.

9. The method according to claim 5, characterized in that Generating the audio corresponding to the video according to the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature includes: Determining audio features of the audio corresponding to the video based on the alignment features, the global features, the split video features, the split text features, and the split audio noise features; Generate audio corresponding to the video according to the audio features.

10. The method according to claim 9, characterized in that The determining, based on the alignment feature, the global feature, the split video feature, the split text feature, and the split audio noise feature, of the audio feature corresponding to the video includes: The alignment features, the global features, the split video features, the split text features and the split audio noise features are processed by an audio feature generation model to obtain audio features of the audio corresponding to the video. The audio feature generation model is used to generate audio features.

11. The method according to any one of claims 1 to 10, characterized in that: Determining video features corresponding to the video includes: Performing frame extraction processing on the video to obtain multiple pictures included in the video; Determine the image features of each image, where the image features of each image are used to characterize the corresponding image; Determine video features corresponding to the video based on the image features of the images.

12. The method according to claim 11, characterized in that Determining the image features of each image includes: For any one of the pictures, the picture feature acquisition model is used to process the picture to obtain the picture feature of the picture. The picture feature acquisition model is used to obtain the picture feature.

13. The method according to any one of claims 1 to 10, characterized in that: Determining a video duration feature corresponding to the video includes: Obtaining the number of frames and frame rate of the video; Determine the duration of the video according to the number of frames and the frame rate of the video; Determine a video duration feature corresponding to the video based on the duration of the video.

14. The method according to claim 13, characterized in that Determining the duration of the video according to the number of frames and the frame rate of the video includes: determining a quotient between the number of frames and the frame rate of the video; When the quotient is an integer, determining the quotient as the duration of the video; When the quotient is a non-integer, the quotient is rounded up to obtain the duration of the video.

15. The method according to claim 13, characterized in that Determining a video duration feature corresponding to the video according to the duration of the video includes: Map the duration of the video to an array; The array is processed through a duration learning model to obtain a video duration feature corresponding to the video, and the duration learning model is used to determine the video duration feature.

16. The method according to claim 15, characterized in that The duration learning model includes a fully connected layer, a first convolutional layer, a second convolutional layer and a third convolutional layer; The processing of the array by the duration learning model to obtain a video duration feature corresponding to the video includes: Inputting the array into a fully connected layer to obtain an output result of the fully connected layer; Inputting the output result of the fully connected layer into the first convolutional layer and the second convolutional layer respectively, to obtain the output result of the first convolutional layer and the output result of the second convolutional layer; Mapping the output of the first convolutional layer through an activation function to obtain an attention weight; Multiply the attention weight by the output result of the second convolutional layer to obtain a dot product result; The dot product result is input into the third convolutional layer to obtain the video duration feature of the video.

17. The method according to any one of claims 1 to 10, characterized in that Determining audio and video synchronization features corresponding to the video includes: The video is processed by an audio-visual synchronization encoder to obtain an audio-visual synchronization feature corresponding to the video, and the audio-visual synchronization encoder is used to obtain the audio-visual synchronization feature.

18. An audio generating device, characterized in that: The device comprises: An acquiring unit is configured to acquire a video and an audio description text of the audio to be generated, wherein the audio description text is a description text of the audio corresponding to the video; a determining unit configured to determine a video feature, a video duration feature, and an audio-visual synchronization feature corresponding to the video, wherein the video feature is used to characterize the video, the video duration feature is used to characterize the duration of the video, and the audio-visual synchronization feature is used to characterize the temporal consistency and semantic relevance of the image and sound of the video; The determining unit is further configured to determine text features corresponding to the audio description text; The generating unit is configured to generate audio corresponding to the video according to the video feature, the video duration feature, the audio-visual synchronization feature and the text feature, wherein the duration of the audio is the same as the duration of the video.

19. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the audio generation method according to any one of claims 1 to 17.

20. A computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to execute the audio generation method according to any one of claims 1 to 17.

21. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the audio generation method according to any one of claims 1 to 17 is implemented.

Citation Information

Patent Citations

  • Silent video mimicking method, electronic equipment and storage medium

    CN118828050A

  • Audio generation method and device, electronic equipment and storage medium

    CN119255028A