Audio generation method, and deep learning model training method and device
By responding to user selections, extracting visual features of the target object in the video and adjusting the volume, the problem of inaccurate audio generation in complex scenarios in existing technologies is solved, achieving high-quality and personalized audio generation effects.
Patent Information
- Application Number
- CN202511341071.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-18
- Publication Date
- 2025-11-18
AI Technical Summary
Existing video-to-audio technologies struggle to generate audio that matches specific targets in complex scenarios and lack user-interactive and personalized audio generation capabilities.
By responding to the user's interactive selection, the visual features of the target object in the video are extracted, an initial audio matching the target object's actions is generated using a deep learning model, and the volume is adjusted according to the proportion of the screen to generate high-quality target audio.
It enables fine-grained audio generation for specific targets in complex scenarios, provides intuitive user interaction capabilities, meets personalized audio generation needs, and improves the quality and realism of audio generation.
Smart Images

Figure CN120977339A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence, in particular to the technical fields of deep learning, contrastive learning and computer vision. More specifically, the present disclosure provides an audio generation method, a training method of a deep learning model, an apparatus, an electronic device, a storage medium and a computer program product. BACKGROUND
[0002] With the development of artificial intelligence technology, video generation has become more and more widespread, and generating matching audio for video has become an important research direction. SUMMARY
[0003] The present disclosure provides an audio generation method, a training method of a deep learning model, an apparatus, an electronic device, a storage medium and a computer program product.
[0004] According to a first aspect, an audio generation method is provided, the method comprising: in response to an interactive selection operation for a target object in a video, extracting visual features of the target object in each video frame of the video; generating an initial audio matching the action of the target object according to the visual features; and adjusting the volume of the initial audio according to the picture proportion of the target object in each video frame of the video to obtain a target audio.
[0005] According to a second aspect, a training method of a deep learning model is provided, the method comprising: obtaining sample data, wherein the sample data comprises a spliced video obtained by splicing video clips of at least two objects and mixed audio obtained by mixing audio clips of at least two objects; calculating the similarity between at least two video clips in the spliced video and at least two audio clips in the mixed audio using a deep learning model; determining the loss of the deep learning model according to the similarity; and adjusting the parameters of the deep learning model according to the loss.
[0006] According to a third aspect, an audio generation apparatus is provided, the apparatus comprising: a visual feature determination module configured to extract visual features of a target object in each video frame of a video in response to an interactive selection operation for the target object in the video; an initial audio generation module configured to generate an initial audio matching the action of the target object according to the visual features; and a target audio generation module configured to adjust the volume of the initial audio according to the picture proportion of the target object in each video frame of the video to obtain a target audio.
[0007] According to a fourth aspect, a training apparatus of a deep learning model is provided, the apparatus comprising: an obtaining module configured to obtain sample data, wherein the sample data comprises a spliced video spliced from video clips of at least two objects and mixed audio mixed from audio clips of the at least two objects; a similarity determining module configured to calculate similarities between the at least two video clips in the spliced video and the at least two audio clips in the mixed audio respectively by using the deep learning model; a loss determining module configured to determine a loss of the deep learning model according to the similarities; and a parameter adjusting module configured to adjust parameters of the deep learning model according to the loss.
[0008] According to a fifth aspect, an electronic device is provided, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method provided by the present disclosure.
[0009] According to a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, the computer instructions being used to cause a computer to execute the method provided by the present disclosure.
[0010] According to a seventh aspect, a computer program product is provided, comprising a computer program stored in at least one of a readable storage medium and an electronic device, and the computer program, when executed by a processor, implements the method provided by the present disclosure.
[0011] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS
[0012] The accompanying drawings are used to better understand the present scheme, and do not limit the present disclosure. Among them:
[0013] Figure 1 is an exemplary system architecture schematic diagram of an audio generation method and a training method of a deep learning model according to an embodiment of the present disclosure;
[0014] Figure 2 is a flowchart of an audio generation method according to an embodiment of the present disclosure;
[0015] Figure 3 is a schematic diagram of an audio generation method according to an embodiment of the present disclosure;
[0016] Figure 4 is a flowchart of a training method of a deep learning model according to an embodiment of the present disclosure;
[0017] Figure 5 is a flowchart of a training method of a deep learning model according to an embodiment of the disclosure;
[0018] Figure 6 is a block diagram of an audio generation device according to an embodiment of the disclosure;
[0019] Figure 7 is a block diagram of a training device of a deep learning model according to an embodiment of the disclosure; and
[0020] Figure 8 is a block diagram of an electronic device for at least one of an audio generation method and a training method of a deep learning model according to an embodiment of the disclosure. DETAILED DESCRIPTION
[0021] Exemplary embodiments of the disclosure are described herein with reference to the accompanying drawings, which are included to provide a thorough understanding of embodiments of the disclosure by a person of ordinary skill in the art, and should not be construed as limiting the scope of the disclosure. Accordingly, persons of ordinary skill in the art will recognize that various modifications and changes can be made to the embodiments described herein without departing from the scope and spirit of the disclosure. Also, descriptions of well-known functions and constructions are omitted for clarity and conciseness.
[0022] Current video-to-audio techniques are mainly based on global understanding of the whole video to generate audio that roughly matches the video content. For example, if there is only one object (e.g., a dog) in the video, a "dog bark" or similar sound may be generated.
[0023] Such a method often has difficulty generating sounds that match a specific target when facing complex scenes (e.g., multiple objects, multiple actions). In addition, existing methods lack the ability to respond to user-specified targets and cannot meet the needs of interactive and personalized audio generation.
[0024] In the technical solutions of the disclosure, the collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solutions comply with relevant laws and regulations and do not violate public order and good customs.
[0025] In the technical solutions of the disclosure, the user's authorization or consent is obtained before the user's personal information is acquired or collected.
[0026] Figure 1 is an exemplary system architecture schematic diagram according to an embodiment of the disclosure, which can apply the audio generation method and the training method of the deep learning model. It should be noted that, Figure 1The examples shown are merely examples of system architectures that can be applied to the embodiments of this disclosure, in order to help those skilled in the art understand the technical content of this disclosure, but do not mean that the embodiments of this disclosure cannot be used in other devices, systems, environments or scenarios.
[0027] like Figure 1 As shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, and 103, a network 104, and a server 105. The network 104 serves as a medium for providing a communication link between the terminal devices 101, 102, and 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0028] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices, including but not limited to smartphones, tablets, laptops, etc.
[0029] Server 105 can be a server that provides various services, such as a backend management server that supports websites browsed by users using terminal devices 101, 102, and 103 (for example only). The backend management server can analyze and process data such as received user requests and feed the processing results back to the terminal devices.
[0030] The audio generation method and deep learning model training method provided in this disclosure can generally be executed by the server 105. Accordingly, the audio generation device and deep learning model training device provided in this disclosure can generally be located in the server 105.
[0031] Figure 2 This is a flowchart of an audio generation method according to an embodiment of the present disclosure.
[0032] like Figure 2 As shown, the audio generation method 200 includes operations S210 to S230.
[0033] In operation S210, in response to an interactive selection operation for a target object in the video, visual features of the target object in each video segment are extracted.
[0034] The video can be silent and can include multiple identical objects, such as multiple cows. Alternatively, it can include multiple different objects, such as at least one cow, at least one sheep, etc. Furthermore, the video can include background elements, such as grass or sky.
[0035] Users can interactively select objects in the video. For example, a user can click on any object in the video and select that object as the target object.
[0036] For example, a pre-trained object segmentation model can be used to segment each object in the current frame (e.g., the first frame of a video) to determine the target region for each object. The target region can be represented by the coordinates of the two opposite vertices of the bounding box. In response to a user's selection, the object selected can be determined based on which target region the coordinates of the user's click belong to.
[0037] After identifying the target object, the target region can be extracted using masking. For example, the mask is an image of the same size as the video frame. In the mask, the pixel value of the target region containing the target object can be 1, while the pixel values of other regions outside the target region are set to 0. By multiplying the video frame image with the target object's mask, a video frame containing only the pixel information of the target region can be obtained.
[0038] After determining the target region of the target object in the current frame, it is also necessary to determine the target region of the target object in each frame of the video. For example, a target tracking algorithm can be used to determine the position of the target object in each frame, and then masking can be used to determine the target region of the target object in each frame, thus obtaining a sequence of target regions of the target object.
[0039] The target region sequence contains only the pixel information of the target object, while other areas (such as grass and sky) are occluded. A visual encoder can be used to extract features from the target region sequence to obtain the visual features of the target object in each video frame. These visual features can include characteristics such as the target object's action and shape.
[0040] According to embodiments of this disclosure, in response to a user's interactive selection operation, a target object specified by the user is determined. Through masking, the visual features of the target object are extracted, which can accurately extract features such as the target object's actions, behaviors, and postures, and eliminate the influence of other objects and backgrounds in the video, which is beneficial for subsequent implementation of interactive, fine-grained object audio.
[0041] In operation S220, an initial audio is generated that matches the action of the target object based on visual features.
[0042] For example, video features can be input into a trained audio-video matching model, which maps visual features to a latent space aligned with audio features, outputting a vector that represents "who the object is and what it is doing." This vector can then be used as a control condition and fed to a diffusion model, which iteratively denoises a segment of random noise into the initial audio.
[0043] Audio-video matching models can be trained using sample videos and sample audio through contrastive learning. Sample videos and sample audio form sample pairs; positive sample pairs are those where the sample videos and sample audio match, while negative sample pairs are those where the sample videos and sample audio do not match.
[0044] For example, in a positive sample pair, the sample video is a video containing the action of the sample object (e.g., a cow running), and the sample audio is the audio corresponding to the sample video containing the sound of the sample object (e.g., the sound of cow hooves). In a negative sample pair, the sample video is a video containing the action of the sample object (e.g., a cow running), and the sample audio is audio that does not correspond to the sample video but contains the sound of the same or different sample objects (e.g., a dog barking, a cow mooing, etc.).
[0045] During training, sample pairs are input into the model. The model extracts visual features of sample objects from sample videos and audio features from sample audio, and then calculates the similarity between the visual and audio features. The model is trained with the constraint that the similarity between visual and audio features in positive samples is relatively high (e.g., tending to 1), and the similarity between visual and audio features in negative samples is relatively high (e.g., tending to 0), so that the model gradually learns the matching rules.
[0046] A trained audio-video matching model can distinguish which audio and video segments match and which do not. Therefore, by inputting visual features into the trained audio-video matching model, the model can project the visual features into a latent space aligned with the audio features, and the output feature vector can serve as audio control conditions.
[0047] The diffusion model is used to progressively denoise an initial noise source, guided by audio control conditions, step by step "reconstructing" the sound content to obtain an audio segment. For example, the initial noise can be a random noise source of the same length as the desired audio, such as the same length as the video. The initial noise and audio control conditions are input into the diffusion model, which performs an iterative denoising process, removing a portion of the noise from the initial noise each time, ultimately converting the noise into audio (the initial audio) that matches the action and form of the target object in the video.
[0048] For example, if the target object (cow) in the video is running, the diffusion model can output a "cow hoof sound" corresponding to "cow running".
[0049] In operation S230, the volume of the initial audio is adjusted according to the proportion of the target object in each video frame to obtain the target audio.
[0050] For example, the initial audio could be a piece of audio with a standard volume, meaning the volume remains the same throughout. However, the distance to the target object in the video might change, such as a cow moving further away or closer. If the volume remains constant regardless of the change in the target object's distance, the audio will lack a sense of space, resulting in an unrealistic user experience.
[0051] Therefore, this embodiment adjusts the volume of the initial audio according to the proportion of the target object in each video frame, producing an effect of near objects appearing larger and far objects appearing smaller, giving the sound a sense of space and making the user experience more realistic.
[0052] For example, the proportion of the target object's area to the overall screen area in each frame can be calculated as the target object's screen proportion. The initial audio volume can then be adjusted based on this proportion. For instance, the initial volume can be multiplied by the screen proportion to obtain the target volume.
[0053] Ultimately, a high-quality audio clip can be generated that matches the actions of the target object in the video and dynamically changes according to the distance of the target object.
[0054] The embodiments of this disclosure respond to the user's interactive selection operation, determine the target object selected by the user, extract the visual features of the target object, and generate initial audio that matches the action of the target object based on the visual features. The volume of the initial audio is dynamically adjusted according to the screen proportion of the target object, which can obtain audio that accurately matches the target object in the video and has a sense of space, thereby improving the audio generation quality.
[0055] Compared to related technologies that generate roughly matching background audio based on global video features and cannot distinguish multiple objects in the video, the embodiments of this disclosure extract the visual features of the target object selected by the user in response to the user's selection operation. This allows the audio generation to focus only on the target object specified by the user, enabling fine-grained, target-specific audio generation in complex scenes containing multiple objects. Furthermore, it provides intuitive user interaction capabilities and meets the needs of personalized audio generation.
[0056] Figure 3 This is a schematic diagram of an audio generation method according to an embodiment of the present disclosure.
[0057] like Figure 3 As shown, this embodiment includes video 301, target region sequence 302, visual encoder 310, audio-video matching model 320, diffusion model 330, initial audio 303, and target audio 304.
[0058] Users can select objects in video 301, for example, by clicking on an object in the video, and that selected object becomes the target object.
[0059] According to embodiments of this disclosure, in response to an interactive selection operation for a target object in a video, a target region of the target object in each video frame of the video is determined; and based on the target region, visual features of the target object in each video frame are extracted.
[0060] For example, a video can include multiple video frames, denoted as V = {V1, V2, ..., VT}, where T is the number of video frames, such as 30 frames. After a user clicks on a target object in a certain frame of the video, the mask of the target object in each frame can be determined, denoted as M = {M1, M2, ..., MT}. Next, video masking can be performed, which can be expressed as Vmasked = V⊙M. This means that based on the original video frames and the mask sequence, the target area of the target object in each frame is preserved, while other areas are masked.
[0061] The video frames processed by video masking can then be input into the mask-guided visual encoder 310 to obtain the visual feature sequence of the target object. The visual feature sequence represents the target object's actions, shapes, etc.
[0062] According to embodiments of this disclosure, visual features can be encoded using an audio-visual matching model to obtain audio control conditions; and a diffusion model can be used to denoise preset noise based on the audio control conditions to obtain initial audio.
[0063] The visual feature sequence of the target object is input into the audio-visual matching model 320 trained by contrastive learning. The audio-visual matching model 320 can project the visual features into a latent space aligned with the audio features, and the output feature vector can be used as an audio control condition.
[0064] In one example, the audio-video matching model 320 can be trained using a spliced video obtained by splicing video clips of at least two objects and a mixed audio obtained by mixing audio clips of at least two objects.
[0065] For example, a video clip containing "cows running" and a video clip containing "dogs opening their mouths" can be spliced together to obtain a spliced video. An audio clip containing "cow hoof sounds" can be mixed with an audio clip containing "dog barks." The spliced video and the mixed audio are then input into an audio-visual matching model. This increases the difficulty of the samples, allowing the model to learn audio-visual matching in more complex scenarios and improving its ability to distinguish and correctly match audio and video.
[0066] After obtaining the audio control conditions generated by the audio-video matching model 320, the audio control conditions can be input into the diffusion model 330. Guided by the audio control conditions, the diffusion model 330 iteratively denoises the preset noise, gradually transforming the preset noise into an initial audio segment that matches the action of the target object in the video. For example, the preset noise can be random noise following a normal distribution with the same length as the video (e.g., T frames), which can be denoted as zT~N(0,1). The initial audio 303 generated by the diffusion model 330 can be denoted as At=Decoder(zT), where At can represent the amplitude of the initial audio 303 in each frame, and t is any frame in T frames.
[0067] According to embodiments of this disclosure, the screen proportion of the target object in each video frame is determined; based on the screen proportion of the target object in each video frame, the volume of the initial audio frame corresponding to the video frame is adjusted to obtain the target audio.
[0068] For example, the percentage of the target object in each frame can be expressed as λt = target area / total frame area. Adjusting the amplitude of the initial audio based on this percentage, the amplitude of the target audio 304 can be expressed as At′ = At × λt. This means that the closer the target object is, the larger its percentage in the frame, and the louder the sound. Therefore, the target audio 304 has a sense of distance depending on the target object, making the user experience more realistic.
[0069] According to embodiments of this disclosure, this disclosure also provides a method for training a deep learning model.
[0070] Figure 4 This is a flowchart of a training method for a deep learning model according to an embodiment of the present disclosure.
[0071] like Figure 4 As shown, the training method 400 of the deep learning model includes operations S410 to S440.
[0072] In operation S410, sample data is acquired, including spliced video obtained by splicing video clips of at least two objects, and mixed audio obtained by mixing audio clips of at least two objects.
[0073] Deep learning models can be audio / video matching models, which may include a video encoder and an audio encoder. The video encoder extracts visual features from the video, and the audio encoder extracts audio features from the audio. The audio / video matching model aims to make the visual and audio features of matching audio / video pairs closer in feature space (higher similarity), while making the visual and audio features of mismatched audio / video pairs farther apart in feature space (lower similarity).
[0074] In order to enable the model to learn audio and video matching in complex scenes containing multiple objects, embodiments of this disclosure select spliced video and mixed audio as sample pairs in sample construction.
[0075] For example, a video clip containing a "cow running" scene can be spliced with a video clip containing a "dog opening its mouth" scene to obtain a spliced video. An audio clip containing the sound of "cow hooves" can be mixed with an audio clip containing the sound of "dog barking." Specifically, the video clip containing the "cow running" scene is matched with the audio clip containing the sound of "cow hooves," and the video clip containing the "dog opening its mouth" scene is matched with the audio clip containing the sound of "dog barking."
[0076] In operation S420, a deep learning model is used to calculate the similarity between at least two video segments in the spliced video and at least two audio segments in the mixed audio.
[0077] The spliced video and mixed audio are input into the audio-video matching model. The model automatically distinguishes video clips and audio clips of different objects, and calculates the similarity between different video clips and different audio clips.
[0078] For example, the similarity between a video clip of a "cow running" scene and an audio clip of "cow hoof sounds" can be calculated, the similarity between a video clip of a "cow running" scene and an audio clip of "dog barking" can be calculated, the similarity between a video clip of a "dog opening its mouth" scene and an audio clip of "cow hoof sounds" can be calculated, and the similarity between a video clip of a "dog opening its mouth" scene and an audio clip of "dog barking" can be calculated.
[0079] In the S430 operation, the loss of the deep learning model is determined based on the similarity.
[0080] The loss of a deep learning model can be calculated by constraining the similarity between matching video and audio segments to be high (e.g., tending to 1) and the similarity between non-matching video and audio segments to be low (e.g., tending to 0).
[0081] For example, positive sample loss can be determined based on the difference between the similarity score of the video clip of the "cow running" scene and the audio clip of the "cow hoof sound" and a value of 1, and the difference between the similarity score of the video clip of the "dog opening its mouth" scene and the audio clip of the "dog barking" scene and a value of 1. Negative sample loss can be determined based on the difference between the similarity score of the video clip of the "cow running" scene and the audio clip of the "dog barking" scene and a value of 0, and the difference between the similarity score of the video clip of the "dog opening its mouth" scene and the audio clip of the "cow hoof sound" scene and a value of 0.
[0082] The overall loss of a deep learning model can be determined based on the positive sample loss and the negative sample loss.
[0083] When operating the S440, adjust the parameters of the deep learning model based on the loss.
[0084] The parameters of the deep learning model can be adjusted based on the overall loss, so that the deep learning model can gradually improve its ability to match the correct audio and video until the training is completed, and the trained deep learning model is obtained.
[0085] According to embodiments of this disclosure, by using a spliced video obtained by splicing video clips from at least two objects and a mixed audio obtained by mixing audio clips from at least two objects to train an audio-video matching model, the model can learn audio-video matching in more complex scenarios, thereby improving the model's ability to distinguish and match correct audio and video.
[0086] Figure 5 This is a schematic diagram of a training method for a deep learning model according to an embodiment of the present disclosure.
[0087] like Figure 5 As shown, spliced video is a video obtained by splicing video clips from at least two objects, and mixed audio is an audio obtained by mixing audio clips from at least two objects. The video clip and audio clip for each object are matched. That is, for each object, the video clip and audio clip are matched.
[0088] According to embodiments of this disclosure, a deep learning model is used to divide a spliced video into at least two video segments corresponding to at least two objects, and to divide a mixed audio into at least two audio segments corresponding to at least two objects; object visual features are extracted from each video segment, and object audio features are extracted from each audio segment; for each video segment, the similarity between the object visual features of the video segment and the object audio features of the multiple audio segments is calculated, which is used as the similarity between the video segment and the at least two audio segments.
[0089] For example, by inputting spliced video and mixed audio into a deep learning model, the model can segment the spliced video into video clips of different objects, such as the first video clip being a clip of the target object "cow" and the second video clip being a clip of the target object "dog". Similarly, the deep learning model can segment the mixed audio into audio clips of different objects, for example, the first audio clip being an audio clip of the target object "cow" and the second audio clip being an audio clip of the target object "dog".
[0090] The visual features of the first video segment, the visual features of the second video segment, the audio features of the first audio segment, and the audio features of the second audio segment can be extracted. Then, the similarity between the visual features of each video segment and the audio features of each audio segment can be calculated.
[0091] According to embodiments of this disclosure, for each video segment, a positive sample loss is determined based on the similarity between the video segment and the target audio segment, and a negative sample loss is determined based on the similarity between the video segment and other audio segments besides the target audio segment, wherein the target audio segment and the video segment correspond to the same object; the loss of the deep learning model is determined based on the positive sample loss and the negative sample loss.
[0092] Video and audio clips of the same object are matched and can form positive sample pairs. Video and audio clips of different objects are not matched and can form negative sample pairs. For example, the first video clip and the first audio clip form a positive sample pair, the first video clip and the second audio clip form a negative sample pair, the second video clip and the first audio clip form a negative sample pair, and the second video clip and the second audio clip form a positive sample pair.
[0093] The positive sample loss is determined based on the difference between the similarity between video and audio segments in a positive sample pair and a first value (e.g., 1). The negative sample loss is determined based on the difference between the similarity between video and audio segments in a negative sample pair and a second value (e.g., 0). The positive and negative sample losses are then summed or weighted to obtain the overall loss of the deep learning model. Based on the overall loss, the parameters of the deep learning model can be updated, gradually improving the model's ability to correctly match audio and video.
[0094] According to embodiments of this disclosure, this disclosure also provides an audio generation apparatus and a training apparatus for a deep learning model.
[0095] Figure 6 This is a block diagram of an audio generation apparatus according to an embodiment of the present disclosure.
[0096] like Figure 6 As shown, the audio generation device 600 includes a visual feature determination module 610, an initial audio generation module 620, and a target audio generation module 630.
[0097] The visual feature determination module 610 is used to extract the visual features of the target object in each video frame of the video in response to an interactive selection operation for a target object in the video.
[0098] The initial audio generation module 620 is used to generate initial audio that matches the action of the target object based on visual features.
[0099] The target audio generation module 630 is used to adjust the volume of the initial audio according to the proportion of the target object in each video frame to obtain the target audio.
[0100] The initial audio generation module 620 includes a control condition determination unit and a noise reduction unit.
[0101] The control condition determination unit is used to encode visual features using the audio-video matching model to obtain audio control conditions. The audio-video matching model is trained using sample videos and sample audio through comparative learning.
[0102] The denoising unit is used to denoise preset noise based on the diffusion model and audio control conditions to obtain the initial audio.
[0103] The target audio generation module 630 includes a screen octane determination unit and a volume adjustment unit.
[0104] The frame proportion determination unit is used to determine the frame proportion of the target object in each video frame.
[0105] The volume adjustment unit is used to adjust the volume of the initial audio frame corresponding to the video frame according to the proportion of the target object in each video frame, so as to obtain the target audio.
[0106] The visual feature determination module 610 includes a target region determination unit and a visual feature determination unit.
[0107] The target region determination unit is used to determine the target region of the target object in each video frame of the video in response to an interactive selection operation for a target object in the video.
[0108] The visual feature determination unit is used to extract the visual features of the target object in each video frame based on the target region.
[0109] Figure 7 This is a block diagram of a training apparatus for a deep learning model according to an embodiment of the present disclosure.
[0110] like Figure 7 As shown, the training device 700 for the deep learning model includes an acquisition module 710, a similarity determination module 720, a loss determination module 730, and a parameter adjustment module 740.
[0111] The acquisition module 710 is used to acquire sample data, wherein the sample data includes spliced video obtained by splicing video clips of at least two objects, and mixed audio obtained by mixing audio clips of at least two objects.
[0112] The similarity determination module 720 is used to calculate the similarity between at least two video segments in the spliced video and at least two audio segments in the mixed audio using a deep learning model.
[0113] The loss determination module 730 is used to determine the loss of the deep learning model based on similarity.
[0114] The parameter tuning module 740 is used to adjust the parameters of the deep learning model based on the loss.
[0115] The similarity determination module 720 includes a segmentation unit, a feature extraction unit, and a similarity calculation unit.
[0116] The segmentation unit is used to divide a spliced video into at least two video segments corresponding to at least two objects, and to divide a mixed audio into at least two audio segments corresponding to at least two objects, using a deep learning model.
[0117] The feature extraction unit is used to extract the visual features of objects in each video segment and the audio features of objects in each audio segment.
[0118] The similarity calculation unit is used to calculate the similarity between the visual features of the video segment and the audio features of the objects of multiple audio segments for each video segment, and to obtain the similarity between the video segment and at least two audio segments.
[0119] The loss determination module 730 includes a positive and negative sample loss determination unit and an overall loss determination unit.
[0120] The positive and negative sample loss determination unit is used to determine the positive sample loss for each video segment based on the similarity between the video segment and the target audio segment, and to determine the negative sample loss based on the similarity between the video segment and other audio segments besides the target audio segment, wherein the target audio segment and the video segment correspond to the same object;
[0121] The overall loss determination unit is used to determine the loss of the deep learning model based on the positive sample loss and the negative sample loss.
[0122] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0123] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0124] like Figure 8As shown, device 800 includes a computing unit 801, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 802 or a computer program loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may also store various programs and data required for the operation of device 800. The computing unit 801, ROM 802, and RAM 803 are interconnected via bus 804. Input / output (I / O) interface 805 is also connected to bus 804.
[0125] Multiple components in device 800 are connected to I / O interface 805, including: input unit 806, such as keyboard, mouse, etc.; output unit 807, such as various types of monitors, speakers, etc.; storage unit 808, such as disk, optical disk, etc.; and communication unit 809, such as network card, modem, wireless transceiver, etc. Communication unit 809 allows device 800 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0126] The computing unit 801 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs at least one of the various methods and processes described above, such as audio generation methods and deep learning model training methods. For example, in some embodiments, at least one of the audio generation methods and deep learning model training methods can be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of at least one of the audio generation methods and deep learning model training methods described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured by any other suitable means (e.g., by means of firmware) to perform at least one of the audio generation method and the training method of the deep learning model.
[0127] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0128] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0129] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0130] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0131] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0132] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0133] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.
[0134] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. An audio generation method, comprising: In response to an interactive selection operation for a target object in a video, visual features of the target object are extracted in each video frame of the video; Based on the visual features, generate initial audio that matches the action of the target object; as well as Based on the proportion of the target object in each video frame of the video, the volume of the initial audio is adjusted to obtain the target audio.
2. The method according to claim 1, wherein, The step of generating initial audio that matches the action of the target object based on the visual features includes: The visual features are encoded using an audio-video matching model to obtain audio control conditions. The audio-video matching model is trained using sample videos and sample audio through comparative learning. Using a diffusion model, the preset noise is denoised based on the audio control conditions to obtain the initial audio.
3. The method according to claim 2, wherein, The audio-video matching model is trained using a spliced video obtained by splicing video clips of at least two objects, and a mixed audio obtained by mixing audio clips of the at least two objects.
4. The method according to claim 1, wherein, The step of adjusting the volume of the initial audio based on the screen proportion to obtain the target audio includes: Determine the percentage of the target object in each video frame of the video; Based on the proportion of the target object in each video frame of the video, the volume of the initial audio frame corresponding to the video frame is adjusted to obtain the target audio.
5. The method according to claim 1, wherein, In response to a selection operation for a target object in the video, extracting the visual features of the target object includes: In response to an interactive selection operation targeting a target object in a video, a target region of the target object is determined in each video frame of the video; Based on the target region, extract the visual features of the target object in each video frame.
6. A method for training a deep learning model, comprising: Acquire sample data, wherein the sample data includes a spliced video obtained by splicing video clips of at least two objects, and a mixed audio obtained by mixing audio clips of the at least two objects; The deep learning model is used to calculate the similarity between at least two video segments in the spliced video and at least two audio segments in the mixed audio. Based on the similarity, determine the loss of the deep learning model; and The parameters of the deep learning model are adjusted based on the loss.
7. The method according to claim 6, wherein, The step of using the deep learning model to calculate the similarity between at least two video segments in the spliced video and at least two audio segments in the mixed audio includes: Using the deep learning model, the spliced video is divided into at least two video segments corresponding to the at least two objects respectively, and the mixed audio is divided into at least two audio segments corresponding to the at least two objects respectively; Extract the visual features of objects from each video segment and extract the audio features of objects from each audio segment; For each video segment, the similarity between the visual features of the object in the video segment and the audio features of the objects in each of the plurality of audio segments is calculated, and this similarity is used as the similarity between the video segment and the at least two audio segments.
8. The method according to claim 7, wherein, The step of determining the loss of the deep learning model based on the similarity includes: For each video segment, a positive sample loss is determined based on the similarity between the video segment and the target audio segment, and a negative sample loss is determined based on the similarity between the video segment and other audio segments besides the target audio segment, wherein the target audio segment and the video segment correspond to the same object; The loss of the deep learning model is determined based on the positive sample loss and the negative sample loss.
9. An audio generation apparatus, comprising: A visual feature determination module is used to extract the visual features of the target object in each video frame of the video in response to an interactive selection operation for a target object in the video. An initial audio generation module is used to generate initial audio that matches the action of the target object based on the visual features; as well as The target audio generation module is used to adjust the volume of the initial audio according to the proportion of the target object in each video frame of the video to obtain the target audio.
10. A training device for a deep learning model, comprising: An acquisition module is used to acquire sample data, wherein the sample data includes a spliced video obtained by splicing video clips of at least two objects, and a mixed audio obtained by mixing audio clips of the at least two objects; A similarity determination module is used to calculate the similarity between at least two video segments in the spliced video and at least two audio segments in the mixed audio using the deep learning model. A loss determination module is used to determine the loss of the deep learning model based on the similarity; and The parameter adjustment module is used to adjust the parameters of the deep learning model according to the loss.
11. An electronic device, comprising: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 8.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 8.
13. A computer program product comprising a computer program stored on at least one of a readable storage medium and an electronic device, the computer program implementing the method according to any one of claims 1 to 8 when executed by a processor.