Audio and video fusion method and device, equipment, storage medium and product

By responding to user actions during video playback and replacing the audio of video characters with user-generated audio, the problem of the inability to integrate with bullet comments is solved. This allows the characters in the video to be transformed into virtual avatars of the user, enhancing the user's immersive interactive experience.

CN121509752APending Publication Date: 2026-02-10MIGU CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511439078.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-09
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

The existing bullet screen method, as additional information outside the video, cannot effectively integrate user-generated information with the internal environment of the video, resulting in a poor user experience.

Method used

During video playback, in response to the user's selection of a character in the video, the audio corresponding to the character in the video is replaced with audio generated based on the user's voice input or selected text, ensuring logical matching between the audio and the video content, and the user's audio is integrated into the video in real time to form an immersive interaction.

Benefits of technology

It enables real-time fusion of user audio and video, transforming users from external observers into active participants in the video scene and enhancing their immersive interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509752A_ABST
    Figure CN121509752A_ABST
Patent Text Reader

Abstract

The invention discloses an audio and video fusion method and device, equipment, a storage medium and a product, and the method comprises the steps: replacing a first audio corresponding to a role in a video with a second audio in response to a selection operation of a user on the role in the video when the video is played; wherein the second audio is generated based on the voice input of the user, or is generated based on the text or audio selected by the user in combination with the timbre of the user. Therefore, in the real-time playing process of the video, the audio of the user is fused into the video in real time, so that a role in the video is avatar as a virtual image of the user, the user obtains immersive interaction experience in the video environment, and the user is converted from an external observer to a participated interactor of a video scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio and video editing, and more particularly to a method, apparatus, device, storage medium, and product for audio and video fusion. Background Technology

[0002] In related technologies, "bullet comments" refer to text information submitted by users via input devices during video playback that is time-related to the video's progress. This text information is superimposed on the video's display interface in a preset display mode (such as scrolling, fixed, or floating) and is presented when the video reaches the corresponding time point. Bullet comments have become a common way for people to express their emotions while watching videos. Existing bullet comment methods include text bullet comments, prop bullet comments, and voice bullet comments. However, these bullet comment methods always exist as supplementary information outside the video; users can only input bullet comments outside the original video to express their emotions. Summary of the Invention

[0003] In view of this, embodiments of this application provide a method, apparatus, electronic device, storage medium, and product for audio and video fusion, which aims to integrate user-generated bullet screen information with the internal environment of the video to improve user experience.

[0004] The technical solution of this application embodiment is implemented as follows: In a first aspect, embodiments of this application provide a method for audio and video fusion, applied to an audio and video playback device including a display screen, comprising: During video playback, in response to the user's selection of a character in the video, the first audio corresponding to the character in the video is replaced with a second audio; wherein the second audio is generated based on the user's voice input, or based on the text or audio selected by the user, combined with the user's voice timbre.

[0005] In the above scheme, prior to the user's selection operation on a character in the video, the method further includes: The display screen is controlled to show a first interface indicating whether to start audio and video fusion. In response to the user's trigger operation to start the audio and video fusion, the display screen is controlled to show a second interface for the user to select characters in the video.

[0006] In the above scheme, before replacing the first audio corresponding to the character in the video with the second audio, the method further includes: Extract the first text information from the first audio and the second text information from the second audio; Determine whether the text structures of the first text information and the second text information match; If a match is found, then the process of replacing the first audio corresponding to the character in the video with the second audio is performed.

[0007] In the above scheme, determining whether the text structures of the first text information and the second text information match includes: The first text information and the second text information are respectively classified into combinations of words corresponding to at least one word type; If the number of word types that are different between the first text information and the second text information is less than a first threshold, it is determined that the text structure of the first text information and the second text information are matched. If the number of word types that are different between the first text information and the second text information is greater than or equal to a first threshold, it is determined that the text structures of the first text information and the second text information do not match.

[0008] The method in the above scheme further includes: If the first audio corresponding to the character cannot be extracted, or if it is determined that the text structure of the first text information does not match that of the second text information, then the third text information is determined based on the current frame of the video. The third text information is used to describe the picture content of the current frame of the video. Determine whether the text structure of the third text information matches that of the second text information; If a match is found, the second audio is added to the video.

[0009] In the above scheme, determining whether the text structure of the third text information matches that of the second text information includes: The third text information and the second text information are respectively divided into combinations of words corresponding to at least one word type, and the weight of each word type in the third text information is determined; Based on each of the word types, determine the number of words corresponding to the word types contained in the third text information and the second text information; Based on the number of words corresponding to the word types contained in the third text information and the second text information, and the weight of each word type in the third text information, the weight of each word type in the second text information is obtained; If the sum of the weights of each word type in the second text information is greater than the second threshold, then the text structure of the third text information is determined to match that of the second text information; if the sum of the weights is less than or equal to the second threshold, then the text structure of the third text information is determined to not match that of the second text information.

[0010] In the above scheme, before replacing the first audio corresponding to the character in the video with the second audio, the method further includes: The display screen is controlled to show a third interface for acquiring the second audio.

[0011] In a second aspect, embodiments of this application provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method described in the first aspect.

[0012] Thirdly, embodiments of this application provide a storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described in the first aspect.

[0013] Fourthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.

[0014] The technical solution provided in this application is applied to an audio and video playback device including a display screen. During video playback, in response to a user's selection of a character in the video, the first audio corresponding to the character is replaced with a second audio. The second audio is generated based on the user's voice input, or based on text or audio selected by the user, combined with the user's voice timbre. Thus, during real-time video playback, the user's audio is integrated into the video in real time, transforming the characters in the video into the user's virtual avatar. This provides the user with an immersive interactive experience within the video environment, transforming the user from an external observer into a participant in the video scene. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the audio and video fusion method according to an embodiment of this application; Figure 2 This is a schematic diagram of the structure of the audio and video fusion apparatus according to an embodiment of this application; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0016] The present application will now be described in further detail with reference to the accompanying drawings and embodiments.

[0017] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.

[0018] In related technologies, "bullet comments" refer to text information submitted by users via input devices during video playback that is time-related to the video's progress. This text information is superimposed on the video's display interface in a preset display mode (such as scrolling, fixed, or floating) and is presented when the video reaches the corresponding time point. Bullet comments have become a common way for people to express their emotions while watching videos. Existing bullet comment methods include text bullet comments, prop bullet comments, and voice bullet comments. However, these bullet comment methods always exist as supplementary information outside the video; users can only input bullet comments outside the original video to express their emotions.

[0019] In various embodiments of this application, during video playback, in response to a user's selection of a character in the video, the first audio corresponding to that character is replaced with a second audio. The second audio is generated based on the user's voice input, or based on text or audio selected by the user, combined with the user's voice timbre. Thus, during real-time video playback, the user's audio is integrated into the video in real time, transforming the characters in the video into the user's virtual avatar. This provides the user with an immersive interactive experience within the video environment, transforming them from an external observer into a participant in the video scene.

[0020] This application provides a method for audio and video fusion, such as... Figure 1 As shown, the method includes: Step 101: During video playback, in response to the user's selection of a character in the video, the first audio corresponding to the character in the video is replaced with a second audio; wherein the second audio is generated based on the user's voice input, or based on the text or audio selected by the user, combined with the user's voice timbre.

[0021] For example, when a user plays a TV series using an audio / video playback device (such as a mobile phone, tablet, or computer), if a favorite actor appears in the series, and that actor interacts with other characters, the user might want to participate in the series, using the character interacting with that actor as their virtual avatar, or directly using that actor as their virtual avatar. Therefore, the user can select a character in the video on the audio / video playback device and replace the first audio corresponding to that character with the second audio the user wants that character to speak. Before replacing the first audio with the second audio, the second audio can be generated based on the user's voice input, or based on the text or audio selected by the user, combined with the user's voice tone. That is, the second audio can be the user's directly input voice (such as voice-based bullet comments), or the second audio can be generated based on the voice in a bullet comment selected by the user (but not sent by the user), using the user's own voice tone; the second audio can also be generated from the user's input text (such as text-based bullet comments), using the user's own voice tone, or the second audio can be generated based on the text in a bullet comment selected by the user (but not sent by the user), using the user's own voice tone. It should be noted that the character can be a person in the video, or an animal, a fictional character, an object, etc. This application embodiment does not limit the specific appearance of the character. The timbre, i.e., the feature parameter characterizing the user's personal timbre, can be extracted from the user's voice by inputting a segment of their own speech in advance using voice timbre feature extraction technology, and stored in the memory provided in this application embodiment. This speech can also be input by the user when audio-video fusion is required. This application does not limit the time or specific method for extracting the user's timbre. Furthermore, the specific method for generating a second audio based on the user's selected text or audio, combined with the user's timbre, can use existing technology. Specifically, it can use Text-to-Speech (TTS) technology, or it can call the API interface provided by external third-party software for generating audio, providing the user's selected text or audio, and the stored user's voice timbre features to the third-party software to obtain the audio with the user's timbre. This application embodiment does not limit this approach.

[0022] Understandably, during real-time video playback, the user's audio is integrated into the video, transforming the characters in the video into the user's virtual avatar. This provides the user with an immersive interactive experience, allowing them to transform from an external observer into a participant in the video scene.

[0023] To lower the barrier to entry for users and optimize the intuitiveness and ease of use of the audio and video fusion interaction process.

[0024] Based on this, in some embodiments, prior to the user's selection operation on a character in the video, the method further includes: The display screen is controlled to show a first interface indicating whether to start audio and video fusion. In response to the user's trigger operation to start the audio and video fusion, the display screen is controlled to show a second interface for the user to select characters in the video.

[0025] In some embodiments, before replacing the first audio corresponding to the character in the video with the second audio, the method further includes: The display screen is controlled to show a third interface for acquiring the second audio.

[0026] For example, when a user plays a TV series using an audio and video playback device, a trigger entry for audio-video fusion can be displayed on the screen when a scene suitable for audio-video fusion occurs. For instance, in a TV series, when an actor interacts with a pet character, a button (i.e., the first interface) can be displayed on the screen to instruct the user to perform audio-video fusion. After the user presses the button, a second interface is displayed on the screen for the user to select a character in the video. At this time, the user can select the pet character. After the character is selected, a third interface can be displayed on the screen to obtain the second audio. Specifically, the third interface is used to instruct the user to input or select bullet comments. Bullet comments can be text bullet comments or voice bullet comments. The user can determine the text content that the pet character is going to say by inputting bullet comments or selecting bullet comments issued by other users on the video, and then say it in the user's own voice.

[0027] It is understood that the button can be a virtual button on a touch screen or a physical button on an audio / video playback device; this application does not limit this. Furthermore, in practical applications, the order of the three steps—user input or selection of comments, user selection of a character in the video, and user pressing the button indicating audio / video fusion—can be arbitrarily combined. Similarly, the order in which the display shows the first, second, and third interfaces can also be arbitrarily combined. That is, the user can first select a character, then press the button representing audio / video fusion, and finally input comments; or they can first input comments, then press the button representing audio / video fusion, and finally select a character, and so on. This application embodiment does not limit the display order of the first, second, and third interfaces on the screen, nor the specific order of the user's operation steps.

[0028] Directly replacing a character's audio may cause the replaced audio to become disconnected from the original character's dialogue, disrupting the video's coherence and plausibility.

[0029] Based on this, in some embodiments, before replacing the first audio corresponding to the character in the video with the second audio, the method further includes: Extract the first text information from the first audio and the second text information from the second audio; Determine whether the text structures of the first text information and the second text information match; If a match is found, then the process of replacing the first audio corresponding to the character in the video with the second audio is performed.

[0030] Here, by comparing the text structure of the text content in the first and second audio recordings, it is ensured that the core content logic of the replaced second audio recording is consistent with that of the original character's audio, thus improving the naturalness and rationality of the audio replacement. The specific method for extracting text information from the audio can utilize existing technologies. Specifically, it can use a neural network model for extracting text information from audio, or it can call the API interface provided by external third-party software for speech-to-text conversion, providing the audio to the third-party software to obtain the text information. This application does not limit the specific methods for extracting the first text information from the first audio recording and the second text information from the second audio recording.

[0031] The matching of the above text structure requires explicit judgment rules to ensure the consistency of the matching results.

[0032] Based on this, in some embodiments, determining whether the text structures of the first text information and the second text information match includes: The first text information and the second text information are respectively classified into combinations of words corresponding to at least one word type; If the number of word types that are different between the first text information and the second text information is less than a first threshold, it is determined that the text structure of the first text information and the second text information are matched. If the number of word types that are different between the first text information and the second text information is greater than or equal to a first threshold, it is determined that the text structures of the first text information and the second text information do not match.

[0033] For example, the text information selected by the user in the bullet comments (second text information) and the text corresponding to the audio of the original character in the video (first text information) are first preprocessed. The preprocessing includes word segmentation, and may also include stop word removal, stemming, etc. The effect of word segmentation is to divide the text information into combinations of words corresponding to at least one word type, and then analyze the language structure of the text information (i.e., determine the word type corresponding to each word in the text information). The word types include at least coordinate phrases (hereinafter referred to as A), attributive phrases (hereinafter referred to as B), verb-object phrases (hereinafter referred to as C), complement phrases (hereinafter referred to as D), and subject-predicate phrases (hereinafter referred to as E).

[0034] Specifically, suppose the content of the first text information is "She gently wiped the pen on the table clean", and the content of the second text information is "He slowly organized the books and notebooks on the bookshelf". After preprocessing including word segmentation, the first and second text information are respectively obtained as "she, gently, wipe, clean, on the table, pen" and "he, slowly, organize, on the bookshelf, books, notebooks". The word types of the first text information include B, C, D, and E, where B includes "gently" and "on the table", C includes "wipe the pen", D includes "wipe clean", and E includes "she wipes". The word types of the second text information include A, B, C, and E, where A includes "books, notebooks", B includes "slowly", C includes "organize books", and E includes "he organizes". It can be seen that the word types in the first text information are B, C, D, and E, while the word types in the second text information are A, B, C, and E. The word types that are different between the first and second text information are A and D, with a quantity of 2. If the first threshold is 3, then the text structure of the first and second text information is determined to match.

[0035] Here, the specific content of the word type and the specific value of the first threshold can be determined according to the actual situation, and this application embodiment does not limit this.

[0036] It is understood that, in this embodiment of the application, determining whether the text structures of the first text information and the second text information match can be done not only by comparing the number of different word types in the first text information and the second text information with a first threshold, but also by comparing the number of identical word types in the first text information and the second text information with a preset threshold. Furthermore, the ratio of the number of identical word types in the first text information and the second text information to the total number of all word types in the second text information can be calculated, and the comparison of this ratio with a preset threshold can be used to determine whether the text structures of the first text information and the second text information match. This embodiment of the application does not limit the specific method for determining whether the text structures of the first text information and the second text information match.

[0037] After confirming that the text structures of the first and second text information match, when replacing the first audio corresponding to the character in the video with the second audio, the first audio should be extracted completely first.

[0038] Based on this, in some embodiments, after the user selects a character in the video, the first audio corresponding to the character can be extracted from the current frame of the video. If the current frame of the video is the beginning segment of the first audio, the timestamp t1 of the current frame is recorded, and a complete first audio sentence that the character is uttering from the current frame of the video is extracted, and the timestamp t2 of the end segment of the first audio is recorded. If the current frame of the video is the middle segment of the first audio, the video can be deduced forward to find the start segment of the first audio, record the timestamp t1 of the video frame corresponding to the start segment, and extract the first audio from t1 to extract the complete first audio, and record the timestamp t2 of the end segment of the first audio; If the current frame of the video is the end segment of the first audio, considering the smoothness of the user watching the video, the audio-video fusion operation is not allowed.

[0039] Here, in practical applications, the complete first audio can be the audio continuously spoken by a certain character in the video or a certain segment of the audio continuously spoken by this character. If there is a scene change in the video or the audio of this character is interrupted by the audio of other characters when this character is speaking, the first audio can also be a set of audio intermittently spoken by a certain character. The embodiments of the present application do not limit the specific content and manifestation form of the first audio. The start segment of the first audio is a period of time after the start of the first audio, the end segment of the first audio is a period of time before the end of the first audio, and the middle segment of the first audio is the remaining time in the first audio. The specific times of the start segment, middle segment and end segment can be determined according to the actual situation, and the embodiments of the present application do not limit this.

[0040] After extracting the complete first audio, the second audio needs to be added to the original position of the first audio.

[0041] Based on this, in some embodiments, the duration of the second audio can be adjusted according to the duration t = t2 - t1 of the complete first audio and the duration t3 of the second audio. Specifically, if t = t3, the first audio is directly replaced with the second audio, and the video after the audio replacement starts to be played.

[0042] If t < t3, the second audio is fast-forwarded at a rate of t3 / t, the first audio is replaced with the fast-forwarded second audio, and the video after the audio replacement starts to be played.

[0043] If t > t3, starting from the video timestamp t1 + (t - t3) / 2, the first audio is replaced with the second audio, and the video after the audio replacement starts to be played; or, the second audio is slow-played at a rate of t3 / t, the first audio is replaced with the slow-played second audio, and the video after the audio replacement starts to be played.

[0044] When the original character audio cannot be extracted (such as the character in the video has no lines) or the text structure corresponding to the original character audio does not match, the effective fusion of the user audio and the video cannot be achieved, which limits the scope of the interaction scenario.

[0045] Based on this, in some embodiments, the method further includes: If the first audio corresponding to the character cannot be extracted, or if it is determined that the text structure of the first text information does not match that of the second text information, then the third text information is determined based on the current frame of the video. The third text information is used to describe the picture content of the current frame of the video. Determine whether the text structure of the third text information matches that of the second text information; If a match is found, the second audio is added to the video.

[0046] Here, by capturing the current frame of the video and matching the text information (i.e., the third text information) corresponding to the content of the current frame with the second text information, the applicable scenarios for user audio can be expanded. Even if the characters in the video do not have original audio or the text structure corresponding to the original character's audio does not match, the user's required audio and video can still be reasonably integrated, thus improving the user experience.

[0047] In practical applications, existing technologies can be used to generate corresponding text information based on the content of the current frame. Specifically, the content of the image can be described in text based on a neural network model, or the API interface provided by a third-party software used to describe the image can be called to provide the image to the third-party software and obtain the text information describing the image. This application does not limit this approach.

[0048] To avoid adding audio and video content that are unrelated, the matching of the third text information with the second text information text structure requires clear judgment rules.

[0049] Based on this, in some embodiments, determining whether the text structure of the third text information matches that of the second text information includes: The third text information and the second text information are respectively divided into combinations of words corresponding to at least one word type, and the weight of each word type in the third text information is determined; Based on each of the word types, determine the number of words corresponding to the word types contained in the third text information and the second text information; Based on the number of words corresponding to the word types contained in the third text information and the second text information, and the weight of each word type in the third text information, the weight of each word type in the second text information is obtained; If the sum of the weights of each word type in the second text information is greater than the second threshold, then the text structure of the third text information is determined to match that of the second text information; if the sum of the weights is less than or equal to the second threshold, then the text structure of the third text information is determined to not match that of the second text information.

[0050] In practical application, the third text information and the second text information are first divided into combinations of words corresponding to at least one word type, and the weight of each word type in the third text information is determined. Specifically, the text structure of the third text information can be: D = [A|w1, B|w2, ​​C|w3, D|w4, E|w5]: [d1, d2, ..., dn], where d1 represents the first word in the third text information, dn represents the nth word in the third text information, A|w1 represents the weight of a coordinate phrase in the third text information, B|w2 represents the weight of a modifier-head phrase in the third text information, C|w3 represents the weight of a verb-object phrase in the third text information, D|w4 represents the weight of a complement phrase in the third text information, E|w5 represents the weight of a subject-predicate phrase in the third text information, and w1 + w2 + w3 + w4 + w5 = 100%. Then, based on the number of words corresponding to each word type contained in the third text information and the weight of each word type in the third text information, the weight of each word type in the second text information is calculated. The specific method for calculating the weight of each word type in the second text information is as follows: Parallel phrase weights: Assuming n is the number of parallel phrases in the second text information and m is the number of parallel phrases in the third text information, then the matching score of parallel phrases F(A) = (n / m) × w1, where if n>=m, then f(A) = w1; Attributive phrase weight: Assuming n is the number of attributive phrases in the second text information and m is the number of attributive phrases in the third text information, then the matching score of attributive phrase f(B) = (n / m) × w2, where if n>=m, then f(B) = w2; Verb-object phrase weight: Assuming n is the number of verb-object phrases in the second text information and m is the number of verb-object phrases in the third text information, then the matching score of verb-object phrases f(C) = (n / m) × w3, where if n>=m, then f(C) = w3; Post-complement phrase weight: Assuming n is the number of post-complement phrases in the second text information and m is the number of post-complement phrases in the third text information, then the matching score of the post-complement phrase f(D) = (n / m) × w4, where if n>=m, then f(D) = w4; Subject-predicate phrase weight: Assuming n is the number of subject-predicate phrases in the second text information and m is the number of subject-predicate phrases in the third text information, then the matching score of subject-predicate phrases f(E) = (n / m) × w5, where if n>=m, then f(E) = w5; Finally, the total weight is the weighted sum of the weights of each word type, i.e., the total weight M(D) = f(A) + f(B) + f(C) + f(D) + f(E). Based on the overall matching score M(D), the degree of matching between the second text information and the current video frame can be determined. A second threshold T is set; when M(D) > T, it is considered a match; otherwise, it is considered a mismatch.

[0051] Here, for example, suppose the content of the third text information is "The waiter is standing behind the bar, with a latte in front of him." After preprocessing, we get "Waiter, standing, behind, bar, in front, with, a, latte." Here, the modifier-head phrase B includes "a latte," the verb-object phrase C includes "with a latte in front of him" and "standing behind the bar," the complement phrase D includes "behind the bar," and the subject-verb phrase E includes "the waiter stands" and "with in front of him." It can be seen that word type A appears 0 times, word type B appears once, word type C appears twice, word type D appears once, and word type E appears twice, for a total of 6. Based on the above example, w1 is 0%, w2 is approximately 17%, w3 is 33%, w4 is 17%, and w5 is 33%. Suppose the content of the second text message is "I'm behind the bar preparing your latte." After preprocessing, we get "[I'm at, bar, behind, preparing, your latte]", where B includes "your latte", C includes "preparing the latte", D includes "behind the bar", and E includes "I'm at the bar". It's clear that word type A appears 0 times, word type B appears 1 time, word type C appears 1 time, word type D appears 1 time, and word type E appears 1 time. Therefore, the second text message... The weight f(A) of the parallel phrases in the text is 0, the weight f(B) of the modifier phrases is 1 / 1 × 17% = 17%, the weight f(C) of the verb-object phrases is 1 / 2 × 33% = 16.5%, the weight f(D) of the complement phrases is 1 / 1 × 17% = 17%, and the weight f(E) of the subject-predicate phrases is 1 / 2 × 33% = 16.5%. Therefore, the total weight M(D) is 67%. If the second threshold T is 60%, then the text structure of the third text information is considered to match that of the second text information.

[0052] The technical solution provided in this application is applied to an audio and video playback device including a display screen. During video playback, in response to a user's selection of a character in the video, the first audio corresponding to the character is replaced with a second audio. The second audio is generated based on the user's voice input, or based on text or audio selected by the user, combined with the user's voice timbre. Thus, during real-time video playback, the user's audio is integrated into the video in real time, transforming the characters in the video into the user's virtual avatar. This provides the user with an immersive interactive experience within the video environment, transforming the user from an external observer into a participant in the video scene.

[0053] In order to implement the method of the embodiments of this application, the embodiments of this application also provide an audio and video fusion apparatus, which corresponds to the above-described audio and video fusion method. The steps in the above-described audio and video fusion method embodiments are also fully applicable to the embodiments of this apparatus.

[0054] like Figure 2 As shown in the figure, this application embodiment provides an audio and video fusion device, which includes a processing module 201.

[0055] The processing module 201 is used to replace the first audio corresponding to the character in the video with a second audio in response to the user's selection operation of the character in the video during video playback; wherein the second audio is generated based on the user's voice input, or based on the text or audio selected by the user and combined with the user's timbre.

[0056] In some embodiments, the processing module 201 is further configured to: control the display screen to display a first interface for indicating whether to start audio and video fusion, and in response to a user's trigger operation to start the audio and video fusion, control the display screen to display a second interface for the user to select characters in the video.

[0057] In some embodiments, the processing module 201 is further configured to: Extract the first text information from the first audio and the second text information from the second audio; Determine whether the text structures of the first text information and the second text information match; If a match is found, then the process of replacing the first audio corresponding to the character in the video with the second audio is performed.

[0058] In some embodiments, the processing module 201 is specifically used for: The first text information and the second text information are respectively classified into combinations of words corresponding to at least one word type; If the number of word types that are different between the first text information and the second text information is less than a first threshold, it is determined that the text structure of the first text information and the second text information are matched. If the number of word types that are different between the first text information and the second text information is greater than or equal to a first threshold, it is determined that the text structures of the first text information and the second text information do not match.

[0059] In some embodiments, the processing module 201 is further configured to: If the first audio corresponding to the character cannot be extracted, or if it is determined that the text structure of the first text information does not match that of the second text information, then the third text information is determined based on the current frame of the video. The third text information is used to describe the picture content of the current frame of the video. Determine whether the text structure of the third text information matches that of the second text information; If a match is found, the second audio is added to the video.

[0060] In some embodiments, the processing module 201 is specifically used for: The third text information and the second text information are respectively divided into combinations of words corresponding to at least one word type, and the weight of each word corresponding to each word type in the third text information is determined; Based on each of the word types, determine the number of words corresponding to the word types contained in the third text information and the second text information; Based on the number of words corresponding to the word types contained in the third text information and the second text information, and the weight of each word corresponding to each word type in the third text information, the weight of each word corresponding to each word type in the second text information is obtained; If the sum of the weights of the words corresponding to each word type in the second text information is greater than the second threshold, then the text structure of the third text information is determined to match that of the second text information; if the sum of the weights is less than or equal to the second threshold, then the text structure of the third text information is determined to not match that of the second text information.

[0061] In some embodiments, the apparatus further includes a control module 202, which, before replacing the first audio corresponding to the character in the video with the second audio, is configured to: The display screen is controlled to show a third interface for acquiring the second audio.

[0062] It should be noted that the audio and video fusion apparatus provided in the above embodiments is only illustrated by the division of the above-described program modules during audio and video fusion. In practical applications, the above processing can be assigned to different program modules as needed, that is, the internal structure of the apparatus can be divided into different program modules to complete all or part of the processing described above. In addition, the audio and video fusion apparatus provided in the above embodiments belongs to the same concept as the embodiments, and its specific implementation process can be found in the method embodiments, which will not be repeated here.

[0063] Based on the hardware implementation of the above program modules, and in order to implement the audio and video fusion method of this application embodiment, this application embodiment also provides an electronic device, such as... Figure 3As shown, the electronic device 300 includes at least one processor 301, a memory 302, a user interface 303, and at least one network interface 304. The various components in the server 300 are coupled together via a bus system 305. It can be understood that the bus system 305 is used to implement communication between these components. In addition to a data bus, the bus system 305 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 3 The general designated all buses as Bus System 305.

[0064] The user interface 303 may include a monitor, keyboard, mouse, trackball, click wheel, buttons, touchpad, or touch screen.

[0065] The memory 302 in this embodiment is used to store various types of data to support the operation of the electronic device 300. Examples of such data include any computer program used to operate on the electronic device 300.

[0066] The audio and video fusion method disclosed in this application can be applied to, or implemented by, processor 301. Processor 301 may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the audio and video fusion method can be completed by integrated logic circuits in the hardware of processor 301 or by instructions in software form. The processor 301 can be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Processor 301 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the embodiments of this application can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software modules can be located in a storage medium, specifically memory 302. Processor 301 reads information from memory 302 and, in conjunction with its hardware, completes the steps provided in the embodiments of this application.

[0067] In an exemplary embodiment, the electronic device 300 may be implemented by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), FPGAs, general-purpose processors, controllers, microcontrollers (MCUs), microprocessors, or other electronic components to perform the aforementioned method.

[0068] It is understood that memory 302 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), EEPROM, ferromagnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM); magnetic surface memory can be disk storage or magnetic tape storage. Volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as Static Random Access Memory (SRAM), Synchronous Static Random Access Memory (SSRAM), Dynamic Random Access Memory (DRAM), Synchronous Dynamic Random Access Memory (SDRAM), Double Data Rate Synchronous Dynamic Random Access Memory (DDRSDRAM), Enhanced Synchronous Dynamic Random Access Memory (ESDRAM), Sync Link Dynamic Random Access Memory (SLDRAM), and Direct Rambus Random Access Memory (DRRAM). The memories described in the embodiments of this application are intended to include, but are not limited to, these and any other suitable types of memory.

[0069] In an exemplary embodiment, this application also provides a storage medium, namely a computer storage medium, specifically a computer-readable storage medium, such as a memory 302 storing a computer program. This computer program can be executed by the processor 301 of the electronic device 300 to complete the steps described in the audio and video fusion method of this application embodiment. The computer-readable storage medium can be a ROM, PROM, EPROM, EEPROM, Flash Memory, magnetic surface memory, optical disc, or CD-ROM, etc.

[0070] In an exemplary embodiment, this application also provides a computer program product, including a computer program that can be executed by a processor 301 of an electronic device 300 to perform the steps described in the method of this application embodiment.

[0071] It should be noted that terms such as "first" and "second" are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0072] Furthermore, the technical solutions described in the embodiments of this application can be combined arbitrarily without conflict.

[0073] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.

Claims

1. A method for audio and video fusion, characterized in that, Applications to audio and video playback devices including displays, including: During video playback, in response to the user's selection of a character in the video, the first audio corresponding to the character in the video is replaced with a second audio; wherein the second audio is generated based on the user's voice input, or based on the text or audio selected by the user, combined with the user's voice timbre.

2. The method according to claim 1, characterized in that, Prior to responding to a user's selection of a character in the video, the method further includes: The display screen is controlled to show a first interface indicating whether to start audio and video fusion. In response to the user's trigger operation to start the audio and video fusion, the display screen is controlled to show a second interface for the user to select characters in the video.

3. The method according to claim 1, characterized in that, Before replacing the first audio corresponding to the character in the video with the second audio, the method further includes: Extract the first text information from the first audio and the second text information from the second audio; Determine whether the text structures of the first text information and the second text information match; If a match is found, then the process of replacing the first audio corresponding to the character in the video with the second audio is performed.

4. The method according to claim 3, characterized in that, Determining whether the text structures of the first text information and the second text information match includes: The first text information and the second text information are respectively classified into combinations of words corresponding to at least one word type; If the number of word types that are different between the first text information and the second text information is less than a first threshold, it is determined that the text structure of the first text information and the second text information are matched. If the number of word types that are different between the first text information and the second text information is greater than or equal to a first threshold, it is determined that the text structures of the first text information and the second text information do not match.

5. The method according to claim 3, characterized in that, The method further includes: If the first audio corresponding to the character cannot be extracted, or if it is determined that the text structure of the first text information does not match that of the second text information, then the third text information is determined based on the current frame of the video. The third text information is used to describe the picture content of the current frame of the video. Determine whether the text structure of the third text information matches that of the second text information; If a match is found, the second audio is added to the video.

6. The method according to claim 5, characterized in that, Determining whether the text structure of the third text information matches that of the second text information includes: The third text information and the second text information are respectively divided into combinations of words corresponding to at least one word type, and the weight of each word type in the third text information is determined; Based on each of the word types, determine the number of words corresponding to the word types contained in the third text information and the second text information; Based on the number of words corresponding to the word types contained in the third text information and the second text information, and the weight of each word type in the third text information, the weight of each word type in the second text information is obtained; If the sum of the weights of each word type in the second text information is greater than the second threshold, then the text structure of the third text information is determined to match that of the second text information; if the sum of the weights is less than or equal to the second threshold, then the text structure of the third text information is determined to not match that of the second text information.

7. The method according to claim 1, characterized in that, Before replacing the first audio corresponding to the character in the video with the second audio, the method further includes: The display screen is controlled to show a third interface for acquiring the second audio.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 6.

9. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.