A method, system and terminal for target video interpretation based on video and audio synthesis
By breaking down the video into characters, scenes, and interactive layers, and matching them with rich audio layers, the problem of poor fit between commentary and content in traditional video commentary methods is solved, thereby improving the quality of the video and the audience experience.
Patent Information
- Application Number
- CN202510296898.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The existing video commentary methods have the problem that the commentary is not well matched with the video content, the audio level is not rich enough, and it is difficult to accurately present the plot and emotions, resulting in a poor viewing experience.
The target video is separated into multiple adjusted sub-videos, and decomposed into character layer, scene layer and interaction layer. The corresponding audio is matched and integrated respectively to generate sub-audio, and commentary is performed according to the sub-audio during playback.
The overall quality and expressiveness of the video have been improved, the audio and video content are more closely aligned, the audience's immersive experience and understanding are enhanced, and the viewing needs of different user groups are met.
Smart Images

Figure CN120091157B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of video and audio processing, and in particular to a method, system and terminal for target video interpretation based on video and audio synthesis. Background Art
[0002] In today's era of rapid digital information development, video, as an important means of information dissemination and entertainment, has been widely used in various fields such as education, entertainment, and advertising. As users' requirements for video content quality and viewing experience continue to increase, how to provide more high-quality, accurate, and engaging video commentary has become a key issue that needs to be addressed.
[0003] Existing video commentary methods have many limitations. For one thing, traditional video commentary often provides a unified, relatively single explanation of the entire video. This commentary is not highly consistent with the video content, making it difficult to accurately present the plot and emotions of the video. This can easily cause comprehension difficulties for viewers, resulting in a poor viewing experience.
[0004] On the other hand, traditional methods are relatively simple in terms of audio selection and use, lacking in rich audio layers. Typically, only single background music or commentary is used, making it difficult for viewers to feel immersive. This limits the overall quality and expressiveness of the video, resulting in a poor viewing experience. Summary of the Invention
[0005] In order to improve the user's viewing experience, the present application provides a target video interpretation method, system and terminal based on video and audio synthesis.
[0006] In a first aspect, the present application provides a method for target video interpretation based on video and audio synthesis, which adopts the following technical solutions:
[0007] A target video interpretation method based on video and audio synthesis, comprising:
[0008] Get the target video;
[0009] Dividing the target video into a plurality of adjusted sub-videos according to a target duration of the target video;
[0010] Decomposing each of the adjusted sub-videos into a character layer, a scene layer, and an interaction layer;
[0011] Matching a first audio according to the character layer, matching a second audio according to the scene layer, and matching a third audio according to the interaction layer;
[0012] fusing the first audio, the second audio, and the third audio to generate a sub-audio;
[0013] Associating the adjusted sub-video with the corresponding sub-audio;
[0014] When the target video is played to a certain adjusted sub-video range, a commentary is performed according to the corresponding sub-audio.
[0015] By employing the above technical solution, the target video is divided into multiple adjustment sub-videos, making editing more flexible. Each adjustment sub-video is relatively independent, allowing for personalized processing such as adding special effects and adjusting colors without affecting other parts, improving editing efficiency and accuracy. Multiple adjustment sub-videos facilitate classification and management. Adjustment sub-videos can be organized according to different themes and plots, making them easier to find and use later. Furthermore, by breaking down the adjustment sub-videos into character, scene, and interaction layers, and matching them with different audio, the audio layer is significantly enriched. The first audio layer matches the character's voice, the second audio layer matches the scene, creating a corresponding atmosphere, and the third audio layer matches the interaction layer, enhancing the interactive effect. The combined sub-audio is more vivid and three-dimensional, improving the overall quality of the video. Matching audio to different layers ensures a closer fit between the audio and the video content. The ability to precisely select appropriate audio based on the characteristics of different elements in the video better conveys the plot and emotion, enhancing the video's expressiveness. Adjusting the close association between the sub-video and the corresponding sub-audio, and providing commentary based on the sub-audio when playing to the corresponding range, can allow viewers to obtain a more coherent and natural audio-visual experience while watching the video. For example, when watching a video on a historical theme, as the video content progresses, the corresponding commentary audio sounds in a timely manner, combined with the sound that matches the characters, scenes, and interactions, making the audience feel as if they are in the scene and more easily immersed in the context created by the video. Precisely matched audio and commentary can better guide the audience to understand the video content. The commentary can explain and illustrate the key information and plot development in the video, helping the audience to better grasp the main theme and details of the video, and improve the smoothness and comprehension of the viewing. This application can flexibly adjust the combination of audio and commentary according to different types of video content. For educational videos, knowledge points can be explained in detail through commentary; for entertainment videos, commentary can increase fun and interactivity. It can meet the viewing needs of different user groups in different scenarios and improve the applicability and attractiveness of the video.
[0016] Optionally, the step of dividing the target video into a plurality of adjusted sub-videos according to the target duration of the target video includes:
[0017] Divide the target video into multiple sub-videos according to the target duration;
[0018] Performing correlation analysis on adjacent sub-videos to obtain analysis results;
[0019] According to the analysis result, adjacent sub-videos are spliced to generate an adjusted sub-video.
[0020] By employing the above technical solution, the target video is first divided into multiple sub-videos of equal length based on the target duration. This initial division is simple and direct, providing a foundation for subsequent detailed processing. It allows operations on the target video to be performed in smaller units, reducing editing complexity and allowing editors to more easily review and process each sub-video individually. Sub-video-based correlation analysis and splicing to generate adjusted sub-videos further enhances editing flexibility. Correlation analysis of adjacent sub-videos provides a deeper understanding of the inherent connections between video content. Accurate analysis is achieved by analyzing factors such as visual content, plot development, and character relationships. Splicing based on the analysis ensures the adjusted sub-videos maintain coherence and logic. After splicing, adjacent video segments transition smoothly without abrupt jumps, enabling viewers to better understand the video's message and enhancing the video's narrative impact. For example, in documentaries, the rational splicing of related sub-videos can make the overall thematic presentation clearer and more coherent.
[0021] Optionally, the step of performing correlation analysis on adjacent sub-videos to obtain analysis results includes:
[0022] Performing correlation analysis on the visual features and audio features at the same time point in the adjacent sub-videos to obtain multiple visual feature similarity values and audio feature similarity values; analyzing the adjacent sub-videos from the beginning to the end;
[0023] Screening time points whose error values of adjacent visual feature similarity values are within an error threshold to form a first time point set, and screening time points whose error values of adjacent audio feature similarity values are within the error threshold to form a second time point set;
[0024] Determining whether the first set of time points and the second set of time points each have consecutive adjacent time points;
[0025] If not, the analysis result is that the adjacent sub-videos do not need to be spliced;
[0026] If yes, if there are continuous adjacent time points in the first time point set, a first adjustment duration is generated based on the continuous adjacent time points, and a first adjusted video range is determined based on the first adjustment duration; if there are continuous adjacent time points in the second time point set, a second adjustment duration is generated based on the continuous adjacent time points, and a second adjusted video range is determined based on the second adjustment duration;
[0027] Determining whether the first adjustment video range and the second adjustment video range have overlapping adjustment ranges;
[0028] If so, the analysis result is that the adjacent sub-videos need to be split and spliced, and the splitting and splicing range is the overlapping adjustment range;
[0029] If not, the analysis result is that the adjacent sub-videos do not need to be spliced.
[0030] By employing this technical solution, the visual and audio features of adjacent sub-videos are analyzed for correlation at the same time point, and a set of eligible time points is selected based on an error threshold. This allows for precise identification of sections with strong visual and audio correlations. For example, when processing adjacent sub-videos from different scenes in a documentary, based on visual feature similarity screening, time intervals with similar colors, scene elements, and other features can be accurately identified. Furthermore, audio features such as background music and ambient sound effects can be identified, preventing abrupt visual or audio disjunctions caused by arbitrary splicing. This ensures a more logical and natural transition in the spliced video, enhancing the overall viewing experience. By simultaneously considering both visual and audio features, the system can more comprehensively assess the correlation between adjacent sub-videos compared to relying solely on a single feature to determine whether to splice. For example, some videos may appear coherent visually but have significant audio differences. If splicing is performed based solely on visual features, the resulting audio and video will be dissonant. This integrated analysis ensures that the spliced video achieves a seamless visual and audio experience, enhancing the overall video's overall cohesion.
[0031] Optionally, the step of fusing the first audio, the second audio, and the third audio to generate a sub-audio includes:
[0032] Get the first interpretation demand;
[0033] Adjusting the fusion order of the first audio, the second audio, and the third audio according to the first interpretation requirement;
[0034] According to the fusion order, the first audio, the second audio, and the third audio are input into a pre-trained word-sentence association model to generate a sub-audio.
[0035] By adopting the above technical solution, focusing on the first step of the commentary requirement allows the audio fusion process to be closely tailored to the specific requirements of the user or project. Different video scenarios and target audiences may have different expectations for audio fusion. For example, in educational videos, the explanation sound (the first audio) may be more prominent, while in entertainment videos, the coordination of ambient sound effects (the second audio) and interactive sound effects (the third audio) may be emphasized. By clarifying the commentary requirements, the audio fusion strategy can be tailored to meet diverse application scenarios and enhance the applicability of video commentary. Adjusting the fusion order based on the first commentary requirement provides great flexibility in audio fusion. Different fusion orders can result in different sub-audio effects. For example, fusion of the first audio before adding the second and third audios versus mixing the second and third audios before fusion with the first audio will produce distinct differences in the sub-audio's primary and secondary sounds and the atmosphere created. This flexible fusion order adjustment based on demand can accurately meet various commentary requirements and provide the most suitable audio support for the video. The sub-audio is generated by inputting the audio with the adjusted fusion order into a pre-trained word-sentence association model, leveraging the model's powerful learning and processing capabilities. After being trained on extensive data, the word-sentence association model understands the semantic, emotional, and acoustic relationships between audio files. This allows it to better coordinate the timbre, volume, rhythm, and other elements of different audio files during the fusion process, making the generated sub-audio more natural and harmonious. The model can intelligently mix and adjust audio based on its content and characteristics, avoiding the potential for conflicting or inharmonious sounds that can arise from simple superposition. For example, during the fusion process, the model can balance the emotional tendencies of different audio files, ensuring that the sub-audio files convey information while creating an appropriate emotional atmosphere, enhancing the quality and expressiveness of the audio.
[0036] Optionally, after generating the sub-audio, the following steps are included:
[0037] Obtain the second interpretation demand;
[0038] According to the second interpretation requirement, matching an interpretation template from an interpretation template library;
[0039] Fill the sub-audio into the commentary template to generate commentary audio.
[0040] By employing the above technical solution, a commentary template is matched from the commentary template library based on the second commentary requirement, fully utilizing a variety of pre-built commentary template resources. The commentary template library contains commentary templates with various styles, tones, and rhythms, covering a variety of common commentary scenarios. By matching templates, a basic framework that best matches the current commentary requirement can be quickly found, greatly improving the efficiency and specificity of commentary audio generation and meeting diverse customization needs. By generating commentary audio based on commentary templates, it is possible to tailor the creation to the characteristics of the sub-audio. For example, if the sub-audio has a brisk and energetic rhythm, the commentary template can guide the generation of a corresponding commentary audio with a slightly faster pace and lively tone. If the sub-audio has a quieter and soothing atmosphere, the commentary audio can also adopt a gentler and more composed style. This ensures a close match between the commentary audio and the sub-audio, providing viewers with a more harmonious and natural audio-visual experience.
[0041] Optionally, the explanation method further includes:
[0042] Obtain the target user's expression image;
[0043] Inputting the expression image into a pre-trained micro-expression analysis model to obtain the expression type;
[0044] Determining whether the expression type is negative; a negative expression type refers to a negative expression type;
[0045] If so, the commentary template and the commentary volume corresponding to the commentary audio are adjusted according to the negative level.
[0046] By employing the above technical solution, the target user's facial expression image is captured and the expression type is derived through a micro-expression analysis model, enabling precise insight into the user's emotional response while watching a video. Micro-expressions, as brief and unconscious manifestations of true human emotions, can be used to capture subtle emotional shifts in the user, understand their true feelings about the video content, and provide a basis for subsequent adjustments. When the expression type is determined to be negative, the corresponding audio commentary template and volume are adjusted based on the level of negativity. This allows the system to tailor the commentary style to the user's specific emotional state to better meet their needs. For example, if the user displays negative emotions, lowering the commentary volume can avoid further aversion, or switching to a gentler, more soothing commentary template to better match the user's current mood, thereby enhancing the user's overall viewing experience and strengthening their interaction and resonance with the video content. This mechanism, which adjusts the commentary in real time based on the user's expression, allows the video commentary to dynamically adapt to changes in the user's mood. As the user watches a video, their emotions may shift as the plot progresses. Through continuous monitoring and adjustments, the audio commentary always matches the user's current emotional state. For example, when users are confused or dissatisfied with the video content, timely adjustments to the commentary can better guide their understanding of the video, alleviate negative emotions, and establish a positive interaction between the commentary and the user's emotions. Properly adjusting the commentary template and volume can more effectively convey the video content. When users are in negative emotions, appropriate commentary can attract their attention, help them better understand the video message, and avoid missing important content due to emotional issues. At the same time, this improved matching helps to enhance the impact of video content on users and optimize the video's dissemination effect.
[0047] Optionally, after adjusting the commentary template and the commentary volume corresponding to the commentary audio according to the negative level, the method further includes:
[0048] Continuously and periodically analyzing the expressions of the target user;
[0049] Determining whether the expression type changes from negative to positive; a positive expression type is a positive expression type;
[0050] If so, adjust to the initial explanation style;
[0051] If not, change the commentary language type and add commentary effects.
[0052] By adopting the above technical solution, when it is detected that the user's expression has not turned positive, the timely change of the commentary language type and the addition of commentary special effects can help re-engage the user's interest and ensure the effective communication of information. Different commentary language types and special effects can stimulate the user's curiosity and encourage them to continue to pay attention to the video content, thereby improving the success rate of information transmission. Optimizing user emotion management: Improving user emotions by adjusting the commentary method reflects concern for user mental health. This ability to actively regulate user emotions can not only enhance the user's viewing experience, but also help users find positive outlets when facing negative emotions, which has a certain effect on promoting mental health.
[0053] In a second aspect, the present application provides a target video interpretation system based on video and audio synthesis, which adopts the following technical solutions:
[0054] A target video interpretation system based on video and audio synthesis, comprising:
[0055] Video acquisition module, used to acquire target video;
[0056] a video processing module, configured to segment the target video into a plurality of adjusted sub-videos according to a target duration of the target video, and decompose each of the adjusted sub-videos into a character layer, a scene layer, and an interaction layer;
[0057] an audio processing module, configured to match a first audio according to the character layer, match a second audio according to the scene layer, match a third audio according to the interaction layer, fuse the first audio, the second audio, and the third audio to generate a sub-audio, and associate the adjusted sub-video with the corresponding sub-audio;
[0058] The commentary module is used to provide commentary based on the corresponding sub-audio when the target video is played to a certain adjusted sub-video range.
[0059] Optionally, the video processing module includes:
[0060] a segmentation unit, configured to equally divide the target video into a plurality of sub-videos according to the target duration;
[0061] A video analysis unit, configured to perform correlation analysis on adjacent sub-videos to obtain analysis results;
[0062] The video splicing unit is used to splice adjacent sub-videos according to the analysis result to generate an adjusted sub-video.
[0063] In a third aspect, the present application provides a terminal that adopts the following technical solution:
[0064] A terminal, comprising:
[0065] A memory storing a target video explanation program based on video and audio synthesis;
[0066] The processor is used to execute the program stored in the memory to implement the steps of the target video interpretation method based on video and audio synthesis.
[0067] In summary, this application has at least the following beneficial effects:
[0068] 1. Based on the target duration of the target video, the target video is divided into multiple adjusted sub-videos. Each adjusted sub-video is then decomposed into a character layer, a scene layer, and an interactive layer. A first audio track is matched to the character layer, a second audio track is matched to the scene layer, and a third audio track is matched to the interactive layer. The first, second, and third audio tracks are then fused to generate a sub-audio track, thereby associating the adjusted sub-video with the corresponding sub-audio track. When the target video plays to a certain adjusted sub-video range, the narration based on the corresponding sub-audio track is used to personalize the sub-video. The first audio track matched to the character layer highlights the character's voice characteristics, the second audio track matched to the scene layer creates a corresponding ambient atmosphere, and the third audio track matched to the interactive layer enhances the interactive effect. The fused sub-audio tracks are more vivid and three-dimensional, improving the overall quality of the video. Matching the audio tracks to different layers ensures a better fit between the audio and video content, allowing for flexible adjustment of the audio and narration pairing based on different video content types. For educational videos, the narration can explain the key points in detail; for entertainment videos, the narration can add interest and interactivity. This can meet the viewing needs of different user groups in different scenarios, improving the applicability and appeal of the video.
[0069] 2. The target user's facial expression image is captured and fed into the micro-expression analysis model to determine the expression type. The model then determines whether the expression type is negative. If so, the corresponding audio commentary template and volume are adjusted based on the degree of negativity. This approach aims to accurately understand the user's emotional response while watching the video. Micro-expressions, as fleeting and unconscious manifestations of true human emotions, can capture subtle emotional shifts in the user, understanding their true feelings about the video content and providing a basis for subsequent adjustments. If the expression type is determined to be negative, the corresponding audio commentary template and volume are adjusted based on the degree of negativity. This allows the system to tailor the commentary style to the user's specific emotional state to better meet their needs. For example, if the user displays negative emotion, lowering the commentary volume can avoid further aversion; or switching to a gentler, more soothing commentary template can better match the user's current mood, thereby enhancing the user's overall viewing experience and strengthening their interaction and resonance with the video content. This mechanism, which adjusts the commentary in real time based on the user's expression, allows the video commentary to dynamically adapt to changing user emotions. As users watch videos, their emotions may shift as the plot unfolds. Through continuous monitoring and adjustment, the audio commentary can always align with the user's current emotional state. For example, when a user is confused or dissatisfied with the video content, timely adjustments to the commentary can better guide their understanding of the video, alleviate negative emotions, and establish a positive interaction between the commentary and the user's emotions. Proper adjustments to the commentary template and volume can more effectively convey the video's content. When users are in a negative mood, appropriate commentary can capture their attention, help them better understand the video message, and avoid missing important content due to emotional issues. At the same time, this improved matching helps to enhance the influence of video content on users and optimize the video's dissemination effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 This is a flowchart of an implementation method of Example 1 of the present application;
[0071] Figure 2 It is a flowchart of the specific steps of S120;
[0072] Figure 3 It is a flowchart of the specific steps of S122;
[0073] Figure 4 It is a flowchart of the specific steps of S150;
[0074] Figure 5 This is a flowchart of another embodiment of the method of the present application;
[0075] Figure 6It is a structural block diagram of an embodiment of the system of the present application. DETAILED DESCRIPTION
[0076] To make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the following will be combined with the appended drawings of the embodiments of the present invention. Figure 1 -Attached Figure 6 The technical solutions in the embodiments of the present invention are clearly and completely described. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.
[0077] The first embodiment of the present application discloses a method for explaining a target video based on video and audio synthesis. Figure 1 As an implementation of the target video interpretation method of the audio-visual synthesis, the target video interpretation method may include S110-S170:
[0078] S110, acquiring a target video;
[0079] S120, dividing the target video into a plurality of adjusted sub-videos according to the target duration of the target video;
[0080] S130, decomposing each adjusted sub-video into a character layer, a scene layer, and an interaction layer;
[0081] S140, matching the first audio according to the character layer, matching the second audio according to the scene layer, and matching the third audio according to the interaction layer;
[0082] S150, fusing the first audio, the second audio, and the third audio to generate a sub-audio;
[0083] S160, associating the adjusted sub-video with the corresponding sub-audio;
[0084] S170: When the target video reaches a certain adjusted sub-video range, a commentary is performed according to the corresponding sub-audio.
[0085] Specifically, the target video can be obtained in any manner, including manual input or downloading from a third-party platform. The target video can be muted or the original video. After obtaining the target video, the target duration (i.e., the playback duration) of the target video is calculated. The target video is then divided into multiple adjusted sub-videos based on the target duration. Each adjusted sub-video is then decomposed into a character layer, a scene layer, and an interaction layer. The character layer includes relevant features such as the image of anthropomorphic features such as people and animals in the adjusted sub-video; the scene layer includes the background of the adjusted sub-video, such as location, time, and weather; and the interaction layer includes the character's actions and various forms of interaction between characters, such as dialogue, eye contact, physical contact, and conflict. Deep learning image recognition models, such as object detection models combined with image classification models, can be used to distinguish between the character and scene layers. Natural language processing models, such as sentiment analysis models combined with semantic understanding models, can be used for character-related analysis in the interaction and character layers. Alternatively, video understanding models, such as 3D convolutional neural network models, can be used to directly distinguish between the layers.
[0086] For example, consider an animated short video about several small animals searching for treasure in a forest. For the character layer, the adjusted sub-video is first imported into a tool that supports video frame extraction. Frame by frame, the images are then input into a YOLOv5 model for detection. The model accurately identifies the small animal characters in each frame, such as a bunny, monkey, or bird, and accurately identifies their specific positions and approximate size ranges. By statistically analyzing the detection results across all frames, we can clearly determine the frequency of each character's appearance in the video. The bunny appears the most frequently, throughout the entire video, suggesting it may be the main character. The bird, on the other hand, appears only sporadically and is likely a supporting role. Next, an image classification model (such as ResNet) is used to refine the classification of the identified small animal characters. The model identifies the rabbit's appearance as white fur and long ears, while the monkey's appearance as brown fur and long tail, allowing for clearer visual distinction between the different characters. At the same time, we extract the voice dialogue content of the characters in the video (if any), and use the semantic understanding model in natural language processing (such as BERT) to analyze the semantics of the characters' speech. For example, the little rabbit always mentions the need to move forward bravely and not give up. From this, we can infer that the little rabbit has the character traits of bravery and perseverance, further improving our understanding of the character.
[0087] At the scene level, the object detection model detects various objects within the forest scene, such as tall trees, colorful flowers, and rocks hidden in the grass. This object information helps us determine the scene's richness. The image classification model further identifies the overall scene category, clearly identifying it as an outdoor forest scene. By analyzing the image's tones and lighting (for example, the overall greenish hue and the dappled light and shadows created by sunlight filtering through leaves indicate a sunny daytime forest environment), it clearly depicts the scene's atmosphere as fresh and vibrant. Using video understanding models (such as C3D) to track scene changes over time, we find that as the animals continue their journey, the scene shifts from an open meadow to a denser, more rugged area with denser trees. By capturing this scene change, the model accurately identifies the scene transition point, determining that the transition occurs approximately halfway through the video. Furthermore, we analyze that the scene transition occurs naturally through a pan, which aligns with the animals' path in their treasure hunt and maintains the video's plot continuity.
[0088] At the interaction layer, the text of the animals' conversations is extracted and fed into a sentiment analysis model (such as one based on the Transformer architecture) to analyze their emotional tendencies during the exchange. For example, when the monkey suggests a break but the rabbit encourages everyone to keep going, the sentiment analysis model determines that the rabbit's words are positive and resolute, while the monkey appears somewhat helpless but still complies. This reflects the emotional states and personality differences between the two characters during the interaction. The semantic understanding model further analyzes the conversation content, clarifying that the purpose of the interaction is to discuss the next steps and that the rabbit encourages everyone to keep going, hoping to find the treasure as quickly as possible. This helps to identify the behavioral motivations behind the interaction. The video understanding model (C3D) observes the animals' interactions with the forest scene, showing the rabbit using branches to cross a stream and the monkey using vines to swing over an obstacle. These characters' interactions with the scene elements are fully captured, and the analysis shows that their purpose is to use objects in the scene to overcome difficulties in the treasure hunt. This also creates a fun and challenging atmosphere, making the video more enjoyable. Moreover, from the perspective of the entire video time period, the rhythm of interaction between the animals and the scene is from slow to fast. As they get closer to the treasure, the interaction becomes more frequent and intense. The model clearly grasps the dynamic changes of this interaction and its role in advancing the plot of the video.
[0089] Reference Figure 2 Specifically, S120 may include S121-S123:
[0090] S121, dividing the target video into multiple sub-videos according to the target duration;
[0091] S122, performing correlation analysis on adjacent sub-videos to obtain analysis results;
[0092] S123: Based on the analysis result, adjacent sub-videos are spliced to generate an adjusted sub-video.
[0093] Reference Figure 3 Specifically, S122 may include S1221-S1228:
[0094] S1221, performing correlation analysis on the visual features and audio features at the same time point in adjacent sub-videos to obtain multiple visual feature similarity values and audio feature similarity values; adjacent sub-videos are analyzed from the beginning to the end;
[0095] S1222, filtering time points where the error values of adjacent visual feature similarity values are within the error threshold to form a first time point set, and filtering time points where the error values of adjacent audio feature similarity values are within the error threshold to form a second time point set;
[0096] S1223, determining whether both the first time point set and the second time point set have consecutive adjacent time points;
[0097] S1224, if not, the analysis result is that the current adjacent sub-videos do not need to be spliced;
[0098] S1225: If yes, if there are consecutive adjacent time points in the first time point set, a first adjustment duration is generated based on the consecutive adjacent time points, and a first adjustment video range is determined based on the first adjustment duration; if there are consecutive adjacent time points in the second time point set, a second adjustment duration is generated based on the consecutive adjacent time points, and a second adjustment video range is determined based on the second adjustment duration;
[0099] S1226, determining whether the first adjustment video range and the second adjustment video range have overlapping adjustment ranges;
[0100] S1227: If yes, the analysis result is that the current adjacent sub-videos need to be split and spliced, and the split and splicing range is the overlap adjustment range.
[0101] S1228: If not, the analysis result is that the current adjacent sub-videos do not need to be spliced.
[0102] For example, consider a video of a campus cultural performance that has been split into multiple sub-videos of equal length. Consider two adjacent sub-videos: Sub-video A, which captures the final portion of a student dance performance and lasts 2 minutes (120 seconds); and Sub-video B, which captures the host's entrance after the dance performance and the beginning of the next program, also lasting 2 minutes (120 seconds). Set the error threshold to 0.2 (the similarity value ranges from 0 to 1 and can be adjusted based on actual conditions).
[0103] First, we analyze the correlation of visual and audio features between the two sub-videos. Regarding visual features, we extract visual features from the ending portion of sub-video A (the first 30 seconds are a key focus, with time points counting forward from the end, such as the first second corresponding to the actual 120th second, the second second to the 119th second, and so on) and the beginning portion of sub-video B (normally counting from the beginning). By calculating the color histogram and comparing the image colors, we can see that the ending of sub-video A is bright overall, with the stage lighting effects of a dance performance and the students wearing colorful costumes. The beginning of sub-video B also appears bright due to the stage lighting, with the background depicting the stage setting, the host standing on the stage in a stylish outfit, and props for the next program nearby, creating an equally rich color palette. Looking at key elements, at the end of sub-video A, the dancers perform their final dance moves and pose for their curtain call, while at the beginning of sub-video B, the host walks to the center stage, while staff prepare props and adjust instruments backstage. The structural similarity algorithm SSIM is used to calculate that the visual feature similarity value corresponding to the 1st second at the end of A (actually the 120th second) and the 1st second at the beginning of B is 0.3. In this way, the visual feature similarity value of each corresponding time point is calculated in turn.
[0104] For audio features, we extracted audio information at each time point at the end of sub-video A and the beginning of sub-video B, analyzing the audio spectrum, volume, and primary sound types. The background music at the end of sub-video A is a lively dance score, interspersed with the gentle footsteps of students and the applause of the audience. The background music at the beginning of sub-video B slows down but maintains a cheerful atmosphere, with the host's voice and the sounds of props being prepared backstage. Using audio feature comparison algorithms such as Mel-Frequency Cepstral Coefficients (MFCCs), we calculated a similarity of 0.5 between the audio features at the end of video A (actually the 116th second) and the beginning of video B (8th second). Similar similarities were calculated for the audio features at each other time point.
[0105] Next, we filter the time points that meet the error threshold. For visual features, we iterate through all calculated visual feature similarity values and select time points where the error between adjacent similarity values falls within the 0.2 error threshold. For example, we find that the time points from the 8th second at the end of A (actually the 113th second) to the 8th second at the beginning of B meet the requirements and form the first time point set. We perform the same screening for audio feature similarity values and find that the time points from the 6th second at the end of A (actually the 115th second) to the 6th second at the beginning of B meet the requirements and form the second time point set.
[0106] Looking at these two sets, we find that in the first time point set (visual features), the time points from the 8th second at the end of A (actually the 113th second) to the 8th second at the beginning of B are continuous and adjacent time points. In the second time point set (audio features), the time points from the 6th second at the end of A (actually the 115th second) to the 6th second at the beginning of B are also continuous and adjacent time points, satisfying the condition that both sets have continuous and adjacent time points.
[0107] Then, based on the consecutive adjacent time points, an adjustment duration is generated and the adjusted video range is determined. For visual features, the first adjustment duration is generated based on the consecutive adjacent time points from the 8th second at the end of A (actually the 113th second) to the 8th second at the beginning of B, which is 26 seconds. Therefore, the first adjusted video range is the video segment from the 113th second onward in sub-video A and from the beginning to the 8th second in sub-video B. For audio features, the second adjustment duration is generated based on the consecutive adjacent time points from the 6th second at the end of A (actually the 115th second) to the 6th second at the beginning of B, which is 12 seconds. The corresponding second adjusted video range is the video portion from the 115th second onward in sub-video A and from the beginning to the 6th second in sub-video B.
[0108] Comparing the two adjustment video ranges, we can find that the interval from the 115th second of sub-video A to the 6th second of sub-video B is included in both ranges, so there is an overlapping adjustment range.
[0109] Therefore, the analysis result shows that the adjacent sub-videos A and B need to be split and spliced. The splitting and splicing range is the overlapping video clips corresponding to the interval from the 115th second of sub-video A to the 6th second of sub-video B. Subsequently, the videos in this range can be spliced using a fade-in and fade-out transition effect to make the transition between the two sub-videos more natural and smooth, and better show the complete scene process of the campus art performance.
[0110] Specifically, for S123 , the adjusted durations of adjacent sub-videos are compared respectively, the sub-video with the longer adjusted duration is used as the splicing object, and the other sub-video is used as the segmentation object; the spliced and segmented video is the adjusted sub-video.
[0111] For example, taking the above example, the first 6 seconds of sub-video B can be spliced to the end of sub-video A; this is equivalent to 6 seconds more of sub-video A and 6 seconds less of sub-video B.
[0112] Specifically for S140: the first audio can be the keywords representing the character layer, such as the character name and character personality, extracted after the feature analysis of the character layer; the second audio can be the keywords representing the scene layer, such as the weather and venue, extracted after the feature analysis of the scene layer; the third audio can be the keywords representing the interaction layer, such as the character's actions and the interaction between characters, extracted after the feature analysis of the interaction layer.
[0113] Reference Figure 4 Specifically, S150 includes S151-S156:
[0114] S151, obtaining the first interpretation requirement;
[0115] S152, adjusting the fusion order of the first audio, the second audio, and the third audio according to the first interpretation requirement;
[0116] S153, inputting the first audio, the second audio, and the third audio into a pre-trained word-sentence association model according to the fusion order to generate a sub-audio;
[0117] S154, obtaining a second interpretation requirement;
[0118] S155, matching an explanation template from an explanation template library according to the second explanation requirement;
[0119] S156: Fill the sub-audio into the commentary template to generate the commentary audio.
[0120] Specifically, for example, a video explaining the theme of a campus sports meeting includes multiple competition events and interactive scenes between different roles such as athletes and spectators. The keywords extracted to represent the role layer include "athlete", "competitive", "referee", "fair and serious", "audience", "enthusiastic cheering", etc. Keywords such as "track and field stadium", "sunny", "football field", "stands", "lively", and "cloudy weather" are extracted from the scene layer. Keywords such as "chasing each other", "handing over the baton", "whistle", "cheer", and "discussion and sharing" are extracted from the interaction layer. The first commentary requirement is to vividly show the enthusiastic atmosphere of the campus sports meeting through audio, first let the audience feel the characteristics of different scenes, then highlight the interaction between the characters, and finally strengthen the unique performance of each character in the sports meeting, so that the entire audio and video images are coordinated, making the audience feel as if they are at the lively sports meeting. According to the commentary requirements, the fusion order is determined as follows: first is the second audio (based on scene layer keywords), which uses audio elements corresponding to keywords such as venue and weather to create the environmental atmosphere of different competition venues, allowing the audience to quickly immerse themselves in the scene of the campus sports meeting; followed by the third audio (based on interaction layer keywords), which reflects the energetic and passionate interaction in the sports meeting by showing the audio corresponding to various interactive behaviors between characters; finally, the first audio (based on role layer keywords) highlights the characteristics and status of different roles such as athletes, referees, and spectators, further strengthening the theme atmosphere of the sports meeting.
[0121] The word-sentence association model is trained based on a large number of audio and text descriptions related to campus activities. It can associate a combination of audio clips that are adapted and coherent based on the input keywords and the set fusion order, and generate sub-audio that is consistent with the theme of the sports meeting.
[0122] First, input the keywords of the second audio (such as "track and field stadium", "sunny", "on the stands", etc.). The model will associate the opening with a clear broadcast announcing the start of the sports meeting, followed by the noise of the crowd on the track and field stadium and the sound of the breeze blowing, accompanied by background music with a light rhythm and bright melody, to create an atmosphere of a lively opening of the track and field stadium under the bright sunshine. The duration is about the first 30 seconds, which serves as the starting prelude of the entire audio.
[0123] Then, the keywords of the third audio ("chasing each other", "passing the baton", "cheers", etc.) were input, and the model associated a segment of athletes chasing each other on the track, with rapidly alternating footsteps, heavier breathing, and the crisp touch of the baton at the moment of handing over. At the same time, the shouts and cheers of the surrounding audiences rose one after another, and the background music became faster and more passionate. The duration is about 25 seconds, naturally continuing the atmosphere of the previous scene and showing the wonderful interactive moments in the sports meeting.
[0124] Finally, we input the keywords of the first audio (such as "athletes," "competitive," "referees," and "fair and serious"). The model then generates the sounds of athletes cheering excitedly after crossing the finish line, the referee announcing the results of the game in a serious and steady voice, and the audience enthusiastically discussing the results. This is paired with background music that is slightly slower but still energetic and celebratory, with a duration of approximately 20 seconds. This highlights the status of different characters in the sports meet and forms a complete sub-audio that conforms to the fusion order.
[0125] The second commentary requirement is to use passionate and easy-to-understand language to provide a detailed introduction to the various competition events, the characters' performances in them, and the highlights of the interactive moments. The tone should also change accordingly according to different scenes and character characteristics. For example, when introducing a tense competition, the speed of speech should be accelerated and the tone of voice should be raised. When describing a relaxing moment, the speed of speech should be slowed down and the tone should be more relaxed and cheerful.
[0126] The commentary template library includes numerous templates suitable for campus events, covering diverse themes such as sports meets, theatrical performances, and campus competitions. Each template features corresponding language structures, common expressions, and suggested tone adjustments tailored to different emotional atmospheres. Based on the theme of the campus sports meet and the needs of the secondary commentary, a suitable commentary template can be found. The template typically begins with a general description of the overall atmosphere and current scene, such as "Today, the campus welcomes a grand sports meet. The atmosphere of tension and excitement permeates the stadiums." The middle section details specific events, such as "On the track and field field, the 100-meter sprint is in full swing. Each athlete, like an arrow from a bow, dashes towards the finish line." The conclusion summarizes the highlights of the event or provides a brief outlook on upcoming competitions, such as "This exciting 100-meter race is just one highlight of the sports meet. There are many more events waiting for you. Let's look forward to it!" Furthermore, corresponding tone inflections are provided for different character interactions and competition progress, providing guidance on how to best convey the appropriate atmosphere through tone. Fill in the commentary based on the structure and prompts of the commentary template, combined with the specific content corresponding to the sub-audio. For example, in the section corresponding to the athletes chasing each other on the track in the previous sub-audio, the commentary audio will say at a fast pace and in an excited tone: "Look! The athletes are engaged in a fierce competition on the 100-meter track. They chase each other and refuse to give in. Every step is full of power, and every arm swing demonstrates the desire for victory. This is the charm of the sports meeting!" At the same time, according to the duration and rhythm of the sub-audio, the commentary sentence speed and pauses are reasonably adjusted to ensure that the commentary audio and sub-audio are synchronized in time and match the emotional atmosphere.
[0127] For other parts of the entire sub-audio, such as the opening introduction of the sports meeting and the subsequent scenes after the end of the game, the corresponding commentary content is also generated in sequence according to the commentary template, and the text content is converted into commentary audio through speech synthesis technology (selecting a speech synthesis engine with clear and contagious tone), which can not only clearly convey the information of the sports meeting but also fully create a warm atmosphere.
[0128] Reference Figure 5 As another embodiment of the target video interpretation method, the interpretation method may further include S210-S250:
[0129] S210, obtaining an expression image of a target user;
[0130] S220, inputting the expression image into a pre-trained micro-expression analysis model to obtain the expression type;
[0131] S230, determining whether the expression type is negative;
[0132] S240, if yes, adjusting the commentary template and the commentary volume corresponding to the commentary audio according to the negative level;
[0133] S250, if not, no response.
[0134] Specifically, negative emotions include sadness, grief, depression, etc. The negativity level refers to an artificially set emotion level, for example, depression is level one, sadness is level two, and sadness is level three.
[0135] For example, while a target user is watching a video, a device (such as a camera mounted on a smart TV or a computer) captures their face at regular intervals (e.g., every 30 seconds) to obtain facial expressions. These images clearly show details such as the user's facial features and changes in facial muscles, providing basic data for subsequent micro-expression analysis.
[0136] The micro-expression analysis model is trained on a large dataset of facial images with labeled expression types. It can accurately identify different types of expressions, such as happiness, surprise, sadness, grief, and depression. The model extracts features from the input user expression images, analyzing subtle facial muscle movements and changes in facial features. It then matches these features with learned expression patterns and outputs the corresponding expression type.
[0137] For example, a user's facial expression image captured on one occasion showed a slightly furrowed brow, a drooping mouth, and a slightly dull look in the eyes. After feeding this image into the micro-expression analysis model, the model analyzed and determined the expression type as "sad."
[0138] The original commentary template may have a relatively plain and objective language style, suitable for explaining video content in a normal emotional state. Now, based on the situation where the negativity level reaches level two, the commentary template is adjusted. For example, some descriptive sentences are made softer and more soothing, and more empathetic words are added. The statement "In this movie, the protagonist encounters a difficult problem and needs to find a solution" is changed to "Alas, the protagonist in this movie is facing such a difficult problem at this moment. I wonder if he can successfully find a solution." Through this linguistic adjustment, the commentary audio is more in line with the user's current slightly depressed emotional state. At the same time, the commentary volume is also adjusted accordingly. Generally speaking, as the negativity level increases, the commentary volume will be appropriately lowered to create a relatively quiet and soothing atmosphere, avoiding excessive noise that exacerbates the user's negative emotions.
[0139] For the "sad" situation with a negative level of level two, the commentary volume is reduced from the original moderate level to about 80% of the original volume, so that the commentary audio is played in a relatively soft manner, making the user feel more comfortable.
[0140] If the expression type output by the micro-expression analysis model is positive (such as happiness, surprise, etc.) or neutral (such as calmness, etc.), then the commentary audio will continue to play according to the normal mode originally set by the system. No additional adjustments will be made to the commentary template and commentary volume, and the original playback status can be maintained.
[0141] In addition, after adjusting the commentary template and volume corresponding to the commentary audio according to the negative level, the target user's expression is continuously and periodically analyzed to further determine whether the target user's expression type has changed to positive. If so, it is adjusted to the initial commentary style; if not, the commentary voice type is changed and commentary special effects are added.
[0142] For example, after adjusting the commentary audio, the system continues to obtain the target user's facial expression image every 20 seconds and inputs it into the micro-expression analysis model for analysis; after a period of observation, if the model analysis finds that the user's mouth corners are raised and the eyes are bright and energetic, and the expression type is judged to have changed to "happy" (positive expression), the system will adjust the commentary audio back to the initial commentary style.
[0143] If analysis reveals that the user's expression remains negative, the original commentary voice, speaking standard Mandarin and speaking young male voice, is replaced with a gentle, female voice that is softer and more approachable, attempting to alleviate the user's negative emotions through different vocal qualities. Soothing background music, such as gentle piano, is also added to the commentary audio to blend in with the voice. At the same time, the commentary speed is appropriately slowed down to make it more calm and peaceful. These adjustments help users gradually relax and improve their emotional state.
[0144] Based on the above method embodiment, the second embodiment of the present application discloses a target video interpretation system based on video and audio synthesis. Figure 6 As an embodiment of the target video interpretation system, the target video interpretation system may include:
[0145] Video acquisition module, used to acquire target video;
[0146] A video processing module, configured to segment a target video into a plurality of adjusted sub-videos according to a target duration of the target video, and decompose each adjusted sub-video into a character layer, a scene layer, and an interaction layer;
[0147] an audio processing module, configured to match a first audio according to a character layer, match a second audio according to a scene layer, match a third audio according to an interaction layer, fuse the first audio, the second audio, and the third audio to generate a sub-audio, and associate the adjusted sub-video with the corresponding sub-audio;
[0148] The commentary module is used to provide commentary based on the corresponding sub-audio when the target video is played to a certain adjusted sub-video range.
[0149] The video processing module includes:
[0150] A segmentation unit, used to divide the target video into multiple sub-videos according to the target duration;
[0151] A video analysis unit, configured to perform correlation analysis on adjacent sub-videos to obtain analysis results;
[0152] The video splicing unit is used to splice adjacent sub-videos according to the analysis result to generate an adjusted sub-video.
[0153] The modules of the target video interpretation system based on video and audio synthesis correspond to the target video interpretation method based on video and audio synthesis, and will not be described in detail here.
[0154] The third embodiment of the present application provides a terminal. As an implementation of the terminal, the terminal may include: a memory and a processor; wherein,
[0155] The memory is used for storing a target video explanation program based on video and audio synthesis;
[0156] The processor is used to execute the program stored in the memory to implement the steps of the target video interpretation method based on video and audio synthesis.
[0157] The memory may be communicatively connected to the processor via a communication bus, and the communication bus may be an address bus, a data bus, a control bus, or the like.
[0158] In addition, the memory may include a random access memory (RAM) and may also include a non-volatile memory (NVM), such as at least one disk storage.
[0159] The processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.
[0160] The above are all preferred embodiments of the present application and are not intended to limit the scope of protection of the present application. Unless otherwise specified, any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features. In other words, unless otherwise specified, each feature is merely an example of a series of equivalent or similar features.
Claims
1. A target video interpretation method based on video and audio synthesis, characterized in that: include: Get the target video; Dividing the target video into a plurality of adjusted sub-videos according to a target duration of the target video; Decomposing each of the adjusted sub-videos into a character layer, a scene layer, and an interaction layer; Matching a first audio according to the character layer, matching a second audio according to the scene layer, and matching a third audio according to the interaction layer; fusing the first audio, the second audio, and the third audio to generate a sub-audio; Associating the adjusted sub-video with the corresponding sub-audio; When the target video is played to a certain adjusted sub-video range, a commentary is performed according to the corresponding sub-audio.
2. The target video interpretation method based on video and audio synthesis according to claim 1, characterized in that: The step of dividing the target video into a plurality of adjusted sub-videos according to the target duration of the target video includes: Divide the target video into multiple sub-videos according to the target duration; Performing correlation analysis on adjacent sub-videos to obtain analysis results; According to the analysis result, adjacent sub-videos are spliced to generate an adjusted sub-video.
3. The target video interpretation method based on video and audio synthesis according to claim 2, characterized in that: The step of performing correlation analysis on adjacent sub-videos to obtain analysis results includes: Performing correlation analysis on the visual features and audio features at the same time point in the adjacent sub-videos to obtain multiple visual feature similarity values and audio feature similarity values; analyzing the adjacent sub-videos from the beginning to the end; Screening time points whose error values of adjacent visual feature similarity values are within an error threshold to form a first time point set, and screening time points whose error values of adjacent audio feature similarity values are within the error threshold to form a second time point set; Determining whether there are consecutive adjacent time points in both the first time point set and the second time point set; If not, the analysis result is that the adjacent sub-videos do not need to be spliced; If yes, if there are continuous adjacent time points in the first time point set, a first adjustment duration is generated based on the continuous adjacent time points, and a first adjusted video range is determined based on the first adjustment duration; if there are continuous adjacent time points in the second time point set, a second adjustment duration is generated based on the continuous adjacent time points, and a second adjusted video range is determined based on the second adjustment duration; Determining whether the first adjustment video range and the second adjustment video range have overlapping adjustment ranges; If so, the analysis result is that the adjacent sub-videos need to be split and spliced, and the splitting and splicing range is the overlapping adjustment range; If not, the analysis result is that the adjacent sub-videos do not need to be spliced.
4. The method for target video interpretation based on video and audio synthesis according to claim 1, characterized in that: The step of fusing the first audio, the second audio, and the third audio to generate a sub-audio includes: Get the first interpretation demand; Adjusting the fusion order of the first audio, the second audio, and the third audio according to the first interpretation requirement; According to the fusion order, the first audio, the second audio, and the third audio are input into a pre-trained word-sentence association model to generate a sub-audio.
5. The method for target video interpretation based on video and audio synthesis according to claim 4, characterized in that: After the sub-audio is generated, the method includes: Obtain the second interpretation demand; According to the second interpretation requirement, matching an interpretation template from an interpretation template library; Fill the sub-audio into the commentary template to generate commentary audio.
6. The method for target video interpretation based on video and audio synthesis according to claim 1, characterized in that: The explanation method further includes: Obtain the target user's expression image; Inputting the expression image into a pre-trained micro-expression analysis model to obtain the expression type; Determining whether the expression type is negative; a negative expression type refers to a negative expression type; If so, the commentary template and the commentary volume corresponding to the commentary audio are adjusted according to the negative level.
7. The method for target video interpretation based on video and audio synthesis according to claim 6, characterized in that: After adjusting the commentary template and the commentary volume corresponding to the commentary audio according to the negative level, the method further includes: Continuously and periodically analyzing the expressions of the target user; Determining whether the expression type changes from negative to positive; a positive expression type is a positive expression type; If so, adjust to the initial explanation style; If not, change the commentary language type and add commentary effects.
8. A target video explanation system based on video and audio synthesis, characterized in that: include: Video acquisition module, used to acquire target video; a video processing module, configured to segment the target video into a plurality of adjusted sub-videos according to a target duration of the target video, and decompose each of the adjusted sub-videos into a character layer, a scene layer, and an interaction layer; an audio processing module, configured to match a first audio according to the character layer, match a second audio according to the scene layer, match a third audio according to the interaction layer, fuse the first audio, the second audio, and the third audio to generate a sub-audio, and associate the adjusted sub-video with the corresponding sub-audio; The explanation module is used to explain the target video according to the corresponding sub-audio when the target video is played to a certain adjusted sub-video range.
9. The target video interpretation system based on video and audio synthesis according to claim 8, characterized in that: The video processing module includes: a segmentation unit, configured to equally divide the target video into a plurality of sub-videos according to the target duration; A video analysis unit, configured to perform correlation analysis on adjacent sub-videos to obtain analysis results; The video splicing unit is used to splice adjacent sub-videos according to the analysis result to generate an adjusted sub-video.
10. A terminal, characterized in that: include: A memory storing a target video explanation program based on video and audio synthesis; A processor is configured to execute the program stored in the memory to implement the steps of the target video interpretation method based on video and audio synthesis as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Multimedia comment system and multimedia comment method
US20140089800A1
Hierarchical segmentation and quality measurement for video editing
US20160358628A1