Short drama translation and explanation method based on multi-mode neural network

Through multimodal neural network technology, the full process of short drama translation and explanation is realized, the accuracy of subtitle recognition and cross-language translation is solved, cultural adaptability is enhanced, and the global communication effect of multimedia content is improved.

CN120302128APending Publication Date: 2025-07-11CHENGDU ROAD YOUYOU TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510480689.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

In the production of multimedia content, especially in the process of short drama translation and commentary, the prior art cannot efficiently and accurately combine multiple information sources to generate cultural adaptive translations that conform to the target language habits and convey original emotional colors, resulting in inaccurate subtitle recognition and poor cross-language translation, affecting the efficiency of content dissemination on a global scale.

Method used

The multimodal neural network is used to extract subtitle text through dynamic peak detection and IQR method, combine optical flow field analysis and scene segmentation to remove subtitle traces, and use LSTM network to analyze plot conflicts, and generate cultural adaptive translation through multimodal neural network fusion across language semantics, and finally perform dubbing synthesis and subtitle suppression.

Benefits of technology

It realizes the full process automation of subtitle extraction to cultural adaptation translation, improves the accuracy of subtitle extraction and the cultural adaptability of translated content, improves the understanding and acceptance of multimedia content under different cultural backgrounds, and greatly improves the viewing experience of global users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302128A_ABST
    Figure CN120302128A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of multimedia, and particularly relates to a short drama translation and explanation method based on a multi-mode neural network, which realizes full-process automatic processing from subtitle extraction to culture adaptation translation by integrating advanced technical means such as an OCR (optical character recognition) model, optical flow field analysis, an LSTM (long short term memory) network and a graph neural network. Particularly, the subtitles in a complex background can be accurately recognized and processed, subtitle traces are effectively removed, plot conflicts are understood through a deep learning algorithm, and then translation content conforming to target culture is generated. The method not only improves the accuracy of subtitle extraction, but also enhances the cultural adaptability of the translated content, so that the multimedia content can be more naturally and smoothly understood and accepted under different cultural backgrounds, and the watching experience of global users is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of multimedia technology, and particularly relates to a method for short play translation and commentary based on a multimodal neural network. Background Art

[0002] In the current field of multimedia content production, especially in the process of short play translation and commentary, there are a series of complex challenges. Existing methods usually rely on traditional OCR technology and basic video processing tools to extract subtitle information and perform simple video editing operations. However, these methods often face the problem of insufficient accuracy, especially when dealing with dynamically changing subtitles and subtitle recognition in complex backgrounds. In addition, in the process of cross-language translation, existing technologies are difficult to fully consider cultural differences and subtle differences in emotional expression, resulting in the final output content being difficult to achieve an ideal communication effect in different cultural backgrounds.

[0003] Currently, a major technical problem is that existing technologies are unable to efficiently and accurately combine multiple information sources when processing multimodal data (such as video, audio, and text) to generate a culturally adapted translation that not only conforms to the target language habits but also conveys the original emotional color. This not only affects the audience's understanding and acceptance of the content but also limits the dissemination efficiency of multimedia content globally. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for short play translation and commentary based on a multimodal neural network, which not only improves the accuracy of subtitle extraction but also enhances the cultural adaptability of the translated content, so as to solve the problems proposed in the above background art.

[0005] To achieve the above object, the present invention adopts the following technical solutions: A method for short play translation and commentary based on a multimodal neural network, comprising the following steps:

[0006] Extract subtitles from the video, determine the main subtitle area through dynamic peak detection, and filter outliers using the IQR method to output subtitle text;

[0007] Utilize the obtained subtitle text, combine optical flow field analysis and scene segmentation to locate subtitle boundaries, identify background texture and color information, and erase the subtitles without a trace according to the background information to restore the original picture;

[0008] Use LSTM to extract temporal features of the video, construct a character relationship network, analyze plot conflicts, and apply a multimodal neural network to fuse cross-language semantics to generate a culturally adapted translation according to the plot conflicts;

[0009] Combine the cultural adaptation translation with the original video content, conduct content analysis, prepare the materials required for dubbing synthesis, complete the dubbing synthesis based on the materials, and add them to the corresponding positions in the video;

[0010] Produce new subtitles and superimpose them on the video. At the same time, design the cover image, select the background music, and finally integrate all elements to output a condensed short video with commentary.

[0011] Technical effects and advantages of the present invention: A short play translation and commentary method based on a multi-modal neural network proposed by the present invention has the following advantages compared with the prior art:

[0012] By integrating advanced technologies such as an OCR model, optical flow field analysis, an LSTM network, and a graph neural network, the present invention realizes the full-process automated processing from subtitle extraction to cultural adaptation translation. In particular, the present invention can accurately identify and process subtitles in complex backgrounds, effectively remove subtitle traces, and understand plot conflicts through deep learning algorithms, thereby generating translation content that fits the target culture. This method not only improves the accuracy of subtitle extraction but also enhances the cultural adaptability of the translation content, enabling multimedia content to be more naturally and smoothly understood and accepted in different cultural backgrounds, greatly improving the viewing experience of global users. Description of the Drawings

[0013] Figure 1 It is a flowchart of a short play translation and commentary method based on a multi-modal neural network of the present invention. Detailed Embodiments

[0014] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. The specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0015] The present invention provides a short play translation and commentary method based on a multi-modal neural network as Figure 1 shown, including the following steps:

[0016] Step 1: Extract subtitles from the video, determine the main subtitle area through dynamic peak detection, and filter outliers using the IQR method to output subtitle text; specifically including:

[0017] Convert the video frame to grayscale, which can simplify the subsequent processing flow and improve the efficiency of subtitle area detection; obtain the brightness value L(x,y) of a single-frame image, where x and y represent the horizontal and vertical coordinates of the pixel point respectively; through grayscale conversion, the computational complexity can be reduced while retaining key information, which helps to locate the subtitle area more efficiently.

[0018] Based on the brightness value L(x,y), calculate the horizontal and vertical projections P_h and P_v of each frame image, and determine the range of the subtitle area through the formulas P_h = ∑_(y = 1)^n * L(x,y) and P_v = ∑_(x = 1)^m * L(x,y), where m and n are the width and height of the image respectively.

[0019] Among them, P_h = ∑_(y = 1)^n * L(x,y) means that for each fixed x coordinate, the brightness values of all pixel points in this column are accumulated to form a horizontal projection map.

[0020] P_v = ∑_(x = 1)^m * L(x,y) accumulates the brightness values of all pixel points in this row for each fixed y coordinate to form a vertical projection map. By analyzing these projection maps, the possible areas where subtitles may exist can be quickly identified. Because subtitles usually have a high brightness contrast, they will appear particularly prominent in the projection maps.

[0021] Apply the dynamic peak detection method within the range to identify the set of local maximum points S as the main subtitle area, and use the IQR method to filter out the outliers in the set. The calculation formula is Q_1 - 1.5 * (Q_3 - Q_1) < S_i < Q_3 + 1.5 * (Q_3 - Q_1), where Q_1 and Q_3 are the first and third quartiles respectively, and S_i represents the elements in the set; this formula is a statistical outlier detection method called the interquartile range (IQR) rule. It defines the normal value range based on the first (Q_1) and third (Q_3) quartiles of the data set, and the points outside this range are regarded as outliers and excluded. According to the filtered point set, extract the text information within the corresponding area and output the final subtitle text.

[0022] Example 1

[0023] Suppose there is a video with a resolution of 1920 * 1080 pixels. According to the above steps:

[0024] First, convert each frame into a grayscale image and calculate the brightness value L(x,y) of each pixel point.

[0025] Next, use the formulas P_h = ∑_(y = 1)^1080 L(x,y) and P_v = ∑_(x = 1)^1920 L(x,y) to calculate the horizontal and vertical projections of each frame image.

[0026] Apply the dynamic peak detection algorithm on the projection image to find potential subtitle regions, and use the IQR rule to filter out outliers. For example, if a value S_i in a point set S is not between Q_1 - 1.5*(Q_3 - Q_1) and Q_3 + 1.5*(Q_3 - Q_1), it is considered an outlier and excluded.

[0027] Finally, perform OCR operations within the region corresponding to the filtered point set to extract high-quality subtitle text.

[0028] In this way, not only can the subtitles in the video be accurately located and extracted, but also unnecessary interferences can be effectively removed, ensuring that the quality of the finally output subtitle text is high and reliable.

[0029] Step 2: Use the obtained subtitle text, combine optical flow field analysis and scene segmentation to locate subtitle boundaries and identify background texture and color information; specifically including:

[0030] Based on the subtitle region determined by the filtered point set, perform optical flow field analysis on the video frame sequence. Calculate the pixel displacement vector D(x, y) between every two adjacent frames through the formula D(x, y) = (L_t(x, y) - L_(t - 1)(x, y)) / Δt, where L_t and L_(t - 1) are the brightness values of the current frame and the previous frame at the position (x, y) respectively, and Δt is the time interval; the core idea of optical flow field analysis is to use brightness changes to infer the movement direction and speed of pixel points, thereby helping to distinguish dynamic subtitles from static backgrounds.

[0031] Use the displacement vector D(x, y), and combine the scene segmentation method to identify the subtitle boundary B according to the formula B = {(x, y)|D(x, y) > T}, where T is a set threshold to distinguish the subtitle region from the background region; separate the pixel points with larger displacements (i.e., the subtitle region) from the pixel points with smaller or no displacements (i.e., the background region), thereby accurately locating the subtitle boundary. The selection of the threshold T needs to comprehensively consider the characteristics of the video content and the relative motion intensity between the subtitles and the background.

[0032] For the area within the subtitle boundary B, the background texture feature F_t and color information C_c are extracted. The texture feature is obtained by calculating the Local Binary Pattern (LBP(x,y)), and the formula is LBP(x,y) = ∑_(i = 0)^7 2^i * s(g_i - g_c), where g_c is the gray value of the central pixel, g_i is the gray value of the i-th pixel around it, and s is the sign function representing the comparison result. Specifically, for each pixel point within the subtitle boundary B, taking it as the central pixel g_c, compare the gray value g_i of its 8 neighboring pixels with the gray value g_c of the central pixel, and generate the corresponding binary pattern. The LBP algorithm can effectively capture the local features of the background texture and provide important references for subsequent background reconstruction.

[0033] According to the background texture feature F_t and color information C_c, a background model M_b is constructed for seamless subtitle erasure operation. This model contains all the necessary information of the subtitle-covered area, such as texture, color distribution, etc. In actual operation, this model can be used to generate new pixels consistent with the original background, thus realizing seamless subtitle erasure.

[0034] Example 2

[0035] Suppose there is a video with a resolution of 1280×720, a frame rate of 30fps, and a time interval Δt of 1 / 30 second. According to the above steps:

[0036] Optical flow field analysis: Process the video frame sequence and calculate the displacement vector D(x,y) of each pixel point. For example, at the position (100,150), the current frame brightness value L_t(100,150) = 200, and the previous frame brightness value L_(t - 1)(100,150) = 190, then the displacement vector is: D(100,150) = (200 - 190) / (1 / 30) = 300.

[0037] Scene segmentation: Set the threshold T = 200. For all pixel points, if D(x,y) > 200, then mark it as the subtitle area. For example, the position (100,150) meets the condition and is thus included in the subtitle boundary B.

[0038] Texture feature extraction: Within the subtitle boundary B, select a pixel point (100,150), whose gray value g_c = 200, and the gray values of its 8 neighboring pixels are [210, 195, 205, 180, 220, 190, 200, 185]. Calculate the LBP value:

[0039] LBP(100,150) = 0*1 + 1*2 + 0*4 + 1*8 + 0*16 + 1*32 + 0*64 + 1*128 = 170.

[0040] Background model construction: Based on the extracted texture features and color information, a background model M_b is generated. During the actual erasure operation, M_b is used to generate new pixels that are consistent with the original background, completing the seamless erasure of the subtitles.

[0041] Step 3: Seamlessly erase the subtitles based on the background information and restore the original picture; specifically including:

[0042] Identify and separate the background texture feature F_t and color information C_c within the subtitle coverage area R_s according to the background model M_b; this information provides the necessary reference data for subsequent seamless erasure operations.

[0043] For each pixel point p in the subtitle coverage area R_s, calculate its similarity S with the surrounding non-subtitle area pixel points p_n through the formula S = 1 - |F_t(p) - F_t(p_n)| / max(F_t), where max(F_t) represents the maximum texture feature difference value within the area; the calculation of the similarity S is based on the difference of the texture feature F_t, and through normalization (i.e., using max(F_t)), the similarity value falls within the range of [0, 1]. The higher the similarity, the more similar the pixel point is to the surrounding non-subtitle area and is suitable for filling.

[0044] Reconstruct the subtitle coverage area R_s using the weighted average method based on the similarity S, and calculate the new pixel value P_new. The formula is P_new = ∑_(n = 1)^N * S_n * P_n / ∑_(n = 1)^N * S_n, where N is the number of non-subtitle area pixels selected for filling, and S_n and P_n are the similarity and the original pixel value of the nth pixel respectively; the weighted average method can effectively integrate the information of multiple reference pixel points to generate a smoother and more realistic background pixel value, thus achieving the seamless erasure of the subtitles.

[0045] Apply the reconstructed pixel value P_new to the corresponding position of the original video frame to restore the original picture. By applying the new pixel value to the original video frame, the subtitle area is perfectly restored, and the entire picture returns to the original state with no obvious repair marks visually.

[0046] Example 3

[0047] Suppose there is a video with a resolution of 1280×720, a frame rate of 30fps, and a time interval Δt of 1 / 30 second. According to the above steps:

[0048] Background information extraction: Suppose the subtitle coverage area R_s is within the range of (100, 150) to (300, 250) of the video frame. Extract the texture feature F_t and color information C_c of this area from the background model M_b.

[0049] Similarity calculation: For a pixel point p(120, 200) within the subtitle coverage area, its texture feature F_t(p) = 0.6. Three reference pixel points p_1, p_2, p_3 in the surrounding non-subtitle area are selected, and their texture features are F_t(p_1) = 0.5, F_t(p_2) = 0.7, and F_t(p_3) = 0.8 respectively.

[0050] Assume max(F_t) = 1, then the similarity calculation is as follows:

[0051] S1 = 1 - |0.6 - 0.5| / 1 = 0.9;

[0052] S2 = 1 - |0.6 - 0.7| / 1 = 0.9S2 = 1 - |0.6 - 0.7| / 1 = 0.9;

[0053] S3 = 1 - |0.6 - 0.8| / 1 = 0.8S3 = 1 - |0.6 - 0.8| / 1 = 0.8.

[0054] Reconstruction by weighted average method:

[0055] Assume the original pixel values of the reference pixel points p_1, p_2, p_3 are P_1 = 120, P_2 = 130, and P_3 = 140 respectively.

[0056] Calculate the new pixel value P_new:

[0057] P_new = (0.9 * 120 + 0.9 * 130 + 0.8 * 140) / (0.9 + 0.9 + 0.8);

[0058] P_new = (108 + 117 + 112) / 2.6 = 132.69Pnew = (108 + 117 + 112) / 2.6 = 132.69.

[0059] Apply the new pixel value: Apply the calculated new pixel value P_new = 132.69 to the position (120, 200) in the original video frame to complete the seamless erasure of the subtitle.

[0060] Step 4: Use LSTM to extract temporal features of the video, construct a character relationship network, and analyze plot conflicts; specifically including:

[0061] Based on the reconstructed video frame sequence, use the LSTM network to extract temporal features from consecutive frames to obtain the behavioral features B_t of each character at different time points, representing the behavioral patterns of the character at different time points. This step provides the basic data for subsequent analysis.

[0062] Using the behavioral feature \(B_t\), calculate the interaction intensity \(I_{ij}\) between any two roles at the same time point. The formula is \(I_{ij}=\sum_{t = 1}^{T}(B_i^t * B_j^t) / T\), where \(i\) and \(j\) are two different roles, and \(T\) is the length of the analysis time period. The numerator part of this formula is the sum of the products of the behavioral features of role \(i\) and role \(j\) at all time points, and the denominator \(T\) is used for averaging to make \(I_{ij}\) fall within a reasonable range.

[0063] Construct a character relationship network \(N_r\) based on the interaction intensity \(I_{ij}\). In this network, nodes represent roles, and edges represent the interaction intensity between roles. High-weight edges mean strong interactions, while low-weight edges mean the opposite. This network structure helps visualize and analyze the relationships between roles.

[0064] Analyze the plot conflict \(C_p\) in the character relationship network \(N_r\) by identifying high-density subgraphs \(H_s\) that appear in the network. These subgraphs usually contain strong and frequent interactions among multiple roles. That is, there are regions with frequent and strong interactions among multiple roles. The density \(D\) is quantified by the formula \(D = |E_s| / (|V_s|*(|V_s|-1) / 2)\), where \(E_s\) and \(V_s\) are the edge set and node set in the subgraph respectively. The density is evaluated by comparing the ratio of the actual number of edges in the subgraph to the theoretically maximum possible number of edges.

[0065] Example 4

[0066] Suppose there is a movie, and after processing, a reconstructed video frame sequence is obtained. Select a 5-minute (\(T = 300\) seconds) segment for analysis:

[0067] Behavioral feature extraction: Use an LSTM network to process this video frame sequence and extract the behavioral features \(B_A^t\) and \(B_B^t\) of the two main characters A and B at each time point.

[0068] Interaction intensity calculation: Suppose at some time points, the behavioral features of the main characters A and B are \(B_A^t=[0.8, 0.6]\) and \(B_B^t=[0.7, 0.5]\) respectively.

[0069] Calculate the interaction intensity \(I_{AB}=(0.8 * 0.7 + 0.6 * 0.5) / 2 = 0.59\).

[0070] Character relationship network construction: Construct a network according to the interaction intensity \(I_{ij}\) of all role pairs, where \(I_{AB}\) is used as the edge weight between A and B.

[0071] Plot conflict analysis: Find a subgraph \(H_s\) in the network that contains A, B, and other roles, and calculate its density \(D\).

[0072] Assume that the subgraph has 3 nodes (including A and B) and 3 edges, then: D = 3 / (3*(3 - 1) / 2) = 1. A high density indicates that this is a plot conflict point because there are frequent and intense interactions among the characters.

[0073] Step Five: According to the plot conflict, apply a multi-modal neural network to fuse cross-lingual semantics and generate a culture-adapted translation; specifically including:

[0074] Based on the plot conflict C_p identified by the high-density subgraph H_s, the key conflict points of the plot are determined. Extract the relevant fragment content C_c involving character dialogues or narrations; including direct dialogue texts, narrations, and other narrative contents that may affect the plot development.

[0075] For the content fragment C_c, apply a multi-modal neural network to analyze its visual V_v and auditory V_a elements, and combine the cross-lingual semantic information S_l to generate a preliminary translation text T_0. The cross-lingual semantic information is obtained by calculating the lexical similarity Sim between different languages. The formula is Sim = cos(θ) = (A·B) / (||A||||B||), where A and B represent the word vectors in two languages, and θ is the angle between A and B; by combining the cross-lingual semantic information S_l, use the lexical similarity Sim to quantify the lexical correspondence relationship between different languages. This step generates the preliminary translation text T_0 as the basis for subsequent culture adaptation.

[0076] Using the preliminary translation text T_0, combine the cultural gene adaptation detection to adjust the translation to match the background knowledge B_k of the target culture. Calculate the cultural fitness F_a = ∑_(i = 1)^nW_i*C_i / n, where W_i is the importance weight of the i-th cultural element, C_i represents the coverage rate of the corresponding cultural element in the translation text, and n is the total number of cultural elements; specifically, calculate the coverage rate C_i of each cultural element in the translation text, and combine its importance weight W_i to obtain the overall cultural fitness F_a. This process ensures that the translated text can be naturally understood and accepted in the context of the target culture.

[0077] Optimize the translation text based on the cultural fitness F_a to generate the final culture-adapted translation T_f. If the cultural fitness of some parts is low, targeted adjustments need to be made until a satisfactory fitness level is achieved.

[0078] Example 5

[0079] Suppose there is a movie clip. After processing, a key plot conflict C_p is identified, involving an intense dialogue between two protagonists A and B. This clip contains the following content:

[0080] Extract relevant fragment content: Determine the plot conflict points and extract the dialogues between protagonists A and B: "Why did you do this?" and "Because I had no choice."

[0081] Multimodal analysis and preliminary translation: Use a multimodal neural network to analyze the visual and auditory elements of the dialogue.

[0082] Assume the target language is French and calculate the lexical similarity Sim: Sim = cos(θ) = (A·B) / (||A||||B||);

[0083] Assume the word vectors of "Why" and "pourquoi" are A = [0.7, 0.5] and B = [0.6, 0.4] respectively: Sim = (0.7*0.6 + 0.5*0.4) / sqrt((0.7^2 + 0.5^2)*(0.6^2 + 0.4^2)) = 0.98.

[0084] Preliminary translation text T_0:

[0085]

[0086] Cultural adaptation detection and adjustment:

[0087] Calculate the cultural fitness F_a: Assume there are three cultural elements with weights W_1 = 0.5, W_2 = 0.3, W_3 = 0.2 and coverage rates C_1 = 0.8, C_2 = 0.6, C_3 = 0.7 respectively:

[0088] F_a = (0.5*0.8 + 0.3*0.6 + 0.2*0.7) / 3 = 0.7. According to the F_a value, fine-tune the translated text to make it more suitable for the target cultural background.

[0089] Final translated text: After optimization, the finally generated culturally adapted translation T_f: "Pourquoi as-tu agi de cette manière? Parce que'il n'y avait pas d'autre option pour moi.".

[0090] Step Six: Combine the culturally adapted translation with the original video content for content analysis and prepare the materials required for dubbing synthesis; specifically including:

[0091] Translate \(T_f\) according to cultural adaptation, match the translated text with the corresponding scene \(S_c\) in the original video content. For each translated segment, find its corresponding scene in the video and determine the timestamp \(T_s\) in the video. The formula is \(T_s=(Start_t + End_t) / 2\), where \(Start_t\) and \(End_t\) are the start and end times of the corresponding scene respectively. Specifically, by obtaining the start time \(Start_t\) and end time \(End_t\) of the corresponding scene, use the formula \(T_s=(Start_t + End_t) / 2\) to calculate the midpoint time as the timestamp to ensure that the translated text can be accurately mapped to the correct position in the video.

[0092] Using the timestamp \(T_s\), segment the video content and extract the audio segment \(A_f\) and visual segment \(V_f\) related to the culturally adapted translation \(T_f\); these segments include audio elements such as character dialogues, background music, environmental sound effects, etc. and the corresponding visual images, providing the necessary materials for subsequent dubbing synthesis.

[0093] Based on the content of the audio segment \(A_f\) and visual segment \(V_f\), analyze the speech intonation requirements \(D_r\), and calculate the emotional intensity \(E_i=\sum_{j = 1}^mP_j*W_j / m\), where \(P_j\) represents the occurrence frequency of the \(j\)-th emotional word, \(W_j\) is its weight, and \(m\) is the total number of emotional words analyzed; through this formula, the emotional intensity contained in the text can be quantified, thus guiding the adjustment of intonation and tone during dubbing.

[0094] Prepare the materials required for the dubbing synthesis of the culturally adapted translation \(T_f\), including the translated text adjusted according to the emotional intensity \(E_i\) and the time information of the corresponding audio segment \(A_f\) and visual segment \(V_f\). These materials will be used in the subsequent dubbing synthesis process to ensure that the final output work is both faithful to the original work and has good audio-visual effects.

[0095] Example 6

[0096] Suppose there is a movie clip, and after processing, a culturally adapted translation \(T_f\) is generated: "Pourquoi as-tu agi de cette manière? Parce que'il n'y avait pas d'autre option pour moi."

[0097] Timestamp calculation: Suppose the start time of the scene corresponding to this clip is \(Start_t = 00:05:10\) and the end time is \(End_t = 00:05:20\), then the timestamp \(T_s\) is calculated as follows:

[0098] \(T_s=(00:05:10 + 00:05:20) / 2 = 00:05:15\).

[0099] Video content segmentation: According to the timestamp T_s, extract the corresponding audio segment A_f (such as the character dialogue "Why did you do this?") and visual segment V_f (such as the expressions and actions of the characters) from the video.

[0100] Emotional intensity calculation: Suppose there are three emotional words with frequencies P_1 = 0.3, P_2 = 0.5, P_3 = 0.2 and weights W_1 = 0.6, W_2 = 0.7, W_3 = 0.8 respectively. Then the emotional intensity E_i is calculated as follows: E_i = (0.3 * 0.6 + 0.5 * 0.7 + 0.2 * 0.8) / 3 = 0.65.

[0101] Dubbing synthesis material preparation: Adjust the translation text according to the emotional intensity E_i: "Pourquoi as - tu agi de cette manière avec une telle intensité? Parce que'il n'y avait pas d'autre option pour moi."

[0102] Prepare the adjusted translation text, the time information of the audio segment A_f and the visual segment V_f for subsequent dubbing synthesis.

[0103] Step Seven: Complete the dubbing synthesis based on the materials and add it to the corresponding position in the video; specifically including:

[0104] Select the matching voice sample library V_b according to the translation text adjusted by the emotional intensity E_i and the corresponding time information; these samples should match the emotional color in the translation text and be able to adapt to the pronunciation characteristics of the target language. The selection process needs to ensure that the voice characteristics (such as timbre, intonation, etc.) of the selected voice samples are coordinated with the overall style of the video content.

[0105] Synthesize the dubbing audio A_s using the elements in the voice sample library V_b, and adjust the pitch Pitch and speed Speed to meet the requirements of the emotional intensity E_i. The formula is Adj = Pitch * Speed / E_i, where Adj represents the adjustment coefficient used to balance the pitch and speed to better adapt to the emotional intensity; ensure that the dubbing audio can accurately convey the emotional color in the text in terms of pitch and speed.

[0106] Based on the time information T_s, the synthesized dubbing audio A_s is accurately added to the corresponding position P_v of the video, and the best alignment point is determined using the synchronization error calculation method. The formula is Err = ∑_(k = 1)^n*(T_k-A_k)^2 / n, where T_k is the target timestamp, A_k is the starting time point of the dubbing audio, and n is the number of time points analyzed. By calculating the synchronization error, the best alignment point that minimizes the synchronization error is found, thereby achieving seamless connection between the dubbing audio and the video screen.

[0107] Example 7

[0108] Suppose there is a movie clip, which is processed to generate a culturally adapted translation T_f: "Pourquoias-tuagidecettemanièreavecunetelleintensité? Parcequ'iln'yavaitpasd'autreoptionpourmoi.", and the calculated sentiment intensity E_i is about 0.65.

[0109] Select the speech sample library: According to the emotion intensity E_i and the content of the translated text, select the appropriate speech sample from the speech sample library V_b to ensure that its emotional expression is consistent with the original text.

[0110] Synthesized dubbing audio: Assuming that the initial pitch Pitch = 1.0 and the speech speed Speed ​​= 1.0, the adjustment coefficient Adj is calculated as follows: Adj = Pitch*Speed / E_i = 1.0*1.0 / 0.65 = 1.54.

[0111] According to the adjustment coefficient Adj, the pitch and speaking speed are appropriately increased to meet the requirements of the emotional intensity E_i to generate the dubbing audio A_s.

[0112] Synchronization processing: Assume that there are three key time points that need to be synchronized, namely T_1 = 00:05:10, T_2 = 00:05:15, T_3 = 00:05:20; the corresponding dubbing audio starting time points are A_1 = 00:05:09, A_2 = 00:05:14, A_3 = 00:05:19, then the synchronization error is calculated as follows:

[0113] Err=((00:05:10-00:05:09)^2+(00:05:15-00:05:14)^2+(00:05:20-00:05:19)^2) / 3=(1^2+1^2+1^2) / 3=1.

[0114] By fine-tuning the starting time point of the dubbing audio, the synchronization error Err is minimized, ensuring perfect synchronization between the dubbing audio and the video image.

[0115] Step Eight: Create new subtitles and superimpose them on the video, and at the same time design the cover image and select the background music; specifically including:

[0116] Based on the time information Ts of the dubbed audio As, create new subtitles Sn that match the culturally adapted translation Tf, ensuring that the display time Start_t and end time End_t of each subtitle paragraph accurately correspond to the dialogue or narrative part in the dubbed audio As. The formula is Duration = End_t - Start_t; calculate the duration of each subtitle paragraph to ensure that the subtitle display duration is appropriate and will not disappear too early or appear late.

[0117] Using the time information Ts and the subtitle text content, superimpose the new subtitles Sn on the video. Adjust the transparency Alpha so that the subtitles are clear and do not obscure key visual elements. The calculation formula is Alpha = 1 - (B_v / Max_B), where B_v is the brightness value of the background video in the subtitle area and Max_B is the maximum brightness value in the video; this can make the subtitles more transparent against a bright background and more obvious against a darker background.

[0118] Design the cover image Ci. Based on the main color distribution Color_d and emotional intensity Ei of the video, select the image elements that best represent the video theme. The color weight calculation formula is W_c = ∑_(x = 1)^X * Color_x * E_x / X, where Color_x represents the xth main color and E_x is its corresponding emotional intensity, and X is the number of colors analyzed; this step helps to generate a cover image that can not only attract the audience's attention but also convey the emotional tone of the film.

[0119] Select the background music Bgm to enhance the overall atmosphere of the video. Select the music segment according to the emotional tone Eb of the video and the color weight W_c of the cover image Ci, and ensure that the music rhythm R is synchronized with the climax part of the video. The formula is Sync = |P_v - P_m|, where P_v is the position of the video climax point and P_m is the position of the background music climax point. Minimize the Sync value to obtain the best synchronization effect. The formula is used to quantify the synchronization error between the video and the music climax point, and minimize the Sync value to achieve the best synchronization effect.

[0120] Example 8

[0121] Suppose there is a movie clip, and after processing, a culturally adapted translation T_f is generated: "Pourquoi as-tu agi de cette manière avec une telle intensité? Parce qu'il n'y avait pas d'autre option pour moi.", and the calculated emotional intensity E_i is approximately 0.65.

[0122] Create new subtitles: According to the time information T_s of the dubbed audio A_s, assume the start time of the subtitle paragraph is Start_t = 00:05:10 and the end time is End_t = 00:05:15. Then the duration Duration of the subtitle paragraph is calculated as follows: Duration = End_t - Start_t = 00:05:15 - 00:05:10 = 5 seconds.

[0123] Adjust subtitle transparency: Assume the brightness value B_v of the background video in the subtitle area is 180, and the maximum brightness value Max_B in the video is 255. Then the subtitle transparency Alpha is calculated as follows: Alpha = 1 - (B_v / Max_B) = 1 - (180 / 255) = 0.29.

[0124] Cover image design: Suppose there are three main colors, and their emotional intensities are E_1 = 0.7, E_2 = 0.5, E_3 = 0.8 respectively. Then the color weight W_c is calculated as follows:

[0125] W_c = (Color_1 * E_1 + Color_2 * E_2 + Color_3 * E_3) / 3 = (0.7 + 0.5 + 0.8) / 3 = 0.67.

[0126] Background music selection: Assume the position of the video climax point P_v = 00:10:00, and the position of the background music climax point P_m = 00:09:55. Then the synchronization error Sync is calculated as follows: Sync = |P_v - P_m| = |00:10:00 - 00:09:55| = 5 seconds.

[0127] Through the above steps, high-quality new subtitles were successfully created and added to the video, a cover image that matches the theme of the film was designed, and background music with good synchronization was selected, ultimately enhancing the overall audio-visual experience.

[0128] Step Nine: Finally, integrate all elements and output a condensed short video with commentary; specifically including:

[0129] Adjust the volume V_b of the background music Bgm according to the background music Bgm and its synchronization value Sync with the climax part of the video to ensure harmonious coexistence with the dubbed audio A_s. The calculation formula is V_b = (L_max - L_min) * (1 - Sync) + L_min, where L_max and L_min are the maximum and minimum volumes allowed for the background music respectively. This formula ensures that when the Sync value is high (i.e., the synchronization is poor), the volume of the background music will be low, and vice versa, thus achieving the best auditory experience.

[0130] Use the new subtitles S_n and the cover image C_i, and combine with the video content for final editing. Condense the original video through the editing rate Rate, and the formula is Rate = T_total / T_final, where T_total is the total duration of the original video, and T_final is the expected duration of the condensed video. This ratio determines the speed of video editing, helps to remove redundant parts, retain key plots, and ensure that the condensed video can convey the main information without losing its attractiveness.

[0131] Generate an opening animation Opening based on the color weight W_c and the emotional intensity E_i. The length L_o of the animation is determined by the formula L_o = α * T_final, where α is a preset scaling factor used to balance the proportion of the opening animation and the main content. Selecting an appropriate value of α can ensure that the opening animation can attract the audience's attention without being too long to affect the display of the main content.

[0132] Integrate all elements including the dubbed audio A_s, the background music Bgm, the new subtitles S_n, the cover image C_i, and the opening animation Opening, and output the condensed short video Final_v with commentary.

[0133] Example 9

[0134] Suppose there is a movie clip, and after processing, a culturally adapted translation T_f is generated: "Pourquoi as-tu agi de cette manière avec une telle intensité? Parce qu'il n'y avait pas d'autre option pour moi.", and the calculated emotional intensity E_i is approximately 0.65.

[0135] Adjustment of the background music volume: Suppose the maximum volume L_max allowed for the background music is 100%, the minimum volume L_min is 20%, and the synchronization error Sync = 0.1. Then the background music volume V_b is calculated as follows:

[0136] V_b = (L_max - L_min) * (1 - Sync) + L_min = (100% - 2 / XMLSchema1) * (1 - 0.1) + 20% = 90%.

[0137] Concentration processing: Assume the total duration of the original video T_total = 60 minutes, and the desired duration of the concentrated video T_final = 10 minutes. Then the editing rate Rate is calculated as follows: Rate = T_total / T_final = 60 minutes / 10 minutes = 6.

[0138] Length of the opening animation: Assume the preset scaling factor α = 0.1. Then the length of the opening animation L_o is calculated as follows: L_o = α * T_final = 0.1 * 10 minutes = 1 minute.

[0139] Final integration: Integrate all elements - the dubbed audio A_s, the background music Bgm (volume set to 90%), the new subtitles S_n, the cover image C_i, and the 1-minute opening animation Opening - together to output the final concentrated short video Final_v.

[0140] Through the above steps, a high-quality concentrated short video was successfully produced, ensuring the harmony and unity of all elements and providing a high-quality audio-visual experience.

[0141] In summary, the present invention realizes the full-process automated processing from subtitle extraction to culture-adapted translation by integrating advanced technologies such as OCR models, optical flow field analysis, LSTM networks, and graph neural networks. In particular, the present invention can accurately identify and process subtitles in complex backgrounds, effectively remove subtitle traces, and understand plot conflicts through deep learning algorithms, and then generate translation content that fits the target culture.

[0142] This method not only improves the accuracy of subtitle extraction but also enhances the cultural adaptability of the translation content, enabling multimedia content to be more naturally and smoothly understood and accepted in different cultural backgrounds, greatly enhancing the viewing experience of global users.

[0143] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A short play translation and commentary method based on a multi-modal neural network, characterized in that, It includes the following steps: Extract subtitles from the video, determine the main subtitle area through dynamic peak detection, filter outliers using the IQR method, and output subtitle text; Utilize the obtained subtitle text, combine optical flow field analysis and scene segmentation to locate subtitle boundaries, identify background texture and color information, and seamlessly erase the subtitles based on the background information to restore the original picture; Use LSTM to extract temporal features of the video, construct a character relationship network, analyze plot conflicts, and apply a multi-modal neural network to fuse cross-language semantics according to the plot conflicts to generate culturally adapted translations; Combine the culturally adapted translation with the original video content, conduct content analysis, prepare materials required for dubbing synthesis, complete dubbing synthesis based on the materials, and add them to the corresponding positions in the video; Produce new subtitles and superimpose them on the video, design a cover image, select background music, and finally integrate all elements to output a condensed short video with commentary; 2. The method for short play translation and commentary based on a multimodal neural network according to claim 1, wherein, The output subtitle text includes: Convert the video frames to grayscale to obtain the luminance value L(x,y) of a single-frame image, where x and y represent the horizontal and vertical coordinates of the pixel points respectively; Based on the luminance value L(x,y), calculate the horizontal and vertical projections P_h and P_v of each frame image, and determine the range of the subtitle area through the formulas P_h = ∑_(y = 1)^n * L(x,y) and P_v = ∑_(x = 1)^m * L(x,y), where m and n are the width and height of the image respectively; Apply the dynamic peak detection method within the range to identify the set of local maximum points S as the main subtitle area, and filter outliers in the set using the IQR method. The calculation formula is Q_1 - 1.5 * (Q_3 - Q_1) < S_i < Q_3 + 1.5 * (Q_3 - Q_1), where Q_1 and Q_3 are the first and third quartiles respectively, and S_i represents the elements in the set; Extract the text information within the corresponding area according to the filtered point set and output the final subtitle text.

3. The method for short play translation and commentary based on a multi-modal neural network according to claim 2, characterized in that, The identification of background texture and color information includes: Based on the subtitle area determined by the filtered point set, conduct optical flow field analysis on the video frame sequence, and calculate the pixel point displacement vector D(x,y) between every two adjacent frames through the formula D(x,y) = (L_t(x,y) - L_(t - 1)(x,y)) / Δt, where L_t and L_(t - 1) are the luminance values of the current frame and the previous frame at the position (x,y) respectively, and Δt is the time interval; Utilize the displacement vector D(x,y), and combine the scene segmentation method to identify the subtitle boundary B according to the formula B = {(x,y)|D(x,y) > T}, where T is a set threshold to distinguish the subtitle area from the background area; For the area within the subtitle boundary B, extract the background texture feature F_t and color information C_c. The texture feature is obtained by calculating the local binary pattern LBP(x,y), and the formula is LBP(x,y) = ∑_(i = 0)^7 2^i * s(g_i - g_c), where g_c is the gray value of the central pixel, g_i is the gray value of the i-th pixel around it, and s is the sign function representing the comparison result; Construct a background model \(M_b\) based on the background texture feature \(F_t\) and color information \(C_c\) for seamless subtitle erasure operation.

4. The method for short play translation and commentary based on a multimodal neural network according to claim 3, wherein, The seamless erasure of subtitles based on background information and restoration of the original picture include: Identify and separate the background texture feature \(F_t\) and color information \(C_c\) within the subtitle-covered area \(R_s\) according to the background model \(M_b\); For each pixel point \(p\) in the subtitle-covered area \(R_s\), calculate its similarity \(S\) with the pixel points \(p_n\) in the surrounding non-subtitle area through the formula \(S = 1-|F_t(p)-F_t(p_n)| / \max(F_t)\), where \(\max(F_t)\) represents the maximum texture feature difference value within the area; Reconstruct the subtitle-covered area \(R_s\) using the weighted average method based on the similarity \(S\), and calculate the new pixel value \(P_{new}\), with the formula \(P_{new}=\sum_{n = 1}^N S_n P_n / \sum_{n = 1}^N S_n\), where \(N\) is the number of non-subtitle area pixels selected for filling, and \(S_n\) and \(P_n\) are the similarity and original pixel value of the \(n\)th pixel respectively; Apply the reconstructed pixel value \(P_{new}\) to the corresponding position of the original video frame to restore the original picture.

5. The method for short drama translation and commentary based on a multimodal neural network according to claim 4, characterized in that The parsing of the plot conflict includes: According to the reconstructed video frame sequence, extract the temporal features of consecutive frames through the LSTM network to obtain the behavior features \(B_t\) of each character at different time points; Using the behavior features \(B_t\), calculate the interaction intensity \(I_{ij}\) between any two characters at the same time point, with the formula \(I_{ij}=\sum_{t = 1}^T(B_i^t\cdot B_j^t) / T\), where \(i\) and \(j\) are two different characters, and \(T\) is the length of the analysis time period; Construct a character relationship network \(N_r\) based on the interaction intensity \(I_{ij}\); Analyze the plot conflict \(C_p\) in the character relationship network \(N_r\) by identifying the high-density subgraphs \(H_s\) in the network, that is, areas with frequent and strong interactions among multiple characters. The density \(D\) is quantified by the formula \(D = |E_s| / (|V_s|*(|V_s|-1) / 2)\), where \(E_s\) and \(V_s\) are the edge set and node set in the subgraph respectively.

6. The method for short play translation and commentary based on a multi-modal neural network according to claim 5, wherein The generation of culture-adapted translation includes: Extract the relevant fragment content \(C_c\) involving character dialogues or narrations according to the plot conflict \(C_p\) identified by the high-density subgraph \(H_s\); For the content fragment \(C_c\), apply a multi-modal neural network to analyze its visual \(V_v\) and auditory \(V_a\) elements, and generate a preliminary translation text \(T_0\) in combination with cross-lingual semantic information \(S_l\). The cross-lingual semantic information is obtained by calculating the lexical similarity \(Sim\) between different languages, with the formula \(Sim=\cos(\theta)=(A\cdot B) / (\|A\|\|B\|)\), where \(A\) and \(B\) represent the word vectors in two languages, and \(\theta\) is the angle between \(A\) and \(B\); Using the preliminary translation text T_0, combined with the meme adaptation detection, adjust the translation to match the background knowledge B_k of the target culture, and calculate the cultural adaptation fitness F_a = ∑_(i = 1)^nW_i*C_i / n, where W_i is the importance weight of the i-th cultural element, C_i represents the coverage rate of the corresponding cultural element in the translation text, and n is the total number of cultural elements; Optimize the translation text based on the cultural adaptation fitness F_a to generate the final culture-adapted translation T_f.

7. A method for short drama translation and commentary based on a multi-modal neural network according to claim 6, characterized in that Preparing the materials required for dubbing synthesis, including: According to the culture-adapted translation T_f, match the translation text with the corresponding scene S_c in the original video content, and determine the timestamp T_s of each translation segment in the video. The formula is T_s = (Start_t + End_t) / 2, where Start_t and End_t are the start and end times of the corresponding scene respectively; Using the timestamp T_s, segment the video content to extract the audio segment A_f and visual segment V_f related to the culture-adapted translation T_f; Based on the content of the audio segment A_f and visual segment V_f, analyze the speech intonation requirements D_r, and calculate the emotional intensity E_i = ∑_(j = 1)^mP_j*W_j / m, where P_j represents the occurrence frequency of the j-th emotional word, W_j is its weight, and m is the total number of emotional words analyzed; Prepare the materials required for the dubbing synthesis of the culture-adapted translation T_f, including the translation text adjusted according to the emotional intensity E_i, and the time information of the corresponding audio segment A_f and visual segment V_f.

8. A method for short play translation and commentary based on a multi-modal neural network according to claim 7, characterized in that Completing the dubbing synthesis based on the materials and adding it to the corresponding position in the video, including: According to the translation text adjusted according to the emotional intensity E_i and the corresponding time information, select the matching voice sample library V_b; Use the elements in the voice sample library V_b to synthesize the dubbing audio A_s, and adjust the pitch Pitch and speed Speed to meet the requirements of the emotional intensity E_i. The formula is Adj = Pitch*Speed / E_i, where Adj represents the adjustment coefficient; Based on the time information T_s, accurately add the synthesized dubbing audio A_s to the corresponding position P_v of the video, and use the synchronous error calculation method to determine the best alignment point. The formula is Err = ∑_(k = 1)^n*(T_k - A_k)^2 / n, where T_k is the target timestamp, A_k is the start time point of the dubbing audio, and n is the number of time points analyzed.

9. A method for short play translation and commentary based on a multimodal neural network according to claim 8, characterized in that, Producing new subtitles and suppressing them onto the video, and at the same time designing the cover image and selecting the background music, including: According to the time information T_s of the dubbing audio A_s, produce new subtitles S_n that match the culture-adapted translation T_f, and ensure that the display time Start_t and end time End_t of each subtitle paragraph accurately correspond to the dialogue or narrative part in the dubbing audio A_s. The formula is Duration = End_t - Start_t; Using the time information \(T_s\) and the subtitle text content, the new subtitle \(S_n\) is superimposed on the video. By adjusting the transparency Alpha, the subtitle is made clear without obscuring key visual elements. The calculation formula is Alpha = 1 - (B_v / Max_B), where B_v is the brightness value of the background video in the subtitle area and Max_B is the maximum brightness value in the video; Design the cover image \(C_i\). Based on the main color distribution Color_d and the emotional intensity \(E_i\) of the video, select the image elements that best represent the video theme. The color weight calculation formula is \(W_c=\sum_{x = 1}^X\frac{Color_x*E_x}{X}\), where Color_x represents the \(x\)-th main color, \(E_x\) is its corresponding emotional intensity, and X is the number of colors analyzed; Select the background music Bgm to enhance the overall atmosphere of the video. Select music segments according to the emotional tone \(E_b\) of the video and the color weight \(W_c\) of the cover image \(C_i\). Ensure that the music rhythm R is synchronized with the climax part of the video. The formula is Sync = |P_v - P_m|, where P_v is the position of the video climax point and P_m is the position of the background music climax point. Minimize the Sync value to obtain the best synchronization effect.

10. A method for short play translation and commentary based on a multi-modal neural network according to claim 9, characterized in that Finally, integrate all elements and output a condensed short video with commentary, including: According to the background music Bgm and its synchronization value Sync with the climax part of the video, adjust the volume \(V_b\) of the background music to ensure harmonious coexistence with the dubbed audio \(A_s\). The calculation formula is \(V_b=(L_{max}-L_{min})*(1 - Sync)+L_{min}\), where \(L_{max}\) and \(L_{min}\) are the maximum and minimum volumes allowed for the background music respectively; Using the new subtitle \(S_n\) and the cover image \(C_i\), combine with the video content for final editing. Condense the original video through the editing rate Rate. The formula is Rate = \(T_{total} / T_{final}\), where \(T_{total}\) is the total duration of the original video and \(T_{final}\) is the desired duration of the condensed video; Based on the color weight \(W_c\) and the emotional intensity \(E_i\), generate an opening animation Opening. The length \(L_o\) of the animation is determined by the formula \(L_o = \alpha*T_{final}\), where \(\alpha\) is a preset scaling factor used to balance the proportion of the opening animation to the main content; Integrate all elements including the dubbed audio \(A_s\), background music Bgm, new subtitle \(S_n\), cover image \(C_i\) and opening animation Opening, and output a condensed short video Final_v with commentary.

Citation Information

Cited By

  • Short drama subtitle translation system based on artificial intelligence

    CN120996056A

  • A short play subtitle translation system based on artificial intelligence

    CN120996056B

  • Method and device for translating and replacing characters in picture

    CN122160555A