AI-driven video content subtitle synchronous translation method and system

Through the AI-driven video content subtitle synchronization translation system, combining video frame acquisition, facial lip action recognition and multi-person speech overlap recognition modules, the precise synchronization of video subtitles is achieved, solving the time deviation problem of traditional systems in multi-person dialogue and voice overlap scenarios, and improving the synchronization accuracy of subtitles and audience experience.

CN119996778APending Publication Date: 2025-05-13CHONGQING MALYA MEDIA CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202411137015.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-19
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

When traditional video subtitle synchronous translation systems handle multi-person conversations and voice overlap scenes, there is a time deviation in subtitle display, which makes it difficult to ensure the accuracy and timeliness of subtitles.

Method used

The AI-driven video content subtitle synchronization translation system is adopted, and through components such as the video frame acquisition module, facial lips movement recognition module, audio acquisition and translation module, multi-person speech overlap recognition module, and other components, it realizes accurate capture and analysis of the lip characteristics and speech speed changes of each character, and performs two-level corrections to synchronize the subtitle timestamp.

Benefits of technology

It effectively reduces the subtitle time deviation caused by speech speed differences and voice overlap, improves the synchronization accuracy of subtitles and the viewer's viewing experience, and ensures the accuracy of subtitles in complex dialogue scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996778A_ABST
    Figure CN119996778A_ABST
Patent Text Reader

Abstract

The invention discloses an AI-driven video content subtitle synchronous translation method and system, and relates to the technical field of video subtitle synchronization, the system is combined with a video frame acquisition module and a face and lip action recognition module, and the system can accurately obtain the lip opening and closing vertical distance and opening and closing times of each role. The data are used for calculating the actual speaking speed, and the actual speaking speed is compared with a traditional speed index to obtain a first calibration difference coefficient. According to the method, the timestamps of the subtitles are effectively adjusted, the time deviation caused by the speech speed difference is reduced, the subtitles and the actual speech are more synchronous, and therefore the accuracy of the subtitles and the film watching experience of audiences are improved. The multi-person talking overlapping recognition module can accurately detect and mark the voice overlapping condition. And if the overlapped voice influence factor D exceeds the abnormal threshold F, the system triggers the second correction instruction to further calibrate the subtitle timestamp, so that the synchronization problem caused by voice overlapping is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video subtitle synchronization, and specifically to an AI-driven video content subtitle synchronization translation method and system. Background Art

[0002] In traditional video content subtitle synchronization translation systems, the time synchronization of subtitles and video content usually relies on direct analysis and processing of audio. However, traditional technologies have great limitations when dealing with multi-person conversations and overlapping voice scenes, which makes it difficult to guarantee the accuracy and timeliness of subtitle display. Everyone's speaking speed is different, and even the same person's speaking speed may change at different times or situations. This difference in speaking speed is often difficult to handle in traditional subtitle synchronization systems, and usually leads to time deviation in subtitle display. Subtitle synchronization technology usually relies on direct analysis of audio signals. If the speaker's speaking speed changes cannot be accurately captured, the subtitles will appear ahead of time or behind time. This time deviation not only affects the audience's viewing experience, but may also lead to misunderstandings and errors in information transmission.

[0003] At the same time, the overlapping of multiple people speaking also affects the accuracy of subtitles. The existing technology lacks effective separation and recognition methods for audio segments of different characters, and cannot meet the needs of high-precision subtitle synchronization translation. Therefore, it is urgent to propose an AI-driven video content subtitle synchronization translation method and system. Through AI-driven technology, the system can accurately capture the lip features of each character, which helps to accurately capture the speaker's speech speed changes. Summary of the invention

[0004] In view of the deficiencies in the prior art, the present invention provides an AI-driven video content subtitle synchronous translation method and system to solve the problems mentioned in the background technology.

[0005] To achieve the above objectives, the present invention is implemented through the following technical solutions: an AI-driven video content subtitle synchronous translation system, including a video frame acquisition module, an audio acquisition and translation module, a facial lip action recognition module, a first correction module, a multi-person speech overlap recognition module and a second correction module;

[0006] The video frame acquisition module is used to acquire the video stream that needs subtitle translation synchronization, and establish a video stream set and a first image sequence set;

[0007] The audio acquisition and translation module is used to extract the audio stream corresponding to the video stream from the video stream set, separate the audio of different characters from the audio stream, use the speaker separation technology Speaker-Diarization, use machine learning and signal processing algorithms to identify the audio segments of different characters, and perform analysis and calculation to obtain the first speech rate index Wxz of the i-th character in the j-th audio segment i,j and first voice translated text;

[0008] The facial lip action recognition module is used to extract lip area features from each frame of the first image sequence set, obtain lip key point data through the 68-point facial marker model in Dlib driven by AI, and extract the lip opening and closing vertical distance features and the start and end time features of each character's speech based on the lip key point data to obtain the lip opening and closing vertical distance Dv and the lip opening and closing times Nv, and perform deep calculation to obtain: the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j , and the second actual speaking speed index Sxz i,j and the first speech rate index Wxz i,j Perform an associated calculation to obtain the first calibration difference coefficient of the i-th character in the j-th audio segment

[0009] The first correction module is used to adjust the first calibration difference coefficient Used to synchronously check the start time and end time of the first voice translation text to obtain the timestamp of the second voice translation text;

[0010] The multi-person speech overlap recognition module is used to extract whether there are multiple voice activity segment features in the jth audio segment of the video stream set during the speech of the i-th character, so as to identify and calculate the overlapping speech segments in the k-th audio segment to obtain the overlapping speech impact factor D. If the overlapping speech impact factor D is greater than the abnormal threshold F, the second correction instruction is triggered;

[0011] The second correction module is used to, after receiving the second correction instruction, correct the overlapped speech impact factor D and the first calibration difference coefficient The second calibration coefficient is calculated and associated The second calibration difference coefficient The second speech translation text timestamp is synchronously corrected to obtain a third speech translation text timestamp.

[0012] Preferably, the video frame acquisition module includes a first acquisition unit and a first preprocessing unit;

[0013] The first acquisition unit is used to acquire a video stream that needs subtitle translation synchronization, and use a video decoder to decode the video stream into original frame image data; according to a set frame rate of 30 frames per second, continuously extract frame images from the video stream and store them as an image sequence;

[0014] The first preprocessing unit is used to preprocess the extracted image sequence, including denoising, image enhancement and scaling, and store the preprocessed image sequence into the first image sequence set.

[0015] Preferably, the audio collection and translation module includes a second collection unit, a second processing unit, an AI calculation unit and a recognition marking unit;

[0016] The second acquisition unit is used to extract an audio stream from the video data set through an audio decoder, and decode the audio stream into an original audio signal;

[0017] The second processing unit is used to remove background noise from the original audio signal by using a Wiener filtering method, and to adjust the gain of the original audio signal;

[0018] The AI ​​calculation unit is used to divide the processed original audio signal into several frames, and calculate each short-time window to obtain the nth Mel frequency cepstral coefficient C n , the nth Mel frequency cepstrum coefficient C n The calculation method includes the following steps:

[0019] S1. Pre-emphasize the original audio signal and calculate the pre-emphasized signal value of the nth sampling point using the following formula

[0020]

[0021] In the formula, x(n) represents the amplitude of the original audio signal at the nth sampling point, x(n-1) represents the amplitude of the original audio signal at the n-1th sampling point, α represents the pre-emphasis coefficient, and the setting value is 0.97;

[0022] S2, divide the original audio signal after pre-emphasis into several frames, each frame is N sampling points long, the frames are set to overlap by 50%, and the pre-emphasis signal value of the nth sampling point is Multiply by the window function w(n), and obtain the windowed signal value y(n)- of the nth sampling point by the following formula:

[0023]

[0024] Where w(n) represents the window function, and n represents the time index in the current window;

[0025] S3, and perform Fourier transformation on the windowed signal value y(n) of the nth sampling point to obtain the frequency domain signal X(k):

[0026]

[0027] Where X(k) represents the frequency signal, which represents the value at the frequency index k; N represents the length of the frame, that is, the number of sampling points contained in each frame; k represents the index value of the frequency component, ranging from 0 to N-1; p represents the imaginary unit, satisfying p 2 = -1, π represents the ratio of pi, and the set value is 3.14159; each time sampling point y(n) is compared with a complex exponential function Multiply, kn is the product of frequency index k and time index n, the complex exponential function represents a complex rotation, and the frequency is proportional to k; this step converts the discrete signal in the time domain into a discrete signal in the frequency domain;

[0028] S4, and obtain the power spectrum P(k) according to the frequency domain signal X(k) by the following fast Fourier transform (FFT) formula:

[0029]

[0030] In the formula, |X(k)| represents the complex amplitude of the frequency domain signal, which reflects the strength of the signal at a specific frequency k;

[0031] S5. Pass the power spectrum P(k) through a set of triangular filters on the Mel scale to obtain the filtered energy to obtain the output value M of the mth filter. m :

[0032]

[0033] In the formula, H m (k) represents the mth filter response value in the Mel filter bank, and is calculated by the output value M of the mth filter. m After logarithmic transformation, the logarithmic energy value log(M m );

[0034] S6, and the logarithmic energy value log(M m ) is converted to the nth Mel frequency cepstrum coefficient C by the following formula n :

[0035]

[0036] Where M represents the number of filters; Represents the phase angle in the cosine function, which is used to determine the contribution of frequency components;

[0037] The identification marking unit is used to calculate the nth Mel frequency cepstral coefficient C of each short time window n The extracted features are clustered by Gaussian mixture model GMM or k-means clustering algorithm to identify audio segments of different speakers; the speaker segment is segmented by hidden Markov model HMM and Viterbi algorithm to detect the boundary of speaker switching;

[0038] Each identified speaker segment is annotated to obtain the start and end time of the jth audio segment, and a unique identifier is given to each speaker segment.

[0039] Preferably, the audio collection and translation module further includes a second calculation unit and a translation unit;

[0040] The second calculation unit is used to extract the speech rate feature of the i-th character in the j-th audio segment, and calculate the first speech rate index Wxz of the i-th character in the j-th audio segment by the following formula: i,j :

[0041]

[0042] Where W i,j represents the total number of words spoken by the i-th character in the j-th video segment; T i,j represents the total speaking time of the i-th character in the j-th video segment;

[0043] The translation unit is used to extract words by performing automatic speech recognition on the j-th audio segment to generate a word sequence, optimize the words through a trained language model, and generate a first speech translation text.

[0044] Preferably, the facial lip action recognition module includes an image processing unit, a lip region extraction unit and a third calculation unit;

[0045] The image processing unit is used to perform grayscale conversion on the first image sequence set, and perform image enhancement and noise suppression processing on the image to obtain a second image sequence set;

[0046] The lip region extraction unit is used to extract features of the lip shapes in the second image sequence set through a 68-point facial marker model in Dlib driven by AI, so as to obtain lip key point data, where the lip key point data includes key point coordinate positions;

[0047] And according to the lip key point data, the lip area in each frame image is identified by the coordinates of the new point U in the upper lip and the new point L in the lower lip, which are (xU, yU) and (xL, yL) respectively;

[0048] The vertical distance Dv between lips opening and closing is calculated by the following formula:

[0049] Dv=|yU―yL|;

[0050] When the lip opening and closing distance Dh of the current frame is greater than the opening and closing threshold X, it is marked as the lip opening and closing state;

[0051] When the lip opening and closing distance Dh of the current frame is identified to be less than or equal to the opening and closing threshold X, it is marked as the lip closed state;

[0052] The third calculation unit is used to traverse and identify all frames, and record each state change from "closed" to "open and closed", and perform statistics to obtain the number of lip opening and closing times Nv; and match the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j , the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j Calculated by the following formula:

[0053]

[0054] In the formula, Nv i,j represents the number of times the ith character's lips open and close in the jth audio segment, T i,j represents the total speaking time of the i-th character in the j-th video segment;

[0055] The second actual speaking speed index Sxz i,j and the first speech rate index Wxz i,j Perform an associated calculation to obtain the first calibration difference coefficient of the i-th character in the j-th audio segment

[0056]

[0057] Preferably, the first correction module comprises a first time calibration unit;

[0058] The first time calibration unit is used to determine the subtitle start time of the first speech translation text corresponding to each video segment. and subtitle end time And according to the first calibration difference coefficient Calculate the new time range and obtain the second start time using the following correction formula and the second end time

[0059]

[0060] Where, ΔT startIndicates the correction amount of the start time, indicating the time that needs to be adjusted due to the change in speech speed so that the subtitle time is aligned with the actual speaking time; Q1 represents the start time correction adjustment coefficient, Q2 represents the end time correction adjustment coefficient, ΔT end Indicates the correction amount of the end time; T i,j V represents the total speaking time of the i-th character in the j-th video segment; base represents the reference speech rate threshold, ΔT end Indicates the correction amount for the end time;

[0061] And the second start time and the second end time Match the second voice translation text to obtain the second voice translation text timestamp.

[0062] Preferably, the multi-speech overlap recognition module includes a voice activity monitoring unit, an overlap activity detection unit and an influence factor calculation unit;

[0063] The voice activity monitoring unit is used to use a voice activity detection (VAD) algorithm to identify the voice activity state of each time frame in the audio segment and mark each time frame as "voice activity" or non-voice activity;

[0064] The overlapping activity detection unit is used to apply multi-channel signal processing technology to each time frame, establish a deep learning recognition model through a convolutional neural network CNN or a long short-term memory network LSTM, detect whether there are multiple voice activities in each time frame, and mark the state of the overlapping voice activities in each time frame;

[0065] The impact factor calculation unit is used to count the number of frames of overlapping speech activities in each time period, identify the overlapping speech time period, and analyze the time proportion characteristics of the overlapping speech, and obtain the overlapping speech impact factor D by the following formula:

[0066]

[0067] In the formula, δ(t) indicates whether there is overlapping speech activity in the tth frame, which is 1 if there is, and 0 if not, YL indicates the volume value of the background multi-person speech in the tth frame, s(t) indicates the volume value weight of the tth frame, and T all Indicates the total number of time frames.

[0068] Preferably, the multi-speech overlap recognition module further includes an evaluation unit, which is used to preset an abnormal threshold F and compare the overlapping speech impact factor D with the abnormal threshold to obtain an evaluation result, including:

[0069] If the overlapping speech impact factor D>abnormal threshold F, it means that the overlapping speech has an abnormal impact on the timestamp, and the second correction instruction is triggered; if the overlapping speech impact factor D≤abnormal threshold, it means that the overlapping speech has a normal impact on the timestamp, and synchronization is performed based on the timestamp of the first speech translation text.

[0070] Preferably, the second correction module comprises an association unit and a second time calibration unit;

[0071] The associating unit is used to associate the overlapping speech impact factor D with the first calibration difference coefficient The second calibration coefficient is calculated by the following correlation formula

[0072]

[0073] Where q represents the weight coefficient, which is used to adjust the influence of the overlapping speech influence factor D;

[0074] The second time calibration unit is used to convert the second calibration difference coefficient Synchronously correcting the timestamp of the second voice translation text;

[0075] Extract the second start time and the second end time The third start time is calculated by the following correction formula and the third end time

[0076]

[0077] And the second start time and the second end time The timestamp of the second voice translation text is matched to obtain a third voice translation text timestamp.

[0078] An AI-driven video content subtitle synchronous translation method comprises the following steps:

[0079] Step 1: Collect the video stream that needs subtitle translation synchronization, establish a video stream set, and use a video decoder to decode the video stream into original frame image data; continuously extract frame images from the video stream according to the set frame rate of 30 frames per second, and store them as an image sequence; pre-process the extracted image sequence, including denoising, image enhancement and scaling, and store the pre-processed image sequence in the first image sequence set;

[0080] Step 2: Extract the audio stream corresponding to the video stream from the video stream set, separate the audio of different characters from the audio stream, use speaker separation technology Speaker-Diarization, use machine learning and signal processing algorithms to identify the audio segments of different characters, and perform analysis and calculation to obtain the first speech rate index Wxz of the i-th character in the j-th audio segment. i,j and first voice translated text;

[0081] Step 3: Extract lip area features from each frame of the first image sequence set, obtain lip key point data through the 68-point facial marker model in Dlib driven by AI, and extract the lip opening and closing vertical distance features and the start and end time features of each character's speech based on the lip key point data to obtain the lip opening and closing vertical distance Dv and the number of lip opening and closing times Nv, and perform deep calculation to obtain: the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j , and the second actual speaking speed index Sxz i,j and the first speech rate index Wxz i,j Perform an associated calculation to obtain the first calibration difference coefficient of the i-th character in the j-th audio segment

[0082] Step 4: Based on the first calibration difference coefficient Used to synchronously check the start time and end time of the first voice translation text to obtain the timestamp of the second voice translation text;

[0083] Step 5: extracting from the video stream set whether there are multiple voice activity segment features during the speech of the i-th character in the j-th audio segment, so as to identify and calculate the overlapping voice segments in the k-th audio segment, so as to obtain the overlapping voice impact factor D. If the overlapping voice impact factor D>abnormal threshold F, the second correction instruction is triggered;

[0084] Step 6: After receiving the second calibration instruction, the overlapping speech impact factor D and the first calibration difference coefficient The second calibration coefficient is calculated and associated The second calibration difference coefficient Perform synchronization correction on the second voice translation text timestamp to obtain the third voice translation text timestamp

[0085] The present invention provides an AI-driven video content subtitle synchronous translation method and system. It has the following beneficial effects:

[0086] (1) This AI-driven video content subtitle synchronization translation system, combined with a video frame acquisition module and a facial lip movement recognition module, can accurately obtain the vertical distance and number of times each character's lips open and close. These data are used to calculate the actual speaking speed and compared with the traditional speaking speed index to obtain the first calibration difference coefficient. This method effectively adjusts the timestamp of the subtitles, reduces the time deviation caused by the difference in speaking speed, and makes the subtitles more synchronized with the actual speech, thereby improving the accuracy of the subtitles and the audience's viewing experience.

[0087] (2) This is an AI-driven video content subtitle synchronization translation method and system. The system introduces a multi-person speech overlap recognition module. Through voice activity monitoring, deep learning recognition model (CNN or LSTM) and overlapping speech impact factor D calculation, it can accurately detect and mark speech overlap. If the overlapping speech impact factor D exceeds the abnormal threshold F, the system will trigger a second correction instruction to further calibrate the subtitle timestamp. This mechanism ensures the accuracy of subtitles in complex dialogue scenarios and avoids synchronization problems caused by speech overlap.

[0088] (3) This is an AI-driven video content subtitle synchronization translation system, which adopts a two-level correction mechanism, including a first correction module and a second correction module. The first correction module adjusts the subtitle timestamp based on the first calibration difference coefficient to cope with the time deviation caused by the change in speech rate; the second correction module performs additional correction after identifying overlapping speech. This multi-level correction method comprehensively considers the impact of speech rate calibration and overlapping speech, ensures that the time synchronization of subtitles and video content is more accurate, and improves the stability and reliability of the subtitle system in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0089] Figure 1 A schematic diagram of the AI-driven video content subtitle synchronization translation method and system structure of the present invention; DETAILED DESCRIPTION

[0090] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0091] Example 1

[0092] See also Figure 1 , the present invention provides an AI-driven video content subtitle synchronous translation system, including a video frame acquisition module, an audio acquisition and translation module, a facial lip action recognition module, a first correction module, a multi-person speech overlap recognition module, and a second correction module;

[0093] The video frame acquisition module is used to acquire the video stream that needs subtitle translation synchronization, and establish a video stream set and a first image sequence set;

[0094] The audio acquisition and translation module is used to extract the audio stream corresponding to the video stream from the video stream set, separate the audio of different characters from the audio stream, use the speaker separation technology Speaker-Diarization, use machine learning and signal processing algorithms to identify the audio segments of different characters, and perform analysis and calculation to obtain the first speech rate index Wxz of the i-th character in the j-th audio segment i,j and first voice translated text;

[0095] The facial lip action recognition module is used to extract lip area features from each frame of the first image sequence set, obtain lip key point data through the 68-point facial marker model in Dlib driven by AI, and extract the lip opening and closing vertical distance features and the start and end time features of each character's speech based on the lip key point data to obtain the lip opening and closing vertical distance Dv and the lip opening and closing times Nv, and perform deep calculation to obtain: the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j , and the second actual speaking speed index Sxz i,j and the first speech rate index Wxz i,j Perform an associated calculation to obtain the first calibration difference coefficient of the i-th character in the j-th audio segment

[0096] The first correction module is used to adjust the first calibration difference coefficient Used to synchronously check the start time and end time of the first voice translation text to obtain the timestamp of the second voice translation text;

[0097] The multi-person speech overlap recognition module is used to extract whether there are multiple voice activity segment features in the jth audio segment of the video stream set during the speech of the i-th character, so as to identify and calculate the overlapping speech segments in the k-th audio segment to obtain the overlapping speech impact factor D. If the overlapping speech impact factor D is greater than the abnormal threshold F, the second correction instruction is triggered;

[0098] The second correction module is used to, after receiving the second correction instruction, correct the overlapped speech impact factor D and the first calibration difference coefficient The second calibration coefficient is calculated and associated The second calibration difference coefficient The second speech translation text timestamp is synchronously corrected to obtain a third speech translation text timestamp.

[0099] In this embodiment, the system combines the video frame acquisition module and the facial lip action recognition module to accurately obtain the vertical distance and number of times each character's lips open and close. These data are used to calculate the actual speaking speed and compare it with the traditional speaking speed index to obtain the first calibration difference coefficient. This method can effectively adjust the timestamps of subtitles, reduce the time deviation caused by differences in speaking speed, and make the subtitles more synchronized with the actual speech.

[0100] By introducing a multi-speaking overlapping recognition module, the system can identify and analyze overlapping speech segments in video streams. Using voice activity monitoring, deep learning recognition models (CNN or LSTM) and overlapping speech impact factor D calculation, the system can accurately detect and mark speech overlap. If the overlapping speech impact factor D exceeds the abnormal threshold F, the system will trigger a second correction instruction to further calibrate the subtitle timestamp to ensure the accuracy of subtitles in complex dialogue scenarios. The system's two-level correction mechanism (first correction module and second correction module) combines speech rate calibration and overlapping speech impact analysis. The first correction module is based on the first calibration difference coefficient The subtitle timestamp is adjusted, and the second correction module performs additional correction after recognizing overlapping speech. This multi-level correction method ensures that the subtitles are more accurately synchronized with the video content.

[0101] Example 2, this example is explained in Example 1, please refer to Figure 1 ,Specifically, the video frame acquisition module includes a first acquisition unit and a first preprocessing unit;

[0102] The first acquisition unit is used to acquire a video stream that needs subtitle translation synchronization, and use a video decoder to decode the video stream into original frame image data; according to a set frame rate of 30 frames per second, continuously extract frame images from the video stream and store them as an image sequence;

[0103] The first preprocessing unit is used to preprocess the extracted image sequence, including denoising, image enhancement and scaling, and store the preprocessed image sequence into the first image sequence set.

[0104] In this embodiment, the design of the first acquisition unit and the first preprocessing unit ensures efficient acquisition and processing of video frames, and provides high-quality input data for the video content subtitle synchronization translation system. Through accurate frame acquisition and image preprocessing, the system can better support subsequent subtitle synchronization correction work, improve the accuracy of subtitle display and the audience's viewing experience.

[0105] Example 3, this example is explained in Example 1, please refer to Figure 1,Specifically, the audio acquisition and translation module includes a second acquisition unit, a second processing unit, an AI calculation unit and a recognition marking unit;

[0106] The second acquisition unit is used to extract an audio stream from the video data set through an audio decoder, and decode the audio stream into an original audio signal;

[0107] The second processing unit is used to remove background noise from the original audio signal by using a Wiener filtering method, and to adjust the gain of the original audio signal;

[0108] The AI ​​calculation unit is used to divide the processed original audio signal into several frames, and calculate each short-time window to obtain the nth Mel frequency cepstral coefficient C n , the nth Mel frequency cepstrum coefficient C n The calculation method includes the following steps:

[0109] S1. Pre-emphasize the original audio signal, and obtain the pre-emphasized signal value x(n) of the nth sampling point by the following formula:

[0110]

[0111] In the formula, x(n) represents the amplitude of the original audio signal at the nth sampling point, x(n-1) represents the amplitude of the original audio signal at the n-1th sampling point, α represents the pre-emphasis coefficient, and the setting value is 0.97;

[0112] S2, divide the original audio signal after pre-emphasis into several frames, each frame is N sampling points long, the frames are set to overlap by 50%, and the pre-emphasis signal value of the nth sampling point is Multiply by the window function w(n), and obtain the windowed signal value y(n) of the nth sampling point by the following formula:

[0113]

[0114] Where w(n) represents the window function, and n represents the time index in the current window;

[0115] S3, and perform Fourier transformation on the windowed signal value y(n) of the nth sampling point to obtain the frequency domain signal X(k):

[0116]

[0117] Where X(k) represents the frequency signal, which represents the value at the frequency index k; N represents the length of the frame, that is, the number of sampling points contained in each frame; k represents the index value of the frequency component, ranging from 0 to N-1; p represents the imaginary unit, satisfying p 2= -1, π represents the ratio of pi, and the set value is 3.14159; each time sampling point y(n) is compared with a complex exponential function Multiply, kn is the product of frequency index k and time index n, the complex exponential function represents a complex rotation, and the frequency is proportional to k; this step converts the discrete signal in the time domain into a discrete signal in the frequency domain;

[0118] S4, and obtain the power spectrum P(k) according to the frequency domain signal X(k) by the following fast Fourier transform (FFT) formula:

[0119]

[0120] In the formula, |X(k)| represents the complex amplitude of the frequency domain signal, which reflects the strength of the signal at a specific frequency k;

[0121] S5. Pass the power spectrum P(k) through a set of triangular filters on the Mel scale to obtain the filtered energy to obtain the output value M of the mth filter. m :

[0122]

[0123] In the formula, H m (k) represents the mth filter response value in the Mel filter bank, and is calculated by the output value M of the mth filter. m After logarithmic transformation, the logarithmic energy value log(M m );

[0124] S6, and the logarithmic energy value log(M m ) is converted to the nth Mel frequency cepstrum coefficient C by the following formula n :

[0125]

[0126] Where M represents the number of filters; Represents the phase angle in the cosine function, which is used to determine the contribution of the frequency component; πn is used to represent the nth Mel frequency cepstrum coefficient C n The phase shift.

[0127] The identification marking unit is used to calculate the nth Mel frequency cepstral coefficient C of each short time window n The extracted features are clustered by Gaussian mixture model GMM or k-means clustering algorithm to identify audio segments of different speakers; the speaker segment is segmented by hidden Markov model HMM and Viterbi algorithm to detect the boundary of speaker switching;

[0128] Each identified speaker segment is annotated to obtain the start and end time of the jth audio segment, and a unique identifier is given to each speaker segment.

[0129] In this embodiment, the Mel frequency cepstral coefficient (MFCC) is calculated by calculating the nth Mel frequency cepstral coefficient C n , the AI ​​computing unit can extract key features of the audio signal, which help to accurately describe the speech content in the audio and capture the spectral characteristics of the speech signal. The use of pre-emphasis and window functions can significantly reduce noise and interference in the audio signal, enhance the clarity of the signal, and provide more stable feature data for subsequent processing. The Fourier transform converts the time domain signal into a frequency domain signal, allowing the AI ​​computing unit to accurately analyze the frequency components of the audio. The power spectrum calculation further provides intensity information of the frequency components, which helps to extract features accurately. Through the Mel-scale filter bank, the AI ​​computing unit can convert the frequency domain signal into an energy feature on the Mel scale, which is more in line with the perceptual characteristics of human hearing and improves the perception of features. Use these clustering algorithms to extract the nth Mel-frequency cepstrum coefficient C n Clustering helps identify audio segments of different speakers. By analyzing the similarity of audio features, the AI ​​computing unit can effectively distinguish different speakers. The HMM and Viterbi algorithms are applied to segment speaker segments, accurately detect the boundaries of speaker switching, and improve the processing capabilities of multi-speaker scenarios.

[0130] Example 4, this example is explained in Example 1, please refer to Figure 1 ,Specifically, the audio acquisition and translation module also includes a second computing unit and a translation unit;

[0131] The second calculation unit is used to extract the speech rate feature of the i-th character in the j-th audio segment, and calculate the first speech rate index Wxz of the i-th character in the j-th audio segment by the following formula: i,j :

[0132]

[0133] Where W i,j represents the total number of words spoken by the i-th character in the j-th video segment; T i,j represents the total speaking time of the i-th character in the j-th video segment;

[0134] The translation unit is used to extract words by performing automatic speech recognition on the j-th audio segment to generate a word sequence, optimize the words through a trained language model, and generate a first speech translation text.

[0135] In this embodiment, by analyzing the total number of words and total duration of the i-th character in the j-th audio segment, the second calculation unit can accurately calculate the speech rate feature. This feature provides information about the character's speaking speed, which helps to accurately synchronize the display time of subtitles and reduce the situation where subtitles lag or advance.

[0136] The translation unit extracts words from the audio through automatic speech recognition technology, and optimizes the words using the trained language model to generate high-quality first-speech translation text. This process not only improves the accuracy of the translation, but also enhances the ability to adapt to complex language environments.

[0137] Example 5. This example is explained in Example 1. Please refer to Figure 1 ,Specifically, the facial lip action recognition module includes an image processing unit, a lip area extraction unit and a third computing unit;

[0138] The image processing unit is used to perform grayscale conversion on the first image sequence set, and perform image enhancement and noise suppression processing to obtain a second image sequence set; this can reduce the impact of background interference and noise on the detection result and ensure recognition accuracy.

[0139] The lip area extraction unit is used to extract features of the lip shapes in the second image sequence set through the 68-point facial marker model in the AI-driven Dlib to obtain lip key point data, which includes key point coordinate positions; this process accurately identifies the shape and movement of the lips, providing reliable data support for subsequent lip opening and closing state analysis.

[0140] And according to the lip key point data, the lip area in each frame image is identified by the coordinates of the new point U in the upper lip and the new point L in the lower lip, which are (xU, yU) and (xL, yL) respectively;

[0141] The vertical distance Dv between lips opening and closing is calculated by the following formula:

[0142] Dv=|yU―yL|;

[0143] When the lip opening and closing distance Dh of the current frame is greater than the opening and closing threshold X, it is marked as the lip opening and closing state;

[0144] When the lip opening and closing distance Dh of the current frame is identified to be less than or equal to the opening and closing threshold X, it is marked as the lip closed state;

[0145] The third calculation unit is used to traverse and identify all frames, and record each state change from "closed" to "open and closed", and perform statistics to obtain the number of lip opening and closing times Nv; and match the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j, the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j Calculated by the following formula:

[0146]

[0147] In the formula, Nv i,j represents the number of times the ith character's lips open and close in the jth audio segment, T i,j represents the total speaking time of the i-th character in the j-th video segment;

[0148] The second actual speaking speed index Sxz i,j and the first speech rate index Wxz i,j Perform an associated calculation to obtain the first calibration difference coefficient of the i-th character in the j-th audio segment

[0149]

[0150] In this embodiment, by calculating the vertical distance Dv between the lips and the opening and closing threshold X, it is possible to effectively determine whether the lips are in an open or closed state. The third calculation unit records the number of state changes from "closed" to "open or closed", and counts the number of lip opening and closing times Nv, thereby providing key dynamic data for speech rate analysis. The third calculation unit calculates the second actual speech rate index Sxz by combining the number of lip opening and closing times Nv and the total speaking time. i,j This index takes into account the actual situation of lip movements, provides more accurate speech rate data, and is compared with the first speech rate index to calculate the first calibration difference coefficient Helps improve subtitle synchronization accuracy.

[0151] Example 6, this example is explained in Example 1, please refer to Figure 1 ,Specifically, the first correction module includes a first time calibration unit;

[0152] The first time calibration unit is used to determine the subtitle start time of the first speech translation text corresponding to each video segment. and subtitle end time And according to the first calibration difference coefficient Calculate the new time range and obtain the second start time using the following correction formula and the second end time

[0153]

[0154] Where, ΔT startIndicates the correction amount of the start time, indicating the time that needs to be adjusted due to the change in speech speed so that the subtitle time is aligned with the actual speaking time; Q1 represents the start time correction adjustment coefficient, Q2 represents the end time correction adjustment coefficient, ΔT end Indicates the correction amount of the end time; T i,j V represents the total speaking time of the i-th character in the j-th video segment; base represents the reference speech rate threshold, ΔT end Indicates the correction amount for the end time;

[0155] And the second start time and the second end time Match the second voice translation text to obtain the second voice translation text timestamp.

[0156] In this embodiment, the first time calibration unit first determines the subtitle start time of the first speech translation text corresponding to it. and subtitle end time And according to the first calibration difference coefficient Calculate the new time range to get the second start time and the second end time The start time and end time are adjusted through the correction formula to adapt to the change of speech speed. The specific correction amount calculation includes the start time correction amount and the end time correction amount, which helps to dynamically adjust the time deviation caused by the difference in speech speed. and the second end time The calculation result can be accurately matched with the second voice translation text, improving the synchronization accuracy of subtitles and video content. This optimization can effectively reduce the phenomenon of subtitles being advanced or delayed, thereby improving the audience's viewing experience. Using the correction adjustment coefficients Q1 and Q2, the start time and end time are flexibly adjusted to cope with the changes in speaking speed in different video segments. This method enhances the system's adaptability in handling scenarios with different speaking speeds and improves the accuracy of overall subtitle synchronization.

[0157] Example 7, this example is explained in Example 1, please refer to Figure 1 ,Specifically, the multi-speaker overlap recognition module includes a voice activity monitoring unit, an overlapping activity detection unit and an influence factor calculation unit;

[0158] The voice activity monitoring unit is used to use a voice activity detection (VAD) algorithm to identify the voice activity state of each time frame in the audio segment and mark each time frame as "voice activity" or non-voice activity;

[0159] The overlapping activity detection unit is used to apply multi-channel signal processing technology to each time frame, establish a deep learning recognition model through a convolutional neural network CNN or a long short-term memory network LSTM, detect whether there are multiple voice activities in each time frame, and mark the state of the overlapping voice activities in each time frame;

[0160] The impact factor calculation unit is used to count the number of frames of overlapping speech activities in each time period, identify the overlapping speech time period, and analyze the time proportion characteristics of the overlapping speech, and obtain the overlapping speech impact factor D by the following formula:

[0161]

[0162] In the formula, δ(t) indicates whether there is overlapping speech activity in the tth frame, which is 1 if there is, and 0 if not, YL indicates the volume value of the background multi-person speech in the tth frame, s(t) indicates the volume value weight of the tth frame, and T all Indicates the total number of time frames.

[0163] In this embodiment, the impact factor calculation unit calculates the impact factor D of the overlapping speech by counting the number of frames of overlapping speech activities, identifying the overlapping speech time period, and analyzing the time proportion characteristics of the overlapping speech. The calculation of this factor comprehensively considers factors such as the existence state, volume value and time frame number of the overlapping speech, provides a quantitative analysis of the impact of the overlapping speech, and helps to further optimize the accuracy of subtitle synchronization.

[0164] Example 8, this example is explained in Example 1, please refer to Figure 1 Specifically, the multi-speech overlap recognition module further includes an evaluation unit, which is used to preset an abnormal threshold F and compare the overlapping speech impact factor D with the abnormal threshold to obtain an evaluation result, including:

[0165] If the overlapping speech impact factor D>abnormal threshold F, it means that the overlapping speech has an abnormal impact on the timestamp, and the second correction instruction is triggered; if the overlapping speech impact factor D≤abnormal threshold, it means that the overlapping speech has a normal impact on the timestamp, and synchronization is performed based on the timestamp of the first speech translation text.

[0166] In the present embodiment, the evaluation unit compares the overlapping voice impact factor D by a preset abnormal threshold F, and can intelligently detect the abnormal impact of overlapping voice on the subtitle timestamp. This abnormal detection mechanism can automatically identify the situation where the overlapping voice has a greater interference with the subtitle synchronization, so as to make adjustments in time. When the overlapping voice impact factor D exceeds the abnormal threshold F, the evaluation unit automatically triggers the second correction instruction. This automated correction mechanism can effectively deal with the timestamp deviation caused by voice overlap, ensuring that the synchronization of subtitles and video content is more accurate. If the overlapping voice impact factor D does not exceed the abnormal threshold, the evaluation unit confirms that the impact of the overlapping voice on the timestamp is within the normal range, and the system will synchronize based on the first voice translation text timestamp. This ensures that under normal circumstances, the synchronization accuracy of the system will not be affected by the overlapping voice, thereby maintaining the consistency and accuracy of the subtitle translation.

[0167] Example 9. This example is explained in Example 1. Figure 1 ,Specifically, the second correction module includes an association unit and a second time calibration unit;

[0168] The associating unit is used to associate the overlapping speech impact factor D with the first calibration difference coefficient The second calibration coefficient is calculated by the following correlation formula

[0169]

[0170] Where q represents the weight coefficient, which is used to adjust the influence of the overlapping speech influence factor D;

[0171] The second time calibration unit is used to convert the second calibration difference coefficient Synchronously correcting the timestamp of the second voice translation text;

[0172] Extract the second start time and the second end time The third start time is calculated by the following correction formula and the third end time

[0173]

[0174] And the second start time and the second end time The timestamp of the second voice translation text is matched to obtain a third voice translation text timestamp.

[0175] In this embodiment, the correlation unit calculates the second calibration coefficient to correlate the overlapping speech impact factor D with the first calibration difference coefficient, thereby obtaining a more accurate second calibration coefficient. The second time calibration unit synchronizes and corrects the timestamp of the second speech translation text according to the second calibration coefficient. In this way, the timestamp of the subtitles can be dynamically adjusted to ensure that the time synchronization between the subtitles and the actual speech is more accurate under the influence of overlapping speech. This method not only takes into account the influence of overlapping speech, but also combines the first calibration difference coefficient. The accuracy of the timestamp is further improved. The third timestamp of the voice translation text is obtained by matching the second start time and the second end time with the second voice translation text timestamp. Such a comprehensive correction scheme ensures the time synchronization accuracy of the subtitle translation text in complex scenarios and improves the consistency and accuracy of the overall subtitle system.

[0176] An AI-driven video content subtitle synchronous translation method comprises the following steps:

[0177] Step 1: Collect the video stream that needs subtitle translation synchronization, establish a video stream set, and use a video decoder to decode the video stream into original frame image data; continuously extract frame images from the video stream according to the set frame rate of 30 frames per second, and store them as an image sequence; pre-process the extracted image sequence, including denoising, image enhancement and scaling, and store the pre-processed image sequence in the first image sequence set;

[0178] Step 2: Extract the audio stream corresponding to the video stream from the video stream set, separate the audio of different characters from the audio stream, use speaker separation technology Speaker-Diarization, use machine learning and signal processing algorithms to identify the audio segments of different characters, and perform analysis and calculation to obtain the first speech rate index Wxz of the i-th character in the j-th audio segment. i,j and first voice translated text;

[0179] Step 3: Extract lip area features from each frame of the first image sequence set, obtain lip key point data through the 68-point facial marker model in Dlib driven by AI, and extract the lip opening and closing vertical distance features and the start and end time features of each character's speech based on the lip key point data to obtain the lip opening and closing vertical distance Dv and the number of lip opening and closing times Nv, and perform deep calculation to obtain: the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j , and the second actual speaking speed index Sxz i,j and the first speech rate index Wxz i,j Perform an associated calculation to obtain the first calibration difference coefficient of the i-th character in the j-th audio segment

[0180] Step 4: Based on the first calibration difference coefficient Used to synchronously check the start time and end time of the first voice translation text to obtain the timestamp of the second voice translation text;

[0181] Step 5: extracting from the video stream set whether there are multiple voice activity segment features during the speech of the i-th character in the j-th audio segment, so as to identify and calculate the overlapping voice segments in the k-th audio segment, so as to obtain the overlapping voice impact factor D. If the overlapping voice impact factor D>abnormal threshold F, the second correction instruction is triggered;

[0182] Step 6: After receiving the second calibration instruction, the overlapping speech impact factor D and the first calibration difference coefficient The second calibration coefficient is calculated and associated The second calibration difference coefficient The second speech translation text timestamp is synchronously corrected to obtain a third speech translation text timestamp.

[0183] The threshold is set to facilitate comparison. The size of the threshold depends on the amount of sample data and the number of bases set by technicians in this field for each group of sample data; as long as it does not affect the proportional relationship between the parameter and the quantized value.

[0184] The above formulas are obtained by collecting a large amount of data for software simulation and selecting a formula that is close to the actual value. The coefficients in the formula are set by technical personnel in this field according to actual conditions. The above is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited to this. Any technical personnel familiar with the technical field within the technical scope disclosed by the present invention, according to the technical solution and the inventive concept of the present invention, make equivalent replacement or change, which should be covered within the protection scope of the present invention.

Claims

1. An AI-driven video content subtitle synchronous translation system, characterized by: It includes a video frame acquisition module, an audio acquisition and translation module, a facial lip action recognition module, a first correction module, a multi-person speech overlap recognition module, and a second correction module; The video frame acquisition module is used to acquire the video stream that needs subtitle translation synchronization, and establish a video stream set and a first image sequence set; The audio acquisition and translation module is used to extract the audio stream corresponding to the video stream from the video stream set, separate the audio of different characters from the audio stream, use the speaker separation technology Speaker-Diarization, use machine learning and signal processing algorithms to identify the audio segments of different characters, and perform analysis and calculation to obtain the first speech rate index Wxz of the i-th character in the j-th audio segment i,j and first voice translated text; The facial lip action recognition module is used to extract lip area features from each frame of the first image sequence set, obtain lip key point data through the 68-point facial marker model in Dlib driven by AI, and extract the lip opening and closing vertical distance features and the start and end time features of each character's speech based on the lip key point data to obtain the lip opening and closing vertical distance Dv and the lip opening and closing times Nv, and perform deep calculation to obtain: the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j , and the second actual speaking speed index Sxz i,j and the first speech rate index Wxz i,j Perform an associated calculation to obtain the first calibration difference coefficient of the i-th character in the j-th audio segment The first correction module is used to adjust the first calibration difference coefficient Used to synchronously check the start time and end time of the first voice translation text to obtain the timestamp of the second voice translation text; The multi-person speech overlap recognition module is used to extract whether there are multiple voice activity segment features in the jth audio segment of the video stream set during the speech of the i-th character, so as to identify and calculate the overlapping speech segments in the k-th audio segment to obtain the overlapping speech impact factor D. If the overlapping speech impact factor D is greater than the abnormal threshold F, the second correction instruction is triggered; The second correction module is used to, after receiving the second correction instruction, correct the overlapped speech impact factor D and the first calibration difference coefficient The second calibration coefficient is calculated and associated The second calibration difference coefficient The second speech translation text timestamp is synchronously corrected to obtain a third speech translation text timestamp.

2. The AI-driven video content subtitle synchronous translation system according to claim 1, characterized in that: The video frame acquisition module includes a first acquisition unit and a first preprocessing unit; The first acquisition unit is used to acquire a video stream that needs subtitle translation synchronization, and use a video decoder to decode the video stream into original frame image data; according to a set frame rate of 30 frames per second, continuously extract frame images from the video stream and store them as an image sequence; The first preprocessing unit is used to preprocess the extracted image sequence, including denoising, image enhancement and scaling, and store the preprocessed image sequence into the first image sequence set.

3. The AI-driven video content subtitle synchronous translation system according to claim 1, characterized in that: The audio collection and translation module includes a second collection unit, a second processing unit, an AI calculation unit and a recognition marking unit; The second acquisition unit is used to extract an audio stream from the video data set through an audio decoder, and decode the audio stream into an original audio signal; The second processing unit is used to remove background noise from the original audio signal by using a Wiener filtering method, and to adjust the gain of the original audio signal; The AI ​​calculation unit is used to divide the processed original audio signal into several frames, and calculate each short-time window to obtain the nth Mel frequency cepstral coefficient C n , the nth Mel frequency cepstrum coefficient C n The calculation method includes the following steps: S1. Pre-emphasize the original audio signal and calculate the pre-emphasized signal value of the nth sampling point using the following formula In the formula, x(n) represents the amplitude of the original audio signal at the nth sampling point, x(n-1) represents the amplitude of the original audio signal at the n-1th sampling point, α represents the pre-emphasis coefficient, and the setting value is 0.97; S2, divide the original audio signal after pre-emphasis into several frames, each frame is N sampling points long, the frames are set to overlap by 50%, and the pre-emphasis signal value of the nth sampling point is Multiply by the window function w(n), and obtain the windowed signal value y(n)- of the nth sampling point by the following formula: Where w(n) represents the window function, and n represents the time index in the current window; S3, and perform Fourier transformation on the windowed signal value y(n) of the nth sampling point to obtain the frequency domain signal X(k): In the formula, X(k) represents the frequency signal, which represents the value at the frequency index k, N represents the length of the frame, that is, the number of sampling points contained in each frame; k represents the index value of the frequency component, ranging from 0 to N-1; p represents the imaginary unit, satisfying p 2 = -1, π represents the ratio of pi, and the set value is 3.14159; each time sampling point y(n) is compared with a complex exponential function Multiply, kn is the product of frequency index k and time index n, the complex exponential function represents a complex rotation, and the frequency is proportional to k; this step converts the discrete signal in the time domain into a discrete signal in the frequency domain; S4, and obtain the power spectrum P(k) according to the frequency domain signal X(k) by the following fast Fourier transform FFT formula: In the formula, |X(k)| represents the complex amplitude of the frequency domain signal, which reflects the strength of the signal at a specific frequency k; S5. Pass the power spectrum P(k) through a set of triangular filters on the Mel scale to obtain the filtered energy to obtain the output value M of the mth filter. m : In the formula, H m (k) represents the mth filter response value in the Mel filter bank, and is calculated by the output value M of the mth filter. m After logarithmic transformation, the logarithmic energy value log(M m ); S6, and the logarithmic energy value log(M m ) is converted to the nth Mel frequency cepstrum coefficient C by the following formula n : Where M represents the number of filters, Represents the phase angle in the cosine function, which is used to determine the contribution of frequency components; The identification marking unit is used to calculate the nth Mel frequency cepstral coefficient C of each short time window n The extracted features are clustered by Gaussian mixture model GMM or k-means clustering algorithm to identify audio segments of different speakers; the speaker segment is segmented by hidden Markov model HMM and Viterbi algorithm to detect the boundary of speaker switching; Each identified speaker segment is annotated to obtain the start and end time of the jth audio segment, and a unique identifier is given to each speaker segment.

4. The AI-driven video content subtitle synchronous translation system according to claim 3, characterized in that: The audio collection and translation module also includes a second calculation unit and a translation unit; The second calculation unit is used to extract the speech rate feature of the i-th character in the j-th audio segment, and calculate the first speech rate index Wxz of the i-th character in the j-th audio segment by the following formula: i,j : Where W i,j represents the total number of words spoken by the i-th character in the j-th video segment; T i,j represents the total speaking time of the i-th character in the j-th video segment; The translation unit is used to extract words by performing automatic speech recognition on the j-th audio segment to generate a word sequence, optimize the words through a trained language model, and generate a first speech translation text.

5. The AI-driven video content subtitle synchronous translation system according to claim 1, characterized in that: The facial lip action recognition module includes an image processing unit, a lip area extraction unit and a third calculation unit; The image processing unit is used to perform grayscale conversion on the first image sequence set, and perform image enhancement and noise suppression processing on the image to obtain a second image sequence set; The lip region extraction unit is used to extract features of the lip shapes in the second image sequence set through the 68-point facial marker model in Dlib driven by AI, so as to obtain lip key point data, the lip key point data including key point coordinate positions, so as to calculate and obtain the lip opening and closing vertical distance Dv: And according to the lip key point data, the lip area in each frame image is identified by the coordinates of the new point U in the upper lip and the new point L in the lower lip, which are (xU, yU) and (xL, yL) respectively; The lip opening and closing vertical distance Dv is calculated by the following formula: Dv=|yU―yL|; When the lip opening and closing distance Dh of the current frame is greater than the opening and closing threshold X, it is marked as the lip opening and closing state; When the lip opening and closing distance Dh of the current frame is identified to be less than or equal to the opening and closing threshold X, it is marked as the lip closed state; The third calculation unit is used to traverse and identify all frames, and record each state change from "closed" to "open and close", and perform statistics to obtain the number of lip opening and closing times Nv; and match the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j , the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j Calculated by the following formula: In the formula, Nv i,j represents the number of times the ith character's lips open and close in the jth audio segment, T i,j represents the total speaking time of the i-th character in the j-th video segment; The second actual speaking speed index Sxz i,j and the first speech rate index Wxz i,j Perform an associated calculation to obtain the first calibration difference coefficient of the i-th character in the j-th audio segment 6. The AI-driven video content subtitle synchronous translation system according to claim 5, characterized in that: The first correction module includes a first time calibration unit; The first time calibration unit is used to determine the subtitle start time of the first speech translation text corresponding to each video segment. and subtitle end time And according to the first calibration difference coefficient Calculate the new time range and obtain the second start time using the following correction formula and the second end time Where, ΔT start Indicates the correction amount of the start time, indicating the time that needs to be adjusted due to the change in speech speed so that the subtitle time is aligned with the actual speaking time; Q1 represents the start time correction adjustment coefficient, Q2 represents the end time correction adjustment coefficient, ΔT end Indicates the correction amount of the end time; T i,j V represents the total speaking time of the i-th character in the j-th video segment; base represents the reference speech rate threshold, ΔT end Indicates the correction amount for the end time; And the second start time and the second end time Match the second voice translation text to obtain the second voice translation text timestamp.

7. The AI-driven video content subtitle synchronous translation system according to claim 6, characterized in that: The multi-speaking overlapping recognition module includes a voice activity monitoring unit, an overlapping activity detection unit and an influence factor calculation unit; The voice activity monitoring unit is used to use a voice activity detection (VAD) algorithm to identify the voice activity state of each time frame in the audio segment and mark each time frame as "voice activity" or non-voice activity; The overlapping activity detection unit is used to apply multi-channel signal processing technology to each time frame, establish a deep learning recognition model through a convolutional neural network CNN or a long short-term memory network LSTM, detect whether there are multiple voice activities in each time frame, and mark the state of the overlapping voice activities in each time frame; The impact factor calculation unit is used to count the number of frames of overlapping speech activities in each time period, identify the overlapping speech time period, and analyze the time proportion characteristics of the overlapping speech, and obtain the overlapping speech impact factor D by the following formula: In the formula, δ(t) indicates whether there is overlapping speech activity in the tth frame, which is 1 if there is, and 0 if not, YL indicates the volume value of the background multi-person speech in the tth frame, s(t) indicates the volume value weight of the tth frame, and T all Indicates the total number of time frames.

8. The AI-driven video content subtitle synchronous translation system according to claim 7, characterized in that: The multi-speech overlap recognition module further includes an evaluation unit, which is used to preset an abnormal threshold F and compare the overlapping speech impact factor D with the abnormal threshold to obtain an evaluation result, including: If the overlapping speech impact factor D>abnormal threshold F, it means that the overlapping speech has an abnormal impact on the timestamp, and the second correction instruction is triggered; if the overlapping speech impact factor D≤abnormal threshold, it means that the overlapping speech has a normal impact on the timestamp, and synchronization is performed based on the timestamp of the first speech translation text.

9. The AI-driven video content subtitle synchronous translation system according to claim 8, characterized in that: The second correction module includes an association unit and a second time calibration unit; The associating unit is used to associate the overlapping speech impact factor D with the first calibration difference coefficient The second calibration coefficient is calculated by the following correlation formula Where q represents the weight coefficient, which is used to adjust the influence of the overlapping speech influence factor D; The second time calibration unit is used to convert the second calibration difference coefficient Synchronously correcting the timestamp of the second voice translation text; Extract the second start time and the second end time The third start time is calculated by the following correction formula and the third end time And the second start time and the second end time The timestamp of the second voice translation text is matched to obtain a third voice translation text timestamp.

10. An AI-driven video content subtitle synchronous translation method, applied to an AI-driven video content subtitle synchronous translation system according to any one of claims 1 to 9, characterized in that: The following steps are involved: Step 1: Collect the video stream that needs subtitle translation synchronization, establish a video stream set, and use a video decoder to decode the video stream into original frame image data; continuously extract frame images from the video stream according to the set frame rate of 30 frames per second, and store them as an image sequence; pre-process the extracted image sequence, including denoising, image enhancement and scaling, and store the pre-processed image sequence in the first image sequence set; Step 2: Extract the audio stream corresponding to the video stream from the video stream set, separate the audio of different characters from the audio stream, use speaker separation technology Speaker-Diarization, use machine learning and signal processing algorithms to identify the audio segments of different characters, and perform analysis and calculation to obtain the first speech rate index Wxz of the i-th character in the j-th audio segment. i,j and first voice translated text; Step 3: Extract lip area features from each frame of the first image sequence set, obtain lip key point data through the 68-point facial marker model in Dlib driven by AI, and extract the lip opening and closing vertical distance features and the start and end time features of each character's speech based on the lip key point data to obtain the lip opening and closing vertical distance Dv and the number of lip opening and closing times Nv, and perform deep calculation to obtain: the second actual speaking speed index Sxz of the i-th character in the j-th audio segment i,j , and the second actual speaking speed index Sxz i,j and the first speech rate index Wxz i,j Perform an associated calculation to obtain the first calibration difference coefficient of the i-th character in the j-th audio segment Step 4: Based on the first calibration difference coefficient Used to synchronously check the start time and end time of the first voice translation text to obtain the timestamp of the second voice translation text; Step 5: extracting from the video stream set whether there are multiple voice activity segment features during the speech of the i-th character in the j-th audio segment, so as to identify and calculate the overlapping voice segments in the k-th audio segment, so as to obtain the overlapping voice impact factor D. If the overlapping voice impact factor D>abnormal threshold F, the second correction instruction is triggered; Step 6: After receiving the second calibration instruction, the overlapping speech impact factor D and the first calibration difference coefficient The second calibration coefficient is calculated and associated The second calibration difference coefficient The second speech translation text timestamp is synchronously corrected to obtain a third speech translation text timestamp.

Citation Information

Cited By

  • Virtual digital human multimedia teaching interaction method and system and storage medium

    CN121888060A

  • A virtual digital human multimedia teaching interaction method and system and a storage medium

    CN121888060B

  • Sound cloning method and system, electronic equipment and medium

    CN121983069A