Method and device for automatically splitting audio of original film for film dubbing

CN122658338APending Publication Date: 2026-08-28CHINA FILM IND GROUP CO LTD BEIJING ARTIFICIAL INTELLIGENCE RESEARCH & APPLICATION BRANCH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610698309.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-20
Publication Date
2026-08-28

AI Technical Summary

Technical Problem

[0004]然而,这种基于固定频带滤波的方案存在明显的技术缺陷

Benefits of technology

首先,通过调取原始影片承载的原始复合音轨数据并对原始复合音轨数据进行基于重叠采样规则的时域切分,以生成时域音频分帧序列;根据时域音频分帧序列构建声学解析初始表征数据表,以及将声学解析初始表征数据表作为输入激励注入至预先训练的声学源深度解耦模型中进行处理,以生成对应的时频域声学掩码权重序列;基于时频域声学掩码权重序列与声学解析初始表征数据表,生成背景声环境能量特征参数序列,解决了传统方案中依靠带通滤波器组按照固定频率区间进行截断分离导致无法处理频率混叠的技术问题。相较于传统手段仅能对固定频带进行粗放处理,本方案通过构建声学解析初始表征数据表并利用预训练的深度解耦模型生成时频域声学掩码权重序列,实现了在时频网格层级对人声与背景声特征的非线性解耦。这种处理方式能够一目了然地识别出重叠频段中的声学成分归属,使得分离出的分轨音频具有较高的纯度,且背景声环境能量特征参数序列中的残留噪声水平较低,缓解了传统方案由于频谱动态重叠导致的音频分离不彻底的瓶颈。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122658338A_ABST
    Figure CN122658338A_ABST
Patent Text Reader

Abstract

The application provides a film dubbing-oriented original film audio automated split processing method and device. The method calls original composite audio track data and performs overlap sampling division to generate a time-domain audio frame sequence and construct an acoustic analysis initial representation data table; injects the same into an acoustic source deep decoupling model to generate a time-frequency domain acoustic mask weight sequence, and combines the initial representation data table to generate a background sound environment energy feature parameter sequence; constructs a role-identified voice split attribute mapping detail table for recording voice attribution; and generates automated film dubbing split audio control instruction stream data for dubbing system analysis based on the mapping detail table and the background sound sequence. Through deep decoupling and role identification mapping, the application solves the technical defects of low audio split purity and inability to realize role-level automatic attribution in traditional schemes, and significantly improves the automation degree and audio track processing quality of film dubbing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of multimedia data processing and speech signal processing technology, and more specifically, to a method and apparatus for automated track splitting of original audio for film dubbing. Background Technology

[0002] With the deepening integration of the global film and television industry, the demand for cross-language translation of films is increasing. High-quality film dubbing not only requires accurate translation but also meticulous separation of the original audio tracks to achieve complete separation of vocals, background music, and ambient sounds. In the film dubbing process, efficient and high-purity audio track separation is a core technical prerequisite for achieving high-quality dubbing and harmonious audio-visual effects in post-production.

[0003] In existing original audio track separation schemes, automated separation logic based on fixed-band filtering is typically used. This scheme first performs full-frequency domain acquisition on the original audio stream; then, it uses a preset bandpass filter bank to truncate and separate the audio according to fixed frequency ranges in an attempt to extract human voices within a specific frequency range; finally, it performs preliminary aggregation and output of the separated frequency band data.

[0004] However, this fixed-band filtering approach has significant technical limitations. Because character voices, background music, and environmental sound effects in film audio exhibit substantial dynamic overlap across the frequency spectrum, simple band filtering struggles to handle frequency aliasing in complex acoustic environments. This results in a large amount of residual background noise remaining in the separated voice tracks, and it fails to achieve automated identification and attribution for different characters. The resulting audio tracks have low purity and lack character-level track attribute information, making it difficult to directly meet the precise voice tracking requirements of high-standard dubbing and translation. Summary of the Invention

[0005] This application provides a method and apparatus for automated track splitting of original audio for film dubbing, in order to at least alleviate the aforementioned technical problems.

[0006] An automated track-splitting method for original film audio in film dubbing includes: Step 1: Retrieve the original composite audio track data carried by the original video and perform temporal segmentation on the original composite audio track data based on the overlap sampling rule to generate a temporal audio frame sequence; construct an acoustic analysis initial characterization data table based on the temporal audio frame sequence; Step 2: The initial acoustic representation data table is injected as input excitation into the pre-trained acoustic source deep decoupling model for processing to generate the corresponding time-frequency domain acoustic mask weight sequence; based on the time-frequency domain acoustic mask weight sequence and the initial acoustic representation data table, the background sound environment energy feature parameter sequence is generated. Step 3: Construct a role-identified voice track attribute mapping detail table to record the voice attribution relationship; Step 4: Based on the character-identified human voice track-splitting attribute mapping details table and the background sound environment energy characteristic parameter sequence, generate automated film dubbing track-splitting audio control command stream data to characterize the predicted optimal track-splitting state of the original film audio and to be parsed by the dubbing system.

[0007] Optionally, an initial acoustic representation data table is constructed based on the temporal audio frame sequence, including: The time-domain audio frame sequence is mapped from the time axis to the frequency axis to generate complex distribution features representing each frame at different frequency nodes. The complex distribution features correspond to multiple time-frequency grid cells in the time-frequency domain, and each time-frequency grid cell corresponds to a complex value at the intersection of a time index and a frequency index. Power spectrum energy distribution data records that reflect the evolution of the original audio energy distribution on the time and frequency axes are generated based on the complex distribution characteristics. Metadata carrying image synchronization timestamps is extracted from a sequence of consecutive frames in the original video to generate the extracted image synchronization metadata identifier; Based on the power spectrum energy distribution data records mapped to the extracted image synchronization metadata identifier, an initial characterization data table for acoustic analysis is generated.

[0008] Optionally, based on the power spectrum energy distribution data records mapped to the extracted image synchronization metadata identifier, an initial acoustic characterization data table is generated, including: The power spectrum energy distribution data is mapped to the linear time base corresponding to the extracted image synchronization metadata identifier to generate a spatiotemporally aligned energy distribution mapping table. Based on this, the spatiotemporal alignment relationship between audio energy fluctuations and continuous image frame sequences is determined. Audio energy fluctuations are used to characterize the dynamic change trend of power spectrum energy distribution data records on the time axis. The spatiotemporally aligned energy distribution mapping table is subjected to feature space projection dimensionality reduction based on principal component analysis to generate a power spectrum energy distribution table with reduced feature dimensionality. The power spectrum energy distribution table after feature dimensionality reduction is subjected to feature dimension compression to generate an initial characterization data table for acoustic analysis.

[0009] Optionally, the initial acoustic representation data table is injected as input excitation into a pre-trained acoustic source deep decoupling model for processing to generate a corresponding time-frequency domain acoustic mask weight sequence. This includes: injecting the initial acoustic representation data table as input excitation into a semantic feature extraction network in the pre-trained acoustic source deep decoupling model to extract global acoustic semantic features; and using the mask estimation branch in the acoustic source deep decoupling model to calculate the human voice survival probability weight value based on the global acoustic semantic features to generate a corresponding time-frequency domain acoustic mask weight sequence.

[0010] Optionally, based on the time-frequency domain acoustic mask weight sequence and the acoustic analytical initial characterization data table, a sequence of background acoustic environment energy feature parameters is generated, including: Based on the time-frequency domain acoustic mask weight sequence and the acoustic analytical initial characterization data table, a time-frequency grid energy modulation distribution data table reflecting the time-frequency energy distribution weight is generated; Based on the time-frequency grid energy modulation distribution data table, the human voice feature components are separated from the initial characterization data table of acoustic analysis; Based on the time-frequency grid energy modulation distribution data table, the environmental background feature components are separated from the initial characterization data table of acoustic analysis; Based on the human voice feature components and the environmental background feature components, a sequence of background sound environment energy feature parameters representing the distribution of background ambient sound and music features is generated.

[0011] Optionally, based on the human voice feature components and the environmental background feature components, a sequence of background sound environment energy feature parameters characterizing the distribution of background ambient sound and music features is generated, including: Nonlinear signal reconstruction based on acoustic residuals is performed on the human voice feature components to generate the reconstructed human voice feature components. The phase polarity alignment algorithm is used to perform feature compensation for the spectral phase of the reconstructed human voice feature components to generate decoupled pure human voice stream data components that represent pure pronunciation features; Based on the decoupled pure human voice stream data components and using the environmental background feature components, energy feature aggregation based on the spectral envelope is performed to generate a sequence of background sound environment energy feature parameters that characterize the distribution of background ambient sound and music features.

[0012] Optionally, a role-identified voice track attribute mapping detail table is constructed to record the voice attribution relationship, including: The decoupled pure human voice stream data components are input into the pre-configured voiceprint feature clustering processing logic; The decoupled pure human voice stream data components are segmented by sliding window using voiceprint feature clustering processing logic to generate audio speech segment distribution blocks; Acoustic statistical modeling of audio speech segment distribution blocks is performed using a pre-built voiceprint feature extractor to extract and generate voiceprint embedding feature attributes of the corresponding speech subject; Based on the voiceprint embedding feature attributes, a role-identified voice track attribute mapping detail table is constructed to record the voice affiliation relationship.

[0013] Optionally, based on the voiceprint embedding feature attributes, a role-identified voice track attribute mapping detail table for recording voice attribution relationships is constructed, including: The voiceprint embedding feature attributes are mapped to their respective feature space to generate mapped voiceprint embedding feature vectors, and the geometric distance between each mapped voiceprint embedding feature vector is calculated to generate geometric distance distribution measurements. The cluster centers corresponding to each audio speech segment distribution block are determined based on the geometric distance distribution measurements. The cluster centers are introduced into the voiceprint comparison space and matched with the preset character identity feature library after feature denoising to determine the associated film character identity tags to which each audio speech segment distribution block belongs. Based on the associated film character identity tags, the decoupled pure voice stream data components are reorganized spatially based on the character affiliation attribute to construct a character-identified voice track attribute mapping detail table for recording voice affiliation relationships.

[0014] Optionally, based on the character-identified human voice track-splitting attribute mapping details table and the background sound environment energy feature parameter sequence, an automated film dubbing track-splitting audio control command stream data is generated to characterize the predicted optimal track-splitting state of the original film audio and can be parsed by the dubbing system, including: The character-identified human voice track attribute mapping details table and background sound environment energy feature parameter sequence are logically encapsulated into a data structure to generate a full-scene track-by-track acoustic timing sequence that includes timestamp index, character identification index and audio track physical feature distribution. Amplitude gain control based on a preset loudness standard is applied to the full-scene track-by-track acoustic timing sequence to generate a gain-controlled acoustic timing sequence. Based on the acoustic timing sequence after gain control, an automated film dubbing track splitting audio control command stream is generated to characterize the predicted optimal track splitting state of the original film audio and can be parsed by the dubbing system.

[0015] Optionally, based on the gain-controlled acoustic timing sequence, an automated film dubbing track splitting audio control command stream data is generated to characterize the predicted optimal track splitting state of the original film audio and can be parsed by the dubbing system, including: Determine the energy balance relationship between different character tracks based on the acoustic timing sequence after gain control; The gain-controlled acoustic timing sequence is encapsulated using a transmission protocol based on the movie dubbing standard to generate encapsulated acoustic transmission data packets. Based on the energy balance relationship, the encapsulated acoustic transmission data packet is physically resampled to generate resampled acoustic data. The resampled acoustic data is then encoded and converted to generate automated film dubbing track-splitting audio control command stream data that characterizes the predicted optimal track splitting state of the original audio and can be parsed by the dubbing system.

[0016] An automated audio track splitting device for film dubbing includes: The first program unit is used to retrieve the original composite audio track data carried by the original film and perform temporal segmentation on the original composite audio track data based on the overlap sampling rule to generate a temporal audio frame sequence; and to construct an acoustic analysis initial characterization data table for low-level audio feature analysis based on the temporal audio frame sequence. The second program unit is used to inject the acoustic analytical initial representation data table as input excitation into the pre-trained acoustic source deep decoupling model for processing, so as to generate the corresponding time-frequency domain acoustic mask weight sequence; based on the time-frequency domain acoustic mask weight sequence and the acoustic analytical initial representation data table, a background sound environment energy feature parameter sequence representing the distribution of background ambient sound and music features is generated. The third program unit is used to construct a role-identified human track attribute mapping detail table for recording the human voice attribution relationship; The fourth program unit is used to generate automated film dubbing track splitting audio control command stream data based on the character-identified human voice track splitting attribute mapping details table and the background sound environment energy characteristic parameter sequence, which is used to characterize the predicted optimal track splitting state of the original film audio and can be parsed by the dubbing system.

[0017] An electronic device includes: a memory and a processor, wherein the memory stores a computer program, and the processor is used to run the computer program to implement an automated film dubbing method with quality assessment and iterative feedback as described above.

[0018] Technical advantages of the technical solution provided in this application First, by retrieving the original composite audio track data from the original video and performing temporal segmentation based on overlap sampling rules, a temporal audio frame sequence is generated. Based on this temporal audio frame sequence, an initial acoustic representation data table is constructed, and this table is used as input excitation to a pre-trained acoustic source deep decoupling model for processing, generating a corresponding time-frequency domain acoustic mask weight sequence. Based on the time-frequency domain acoustic mask weight sequence and the initial acoustic representation data table, a sequence of background sound environment energy characteristic parameters is generated. This solves the technical problem in traditional schemes where relying on bandpass filter banks to truncate and separate frequencies within fixed frequency ranges leads to an inability to handle frequency aliasing. Compared to traditional methods that can only perform coarse processing on fixed frequency bands, this scheme achieves nonlinear decoupling of human voice and background sound features at the time-frequency grid level by constructing an initial acoustic representation data table and using a pre-trained deep decoupling model to generate a time-frequency domain acoustic mask weight sequence. This processing method can clearly identify the acoustic components in the overlapping frequency bands, resulting in high purity of the separated audio tracks and low residual noise levels in the background sound environment energy characteristic parameter sequence, thus alleviating the bottleneck of incomplete audio separation caused by dynamic spectral overlap in traditional schemes.

[0019] Secondly, by constructing a role-identified voice track attribute mapping detail table for recording voice attribution relationships, this solution addresses the technical deficiency pointed out in the background section: the inability of existing technologies to achieve automated identification and attribution recording for different roles. Compared to traditional solutions that only provide preliminary summary output of audio segment data without character attribute information, this solution uses a specially constructed role-identified mapping detail table to digitally record the attribution relationships of voices. This relay-style logic ensures that the separated voices are no longer isolated audio signals but possess clear role-level track attributes, enabling the dubbing system to clearly identify the character identity corresponding to each voice segment and meeting the technical requirements for accurate recall of specific character voice trajectories in high-standard dubbing.

[0020] Finally, by using a character-identified voice track attribute mapping detail table and a background sound environment energy feature parameter sequence, an automated film dubbing track-splitting audio control command stream is generated to characterize the predicted optimal track splitting state of the original film's audio and is parseable by the dubbing system. This solves the problem that traditional solutions produce outputs that are difficult to directly interface with the dubbing process and have a low degree of automation. Compared to the preliminary summary data output by traditional methods, this solution encapsulates the track attribute details and background sound parameters into a control command stream that can be parsed by the system. This processing method enables the track splitting results to directly drive subsequent dubbing stages in the form of commands, achieving full-process automation from the original composite audio track to the output of the optimal track splitting state, significantly improving the intelligence level of audio track processing and dubbing efficiency. Attached Figure Description

[0021] Figure 1 This is a flowchart illustrating an automated film dubbing and translation method with quality assessment and iterative feedback, as described in an embodiment of this application.

[0022] Figure 2 This is a structural block diagram of an automated movie dubbing and translation device with quality assessment and iterative feedback, according to an embodiment of this application.

[0023] Figure 3 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation

[0024] like Figure 1 As shown in the figure, an embodiment of this application provides a method for automated track splitting of original audio for film dubbing, comprising: Step 1: Retrieve the original composite audio track data carried by the original video and perform temporal segmentation on the original composite audio track data based on the overlap sampling rule to generate a temporal audio frame sequence; construct an acoustic analysis initial characterization data table based on the temporal audio frame sequence; Step 2: The initial acoustic representation data table is injected as input excitation into the pre-trained acoustic source deep decoupling model for processing to generate the corresponding time-frequency domain acoustic mask weight sequence; based on the time-frequency domain acoustic mask weight sequence and the initial acoustic representation data table, the background sound environment energy feature parameter sequence is generated. Step 3: Construct a role-identified voice track attribute mapping detail table to record the voice attribution relationship; Step 4: Based on the character-identified human voice track-splitting attribute mapping details table and the background sound environment energy characteristic parameter sequence, generate automated film dubbing track-splitting audio control command stream data to characterize the predicted optimal track-splitting state of the original film audio and to be parsed by the dubbing system.

[0025] Optionally, an initial acoustic representation data table is constructed based on the temporal audio frame sequence, including: The time-domain audio frame sequence is mapped from the time axis to the frequency axis to generate complex distribution features representing each frame at different frequency nodes. The complex distribution features correspond to multiple time-frequency grid cells in the time-frequency domain, and each time-frequency grid cell corresponds to a complex value at the intersection of a time index and a frequency index. Power spectrum energy distribution data records that reflect the evolution of the original audio energy distribution on the time and frequency axes are generated based on the complex distribution characteristics. Metadata carrying image synchronization timestamps is extracted from a sequence of consecutive frames in the original video to generate the extracted image synchronization metadata identifier; Based on the power spectrum energy distribution data records mapped to the extracted image synchronization metadata identifier, an initial characterization data table for acoustic analysis is generated.

[0026] Preferably, for the mapping processing of the time-domain audio frame sequence from the time axis to the frequency axis, windowing preprocessing is first performed on each time-domain audio frame in the time-domain audio frame sequence to suppress the spectral leakage problem caused by signal abrupt changes at both ends of the audio frame. In this application, the time-domain audio frame sequence is the sequence data generated after the previous step performs time-domain segmentation based on the overlapping sampling rule on the original composite audio track data. Each time-domain audio frame retains overlapping sampling points to achieve signal continuity between adjacent frames and avoid information discontinuity in the time dimension during subsequent time-frequency conversion. The specific implementation process of windowing preprocessing is to perform progressive attenuation processing on the sampling amplitude at both ends of the time-domain audio frame, so that the amplitude at both ends of the frame smoothly transitions to zero, avoiding abrupt changes caused by hard truncation of the signal, and thus generating false frequency components in the subsequent frequency conversion process, achieving the authenticity and reliability of the frequency dimension features. Each time-domain audio frame after windowing preprocessing is used as the direct input object for subsequent time-frequency conversion processing, providing a preprocessing basis for generating accurate complex distribution features.

[0027] Preferably, time-frequency conversion processing is performed on the windowed preprocessed time-domain audio frames to map the time-domain audio frame sequence from the time axis to the frequency axis, generating complex distribution characteristics of each frame at different frequency nodes. The core function of time-frequency conversion processing is to convert a one-dimensional time-domain amplitude signal with time as the independent variable into a two-dimensional frequency-domain component signal with frequency as the independent variable, while completely preserving the phase information of the audio signal. After the time-frequency conversion processing is completed, each windowed preprocessed time-domain audio frame generates a set of complex values ​​corresponding one-to-one with preset frequency nodes. Each complex value contains a real component and an imaginary component, where the real component represents the amplitude of the in-phase component of the audio signal at that frequency node, and the imaginary component represents the amplitude of the quadrature component of the audio signal at that frequency node. By combining the real and imaginary components, the amplitude and phase shift information of the audio signal at that frequency node can be completely restored. The complex values ​​corresponding to all time-domain audio frames together constitute a complete complex distribution feature, covering the entire time length and preset full-frequency range of the original composite audio track data, providing a complete feature basis that includes both amplitude and phase information for the subsequent fine separation of acoustic components.

[0028] Preferably, the generated complex distribution features are mapped to a two-dimensional time-frequency grid space to form multiple time-frequency grid units that correspond one-to-one with the complex distribution features. The time-frequency grid space has two mutually orthogonal dimensions: a time dimension and a frequency dimension. The scale of the time dimension corresponds one-to-one with the frame number of the time-domain audio frame sequence, and each scale corresponds to a unique time index, which can directly locate the absolute time position of the corresponding frame in the original composite audio track data. The scale of the frequency dimension corresponds one-to-one with the preset frequency nodes in the time-frequency conversion process, and each scale corresponds to a unique frequency index, which can directly locate the frequency value of the corresponding node within the full frequency band. Each time-frequency grid unit corresponds to the intersection of a time index in the time dimension and a frequency index in the frequency dimension. The value stored in each time-frequency grid unit is the complex value at the corresponding time and frequency position in the complex distribution features. By constructing time-frequency grid units, the originally continuous time-frequency domain signal can be discretized into the smallest unit that can be accurately located and processed independently. Subsequent acoustic mask weight calculation and human voice and background sound separation operations are all performed using time-frequency grid units as the smallest execution unit. This is different from the coarse processing of the entire frequency range by traditional fixed frequency band filtering, and achieves fine control of time and frequency in two dimensions, which is suitable for the high requirements of audio track separation purity in movie dubbing scenarios.

[0029] Preferably, based on the generated complex distribution characteristics, power spectrum energy calculation processing is performed to generate a power spectrum energy distribution data record reflecting the evolution of the original audio energy distribution along the time and frequency axes. The power spectrum energy calculation processing targets each complex value in the complex distribution characteristics. In specific implementation, the signal amplitude corresponding to the complex value is first calculated using the real and imaginary components of the complex value. Then, the signal amplitude is squared to obtain the signal energy value corresponding to that time-frequency position. By orderly arranging and combining the signal energy values ​​corresponding to all time-frequency grid units according to the correspondence between time and frequency indices, a complete power spectrum energy distribution data record can be generated. The power spectrum energy distribution data record fully characterizes the audio energy magnitude at each time point and frequency node in the original composite audio track data, clearly reflecting the energy distribution patterns of different acoustic components. For example, the energy of human voices in movies is mainly concentrated in the mid-low frequency range and exhibits a discontinuous distribution characteristic corresponding to the dialogue content in time, while the energy distribution of background music is more continuous, and the energy distribution of ambient sounds changes dynamically with the film scene. Power spectrum energy distribution data records are the core foundational data for constructing the initial characterization data table for acoustic analysis, and also the core input basis for subsequent deep decoupling models of acoustic sources, providing core feature support in the energy dimension for the decoupling and separation of different acoustic sources.

[0030] Preferably, normalization preprocessing is performed on the generated power spectrum energy distribution data record to eliminate feature distribution deviations caused by overall loudness differences in the original composite audio track data. Since the original film source formats accessed in film dubbing scenarios are diverse, the overall loudness of different films and different segments of the same film may fluctuate significantly. Directly using the original energy values ​​for subsequent processing would lead to a shift in feature distribution, thus affecting the processing effect of the subsequent acoustic source deep decoupling model. The specific execution process of normalization preprocessing is as follows: first, the maximum and minimum values ​​of all signal energy values ​​in the power spectrum energy distribution data record are statistically analyzed; then, linear mapping processing is performed on each signal energy value to uniformly map all signal energy values ​​to a preset fixed value range, ensuring that audio data of different loudnesses have a consistent feature distribution range. The power spectrum energy distribution data record after normalization preprocessing not only completely preserves the relative distribution relationship of the original audio energy in the time-frequency domain but also eliminates feature interference caused by absolute loudness differences, improving the stability and consistency of subsequent feature processing and adapting to the audio processing needs of film source materials with different formats and loudness standards.

[0031] Preferably, metadata carrying the image synchronization time stamp is extracted from the continuous frame sequence of the original film to generate the extracted image synchronization metadata identifier. In this application, the continuous frame sequence and the original composite audio track data come from the same original film, ensuring the homogeneity of audio and video data and the consistency of the time base. Each frame corresponds to a unique image synchronization time stamp, which represents the absolute playback time position of the frame in the original film, and uses the same time measurement base as the time axis of the original composite audio track data, achieving the underlying consistency of the audio and video time axis. The specific execution process of metadata extraction is as follows: according to the playback order of the continuous frame sequence, the header information of each frame is read sequentially, the image synchronization time stamp corresponding to the frame is parsed from the header information, and the auxiliary metadata information such as the frame number and content type of the frame is recorded simultaneously. The parsed image synchronization time stamp and auxiliary metadata information are integrated and encapsulated to generate an image synchronization metadata identifier corresponding to each frame. Image synchronization metadata identification provides a standard time reference for the spatiotemporal alignment of subsequent audio features and video frames. It can solve the problem of audio-visual time deviation caused by traditional audio track splitting schemes that only process audio data and ignore video synchronization information. It enables the audio data after track splitting to fully match the rhythm of the original film's video, meeting the high requirements for audio-visual synchronization accuracy in the process of film dubbing.

[0032] Preferably, all generated image synchronization metadata identifiers undergo continuity and consistency verification to eliminate audio-visual alignment deviations caused by timescale jumps and repetitions in the frame sequence. The verification process targets all generated image synchronization metadata identifiers. Specifically, following the playback order of the continuous frame sequence, the time interval between the image synchronization timestamps corresponding to two adjacent image synchronization metadata identifiers is verified sequentially. It is determined whether this time interval is consistent with the standard frame interval corresponding to the preset frame rate of the original film. Simultaneously, it is verified whether each image synchronization timestamp is within the total playback duration of the original film and whether there are duplicate timestamp values. For image synchronization metadata identifiers with abnormal timestamps, linear interpolation is used to correct the corresponding image synchronization timestamps, ensuring that the image synchronization timestamps corresponding to all image synchronization metadata identifiers form a continuous, uniform, and non-jumping timeline sequence. After verification and correction, the image synchronization metadata identifiers possess a time reference that perfectly matches the timeline of the original composite audio track data, providing a reliable synchronization basis for the spatiotemporal alignment of subsequent audio features and frame sequences. This further reduces the risk of audio-visual desynchronization and improves the usability of the track splitting processing results.

[0033] Optionally, based on the power spectrum energy distribution data records mapped to the extracted image synchronization metadata identifier, an initial acoustic characterization data table is generated, including: The power spectrum energy distribution data is mapped to the linear time base corresponding to the extracted image synchronization metadata identifier to generate a spatiotemporally aligned energy distribution mapping table. Based on this, the spatiotemporal alignment relationship between audio energy fluctuations and continuous image frame sequences is determined. Audio energy fluctuations are used to characterize the dynamic change trend of power spectrum energy distribution data records on the time axis. The spatiotemporally aligned energy distribution mapping table is subjected to feature space projection dimensionality reduction based on principal component analysis to generate a power spectrum energy distribution table with reduced feature dimensionality. The power spectrum energy distribution table after feature dimensionality reduction is subjected to feature dimension compression to generate an initial characterization data table for acoustic analysis.

[0034] Preferably, a unified linear time reference is first constructed that perfectly matches the playback progress of the original video. This provides a unified time scale for the mapping and alignment of power spectrum energy distribution data records and extracted image synchronization metadata identifiers. The starting point of the unified linear time reference coincides perfectly with the start playback time of the original video, and the ending point coincides perfectly with the total playback duration of the original video. The scale precision of the time axis is consistent with the frame interval of the temporal audio frame sequence, ensuring that each time node corresponding to a temporal audio frame can find a unique corresponding scale position in the unified linear time reference. Simultaneously, the image synchronization time stamp carried in the extracted image synchronization metadata identifier for each frame can also find a unique corresponding scale position in the unified linear time reference. This eliminates the time axis mismatch problem caused by differences in sampling frequency and frame rate between audio frames and frame frames, providing a unified time reference standard for subsequent spatiotemporal mapping processing.

[0035] Preferably, the power spectrum energy distribution data is mapped to the linear time base corresponding to the extracted image synchronization metadata identifier to generate a spatiotemporally aligned energy distribution mapping table. Each set of energy data in the power spectrum energy distribution data record corresponds to the time index of a time-domain audio frame in the time-domain audio frame sequence. First, the absolute time position corresponding to each time index is mapped to a unified linear time base to determine the corresponding scale interval of the set of energy data in the unified linear time base. Then, the image synchronization time stamp corresponding to each frame in the extracted image synchronization metadata identifier is mapped to the unified linear time base to determine the corresponding scale interval of the frame in the unified linear time base. Subsequently, the energy data within the same scale interval are bound and associated with the frame identifier, so that each set of energy data carries the metadata information of the corresponding frame. Finally, according to the time order of the unified linear time base, all the bound and associated energy data are arranged in order to generate a spatiotemporally aligned energy distribution mapping table.

[0036] Preferably, based on the spatiotemporally aligned energy distribution mapping table, the spatiotemporal alignment relationship between audio energy fluctuations and continuous frame sequences is determined. Audio energy fluctuations are used to characterize the dynamic change trend of power spectrum energy distribution data recorded on the time axis. First, based on the spatiotemporally aligned energy distribution mapping table, the magnitude of energy value changes within a continuous time scale is statistically analyzed, and the time nodes where energy values ​​change abruptly are extracted to generate an audio energy fluctuation node sequence. Then, based on the extracted image synchronization metadata identifier, the time nodes where screen content changes in the continuous frame sequence are extracted to generate a frame switching node sequence. Subsequently, the audio energy fluctuation node sequence and the frame switching node sequence are matched and compared on a unified linear time reference to determine the time deviation of the audio energy fluctuation node corresponding to each frame switching node, and at the same time, the audio energy fluctuation range corresponding to the playback time interval of each frame is determined, thus forming a complete spatiotemporal alignment relationship between audio energy fluctuations and continuous frame sequences. This spatiotemporal alignment relationship can clearly characterize the correspondence between changes in film scene and changes in audio energy, providing a scene-dimensional reference for subsequent acoustic source decoupling processing, which is different from the processing logic of traditional audio processing solutions that only focus on audio features and ignore the scene correlation.

[0037] Preferably, a centralized preprocessing is performed on the spatiotemporally aligned energy distribution map to generate a standardized energy distribution data table adapted to the requirements of principal component analysis. The centralized preprocessing targets all energy values ​​in the spatiotemporally aligned energy distribution map. First, the energy values ​​are grouped according to the frequency dimension. Each group of energy values ​​corresponds to the energy data at the same frequency index under all time scales in a unified linear time base. Then, for each group of energy values ​​in the frequency dimension, the arithmetic mean of all energy values ​​in that group is calculated to obtain the energy mean corresponding to that frequency dimension. Subsequently, the energy mean corresponding to that frequency dimension is subtracted from each energy value in that frequency dimension, making the mean of energy values ​​in each frequency dimension zero, eliminating the feature weight bias caused by the inherent differences in energy amplitude between different frequency dimensions. The standardized energy distribution data table generated after centralized preprocessing can avoid the problem in subsequent principal component analysis where high-energy low-frequency components excessively dominate the feature space, causing the mid-to-low frequency features related to human voices to be submerged. This achieves feature dimensionality reduction while retaining the core effective information related to distinguishing between human voices and background sounds.

[0038] Preferably, a feature space projection dimensionality reduction process based on principal component analysis is performed on the standardized energy distribution data table to generate a feature-dimension-reduced power spectrum energy distribution table. First, the covariance matrix between frequency dimensions is calculated based on the standardized energy distribution data table. The rows and columns of the covariance matrix correspond to the indices of the frequency dimensions. The value at each intersection position in the covariance matrix corresponds to the covariance between energy data from two different frequency dimensions, representing the degree of correlation between the energy changes in the two frequency dimensions. Then, the covariance matrix is ​​solved to obtain all corresponding eigenvalues ​​and eigenvectors. Each eigenvalue corresponds to an eigenvector, and the magnitude of the eigenvalue represents the amount of original data information carried by the eigenvector. Subsequently, the eigenvectors are sorted in descending order of eigenvalues, and the top-ranked eigenvectors are selected to form a projection matrix. The number of columns in the projection matrix is ​​the number of feature dimensions after dimensionality reduction. Finally, the standardized energy distribution data table is projected onto this projection matrix to obtain the feature data after dimensionality reduction. This feature data is then associated and integrated with the corresponding spatiotemporal alignment relationship and frame data information to generate the feature-dimension-reduced power spectrum energy distribution table. This process can remove redundant information and noise components from the original high-dimensional features, significantly reduce the dimensionality of the feature data while retaining the core acoustic distinguishing features, reduce the computational cost of the subsequent acoustic source deep decoupling model, and improve the model's ability to distinguish different acoustic sources.

[0039] Preferably, feature normalization preprocessing is performed on the power spectrum energy distribution table after feature dimensionality reduction to generate a normalized feature data table adapted to the requirements of dimensionality compression processing. The feature normalization preprocessing object is the power spectrum energy distribution table after feature dimensionality reduction. First, normalization processing is performed on the values ​​of each feature dimension after dimensionality reduction, and all values ​​in each feature dimension are linearly mapped to a preset fixed value range to eliminate the difference in numerical amplitude between different feature dimensions and make the numerical distribution of all feature dimensions of the same order of magnitude. Then, the normalized feature data is sorted and verified according to the time order of a unified linear time base to ensure that the time order of the feature data is completely consistent with the playback order of the original film and there are no time misalignments or missing data problems. Subsequently, the spatiotemporal alignment relationship, frame metadata identifier, and time index information corresponding to each set of feature data are structurally encapsulated so that each set of feature data forms a complete data unit containing acoustic features, spatiotemporal information, and image association information, and finally, a normalized feature data table is generated. This preprocessing process ensures that the input data for subsequent feature dimension compression processing has a uniform dimensional specification and numerical distribution, avoiding deviations in the compression processing results caused by uneven data distribution.

[0040] Preferably, feature dimension compression processing is performed on the normalized feature data table to generate an initial acoustic analysis representation data table. The feature dimension compression processing takes the normalized feature data table as input. First, according to a preset time window length, the normalized feature data table is divided into sliding windows on a unified linear time base. Each sliding window corresponds to a continuous playback time interval in the original video, and a preset overlap interval is maintained between adjacent sliding windows to avoid loss of feature information at window boundaries. Then, for all normalized feature data within each sliding window, feature aggregation processing is performed to extract the statistical features of each feature dimension within that window, including the mean, variance, peak, and trough values ​​of that dimension. Simultaneously, the corresponding frame metadata identifier, spatiotemporal alignment parameters, and time interval information are retained. Subsequently, the aggregation processing results corresponding to all sliding windows are arranged in chronological order to form a structured data table. Each row of this data table structure corresponds to the time index of a sliding window, and each column corresponds to an aggregated feature dimension, ultimately generating the initial acoustic analysis representation data table. The initial acoustic analysis data table fully integrates the core acoustic features of the original audio, the spatiotemporal alignment information of audio and video, and the film scene association information. It can be directly used as the input of the subsequent acoustic source deep decoupling model, providing basic feature data with both discriminative and synchronous characteristics for audio track separation processing in film dubbing scenarios. Unlike traditional solutions that only use pure audio features, it can effectively improve the accuracy of subsequent audio separation and the audio-visual synchronization effect.

[0041] Optionally, the initial acoustic representation data table is injected as input excitation into a pre-trained acoustic source deep decoupling model for processing to generate a corresponding time-frequency domain acoustic mask weight sequence. This includes: injecting the initial acoustic representation data table as input excitation into a semantic feature extraction network in the pre-trained acoustic source deep decoupling model to extract global acoustic semantic features; and using the mask estimation branch in the acoustic source deep decoupling model to calculate the human voice survival probability weight value based on the global acoustic semantic features to generate a corresponding time-frequency domain acoustic mask weight sequence.

[0042] Preferably, the pre-training of the acoustic source deep decoupling model is completed first, providing a processing foundation adapted to the film dubbing scenario for subsequent feature extraction and mask generation. The acoustic source deep decoupling model adopts a two-branch serial structure design, with two core components: a semantic feature extraction network and a mask estimation branch. The output of the semantic feature extraction network is directly connected to the input of the mask estimation branch, forming a forward propagation processing link. The model's pre-training process is completed using an annotated audio dataset from the film dubbing scenario. The samples in the dataset are all from professionally tracked film audio data. Each sample contains input features consistent with the format of the initial acoustic analysis representation data table, as well as corresponding annotated real mask labels for the human voice region and the background sound region. Through the pre-training process, the model learns the feature distribution differences of human voice, background music, and environmental sound effects in film audio, as well as the correlation between human voice features and temporal context and film scene. This differs from the training logic of general speech separation models that are only adapted to everyday dialogue scenarios, and can better adapt to the complex acoustic environment of film dubbing scenarios.

[0043] Preferably, the initial acoustic representation data table is subjected to model input adaptation preprocessing to generate a standardized input feature sequence that conforms to the input specification of the semantic feature extraction network. The preprocessing object is the acoustic analysis initial representation data table generated in the previous steps. First, the acoustic analysis initial representation data table is structured and parsed to extract four core data categories: temporal acoustic feature data, time index information, audio-visual synchronization metadata identifier, and spatiotemporal alignment parameters. Then, the extracted temporal acoustic feature data is dimensionally aligned according to the input dimension requirements preset by the semantic feature extraction network. Feature sequences with insufficient dimensions are padded with zero values, and feature sequences with excessive dimensions are pruned according to feature contribution, so that all feature sequences have a uniform dimension specification. Subsequently, amplitude normalization is performed on the dimension-aligned feature sequences to map all feature values ​​to a preset fixed value range, eliminating feature amplitude differences between different films and segments. Finally, the normalized feature sequences are structured and encapsulated with the corresponding time index information, audio-visual synchronization metadata identifier, and spatiotemporal alignment parameters to generate a standardized input feature sequence, which serves as the direct input to the semantic feature extraction network.

[0044] Preferably, the standardized input feature sequence is injected into the semantic feature extraction network of a pre-trained acoustic source deep decoupling model to perform shallow local acoustic feature extraction processing, thereby generating shallow local acoustic features that capture local temporal correlations. The front end of the semantic feature extraction network is equipped with multi-layer cascaded one-dimensional convolutional units and gated linear units, wherein the convolution window size of the one-dimensional convolutional unit is adapted to the smallest temporal unit of human voice pronunciation in movie audio, which can perform correlation extraction of acoustic features of adjacent time slices. In the processing, the standardized input feature sequence is first fed into the first-layer one-dimensional convolutional unit. Local convolution operations are performed on the acoustic features within each time window to capture the changes in acoustic features between adjacent time slices and extract primary features that characterize the local spectral details of the audio. The primary features are then fed into a gated linear unit, where a gating mechanism filters the primary features, retaining feature components related to the human voice spectral features and suppressing invalid feature components related to background noise and other noise. Subsequently, the gated features are fed into subsequent cascaded convolutional units and gated linear units, and the local feature extraction and gating operations are repeated to gradually refine the local expressive power of the features, ultimately generating shallow local acoustic features that cover the local spectral details of the human voice and the information related to adjacent temporal sequences.

[0045] Preferably, based on the generated shallow local acoustic features, a bidirectional temporal coding layer in the semantic feature extraction network is used to perform global context dependency modeling to extract global acoustic semantic features. The bidirectional temporal coding layer adopts a bidirectional temporal coding structure, which can simultaneously capture the forward and backward temporal dependencies in the audio sequence, adapting to the contextual association characteristics of human voice statements in movie dialogues. In the processing, shallow local acoustic features are first input into a bidirectional temporal coding layer. For each time position, the features are simultaneously fused with forward feature information from all previous time slices and backward feature information from all subsequent time slices to establish a global temporal correlation between all time slices in the entire audio sequence, capturing the global evolution of human voice features throughout the film's duration. Then, combined with the audio-visual synchronization metadata identifier and spatiotemporal alignment parameters carried in the initial acoustic analysis representation data table, the weights of features corresponding to time points where scene changes occur, characters appear, and disappear are adjusted. The feature weights of time slices where characters appear are strengthened, while the feature weights of time slices without characters are weakened. Finally, a global acoustic semantic feature is generated that simultaneously covers local acoustic details, global temporal context dependencies, and film scene correlation information. This feature can completely represent the global distribution pattern of different acoustic sources in the film audio, providing a comprehensive feature basis for subsequent mask estimation.

[0046] Preferably, the extracted global acoustic semantic features are input into the mask estimation branch of the acoustic source deep decoupling model, and time-frequency dimension feature mapping processing is performed to generate time-frequency grid-related semantic features. The mask estimation branch has a feature mapping layer at its front end to map the high-dimensional global acoustic semantic features back to a low-dimensional feature space that corresponds one-to-one with the original time-frequency grid cells. During processing, the global acoustic semantic features are first transformed into feature groups corresponding one-to-one with the original time and frequency indices, with each feature group corresponding to a time-frequency grid cell. Then, each feature group is upsampled to restore its time-frequency resolution, ensuring that the dimension of each feature group perfectly matches the original feature dimension of the corresponding time-frequency grid cell. Subsequently, all feature groups are arranged in order according to the correspondence between time and frequency indices, ensuring that each feature group forms an accurate one-to-one correspondence with the corresponding time-frequency grid cell in the original acoustic analysis initial representation data table. Finally, time-frequency grid-related semantic features are generated, where each feature group can directly characterize the acoustic component attributes within the corresponding time-frequency grid cell.

[0047] Preferably, based on the generated time-frequency grid-related semantic features, the probability calculation layer in the mask estimation branch is used to perform human voice attribution probability calculation to generate a human voice survival probability weight value corresponding to each time-frequency grid cell. The probability calculation layer performs attribution classification calculation on the semantic features related to each time-frequency grid cell based on the feature distribution boundaries of human voice and background sound learned during pre-training. In the processing, firstly, for each time-frequency grid cell, the matching degree between the associated semantic features of the time-frequency grid and the distribution center of the human voice features, and the matching degree between the associated semantic features and the distribution center of the background sound features are calculated. Then, based on the ratio of the two matching degrees, the preliminary probability value of the acoustic component in the time-frequency grid cell belonging to human voice is calculated. The value of the preliminary probability value is between 0 and 1. Subsequently, a smoothing filter is performed on the preliminary probability values ​​of adjacent time-frequency grid cells to eliminate isolated probability abrupt changes in the time-frequency domain. Because human voices in movie audio exhibit a continuous distribution in the time-frequency domain, isolated probability abrupt changes are likely interference from background noise. Finally, a human voice survival probability weight value is generated for each time-frequency grid cell. The closer the weight value is to 1, the higher the probability that the acoustic component in the corresponding time-frequency grid cell belongs to human voice; the closer it is to 0, the higher the probability that it belongs to background sound.

[0048] Preferably, based on the voice survival probability weight values ​​corresponding to all time-frequency grid units, an ordered integration process is performed along the time-frequency dimension to generate a complete time-frequency domain acoustic mask weight sequence. During the process, firstly, according to the correspondence between the time index and frequency index in the original acoustic analysis initial characterization data table, the voice survival probability weight values ​​corresponding to all time-frequency grid units are arranged and combined to form a two-dimensional time-frequency mask weight matrix. The row dimension of this matrix corresponds to the time index, and the column dimension corresponds to the frequency index. The values ​​at the intersection of the row and column are the voice survival probability weight values ​​of the time-frequency grid units corresponding to the time index and frequency index. Then, according to the chronological order of the original audio timeline, the two-dimensional time-frequency mask weight matrices corresponding to all time slices are sequentially spliced ​​together to form a three-dimensional weight sequence covering the complete playback duration of the original film. Finally, the spliced ​​weight sequence undergoes a temporal consistency check to ensure that the sequence length completely matches the time length of the original composite audio track data, and that the time index corresponds one-to-one with the frame number of the original time-domain audio frame sequence, ultimately generating the time-frequency domain acoustic mask weight sequence. This sequence can be directly used for subsequent separation processing of human voice and background sound, providing a quantitative weight basis for the acoustic component attribution of each time-frequency grid unit, and realizing refined acoustic source decoupling at the time-frequency grid level.

[0049] Optionally, based on the time-frequency domain acoustic mask weight sequence and the acoustic analytical initial characterization data table, a sequence of background acoustic environment energy feature parameters is generated, including: Based on the time-frequency domain acoustic mask weight sequence and the acoustic analytical initial characterization data table, a time-frequency grid energy modulation distribution data table reflecting the time-frequency energy distribution weight is generated; Based on the time-frequency grid energy modulation distribution data table, the human voice feature components are separated from the initial characterization data table of acoustic analysis; Based on the time-frequency grid energy modulation distribution data table, the environmental background feature components are separated from the initial characterization data table of acoustic analysis; Based on the human voice feature components and the environmental background feature components, a sequence of background sound environment energy feature parameters representing the distribution of background ambient sound and music features is generated.

[0050] Preferably, the time-frequency domain acoustic mask weight sequence and the initial acoustic representation data table are subjected to spatiotemporal dimension alignment verification and standardization preprocessing to generate a mask-feature pairing dataset with perfectly matched dimensions. Both the time-frequency domain acoustic mask weight sequence and the initial acoustic representation data table are generated based on the composite audio track data of the same original film. They share a unified linear time reference and frequency node division rules. However, there may be issues with temporal offset and dimension length deviation during the generation process. Therefore, the time index sequence and frequency index sequence of the two sets of data are first matched and verified one by one. For sequences with inconsistent time lengths, edge padding or tail pruning is used to unify the time dimension length. For sequences with inconsistent frequency node numbers, linear interpolation is used to unify the frequency dimension node number. This ensures that the intersection of each time index and each frequency index in the two sets of data can correspond to a unique time-frequency grid unit, and each time-frequency grid unit has a corresponding human voice survival probability weight value and original acoustic feature data. Simultaneously, the audio-visual synchronization metadata identifier and spatiotemporal alignment parameters carried in the initial acoustic analysis representation data table are synchronously mapped to the mask-feature pairing dataset, ensuring that the time reference of the pairing data is completely aligned with the frame sequence of the original film, providing the input basis for dimensional matching and temporal synchronization for subsequent time-frequency grid energy modulation processing.

[0051] Preferably, based on the mask-feature pairing dataset, a dual-channel weighted modulation operation at the time-frequency grid cell level is performed to generate an initial modulation weight matrix. The core of the dual-channel weighted modulation operation is to generate modulation coefficients for the human voice channel and background sound channel for each time-frequency grid cell. During the operation, the human voice survival probability weight value corresponding to each time-frequency grid cell in the mask-feature pairing dataset is used as the modulation coefficient for the human voice channel of that grid cell; the result obtained by subtracting the human voice survival probability weight value from a fixed value of 1 is used as the modulation coefficient for the background sound channel of that grid cell. The modulation coefficient for the human voice channel controls the degree of preservation of human voice features within the corresponding time-frequency grid cell. The closer the modulation coefficient is to 1, the higher the probability that the acoustic features within that grid cell belong to human voice, and the higher the proportion of features preserved. The modulation coefficient for the background sound channel controls the degree of preservation of background sound features within the corresponding time-frequency grid cell. The closer the modulation coefficient is to 1, the higher the probability that the acoustic features within that grid cell belong to background sound, and the higher the proportion of features preserved. The modulation coefficients of the human voice channel and the background sound channel corresponding to all time-frequency grid units are arranged in order according to the correspondence between time index and frequency index to form a two-dimensional initial modulation weight matrix. The row dimension of this matrix corresponds to the time index, and the column dimension corresponds to the frequency index. Each element at the intersection of row and column contains a set of dual-channel modulation coefficients, providing a quantitative modulation basis for subsequent feature separation.

[0052] Preferably, a time-frequency dual-dimensional smoothing constraint is applied to the initial modulation weight matrix to generate an optimized modulation weight matrix, which is then integrated to generate a time-frequency grid energy modulation distribution data table. Human voices in movie audio exhibit temporal continuity, while background music and ambient sounds exhibit frequency domain continuity. Isolated modulation coefficient abrupt changes are likely computational deviations caused by noise interference, leading to issues such as noise and audio dropouts in the subsequently separated audio. Therefore, a smoothing constraint is applied to the initial modulation weight matrix. During the process, a sliding window smoothing filter is first applied to the initial modulation weight matrix along the time dimension. For each time position, the modulation coefficients are weighted and averaged with the modulation coefficients of their immediate and adjacent time positions to eliminate isolated abrupt changes in the time dimension. Then, a sliding window smoothing filter is applied to the time-smoothed weight matrix along the frequency dimension. For each frequency position, the modulation coefficients are weighted and averaged with the modulation coefficients of their immediate and adjacent frequency positions to eliminate isolated abrupt changes in the frequency dimension, ultimately generating the optimized modulation weight matrix. The optimized modulation weight matrix is ​​then structurally encapsulated with the corresponding time index information, frequency index information, audio-visual synchronization metadata identifier, and spatiotemporal alignment parameters, so that the dual-channel modulation coefficients of each time-frequency grid unit are bound to the corresponding original acoustic features, time position, and picture frame information. Finally, according to the ordered arrangement rules of time and frequency, a time-frequency grid energy modulation distribution data table is generated.

[0053] Preferably, based on the human voice channel modulation coefficients in the time-frequency grid energy modulation distribution data table, a weighted extraction process for human voice features is performed on the initial acoustic characterization data table to generate an initial human voice feature matrix. The weighted extraction process uses the time-frequency grid cell as the smallest unit of execution, and the processing object is the original acoustic feature data corresponding to each time-frequency grid cell in the initial acoustic characterization data table. The processing action involves multiplying this original acoustic feature data with the corresponding human voice channel modulation coefficients in the time-frequency grid energy modulation distribution data table to obtain the human voice feature component fragments corresponding to that time-frequency grid cell. This operation process can finely filter the acoustic features within each time-frequency grid cell, retaining feature components related to human voice and suppressing feature components related to background noise, thus fundamentally alleviating the problem of frequency aliasing that traditional fixed-band filtering cannot handle. All human voice feature component fragments corresponding to all time-frequency grid units are arranged in an ordered manner according to the corresponding order of time index and frequency index to form an initial human voice feature matrix that perfectly matches the original acoustic feature dimensions. This matrix completely preserves the time-frequency domain feature distribution related to human voice, providing basic data for the subsequent generation of complete human voice feature components.

[0054] Preferably, the initial voice feature matrix is ​​subjected to feature dimension restoration and phase information completion processing to generate voice feature components. The initial voice feature matrix is ​​time-frequency domain feature data after weighted filtering. It is restored to a complete feature structure corresponding to the original composite audio track data, and the phase information of the original audio is completed to avoid phase distortion in the subsequent audio reconstruction process. In the processing, the time-frequency domain amplitude features in the initial voice feature matrix are first matched and bound with the phase information in the complex distribution features of the corresponding time-frequency grid unit in the acoustic analysis initial characterization data table to restore the complex structure of the voice features in each time-frequency grid unit, ensuring that the voice features contain complete amplitude and phase information. Then, the restored complex voice features are subjected to temporal continuity verification. The time length of the feature sequence is verified to be completely consistent with the time length of the original composite audio track data, and the time index corresponds one-to-one with the frame number of the time-domain audio frame sequence to eliminate the problem of temporal misalignment. At the same time, the audio-visual synchronization metadata identifier and spatiotemporal alignment parameters are synchronously bound to the restored voice features to ensure that the time axis of the voice features is completely aligned with the frame sequence of the original film. After the above processing, the final human voice feature components, which cover the complete time-frequency domain amplitude characteristics, phase information, time synchronization information, and image association information, are generated.

[0055] Preferably, based on the background sound channel modulation coefficients in the time-frequency grid energy modulation distribution data table, a weighted extraction process for background sound features is performed on the initial acoustic characterization data table to generate an initial background feature matrix. The weighted extraction process also uses the time-frequency grid cell as the smallest unit of execution, forming a complementary processing logic with the human voice feature extraction. The processing object is the original acoustic feature data corresponding to each time-frequency grid cell in the initial acoustic characterization data table. The processing action involves multiplying this original acoustic feature data with the background sound channel modulation coefficients of the corresponding time-frequency grid cell in the time-frequency grid energy modulation distribution data table to obtain the environmental background feature component fragment corresponding to that time-frequency grid cell. This operation process can perform reverse filtering on the acoustic features within each time-frequency grid cell, retaining feature components related to background music and environmental sound effects while suppressing feature components related to human voices, thus achieving a fine decoupling between background sound features and human voice features. All environmental background feature component fragments corresponding to all time-frequency grid cells are arranged in an orderly manner according to the corresponding order of time index and frequency index to form an initial background feature matrix that perfectly matches the original acoustic feature dimensions. This matrix completely preserves the time-frequency domain feature distribution related to background sound, providing basic data for the subsequent generation of complete environmental background feature components.

[0056] Preferably, the initial background feature matrix is ​​subjected to human voice aliasing residue suppression and feature integrity verification to generate environmental background feature components. In time-frequency regions where there is severe frequency aliasing between human voice and background sound, weak human voice feature components may remain in the initial background feature matrix, which will cause human voice residue in the subsequently separated background audio track, failing to meet the requirements for the purity of the background audio track in film dubbing. Therefore, aliasing residue suppression processing is performed. During the processing, based on the time-frequency domain energy distribution of the human voice feature components, the time-frequency grid units with human voice energy higher than a preset threshold are located. Secondary attenuation suppression is performed on the feature components of the corresponding time-frequency grid units in the initial background feature matrix to further eliminate residual human voice features. Then, feature integrity verification is performed on the background features after secondary suppression to verify that the time length and frequency coverage of the feature sequence are completely consistent with the original composite audio track data, ensuring that the continuous segments of background music and environmental sound effects will not have truncation or distortion problems. At the same time, the audio-visual synchronization metadata identifier and spatiotemporal alignment parameters are synchronously bound to the processed background features to ensure that the time axis of the environmental background feature components is completely synchronized with the scene switching of the original film. After the above processing, the environmental background feature components with no obvious human voice residue and complete time-frequency domain characteristics and synchronization information are finally generated.

[0057] Optionally, based on the human voice feature components and the environmental background feature components, a sequence of background sound environment energy feature parameters characterizing the distribution of background ambient sound and music features is generated, including: Nonlinear signal reconstruction based on acoustic residuals is performed on the human voice feature components to generate the reconstructed human voice feature components. The phase polarity alignment algorithm is used to perform feature compensation for the spectral phase of the reconstructed human voice feature components to generate decoupled pure human voice stream data components that represent pure pronunciation features; Based on the decoupled pure human voice stream data components and using the environmental background feature components, energy feature aggregation based on the spectral envelope is performed to generate a sequence of background sound environment energy feature parameters that characterize the distribution of background ambient sound and music features.

[0058] Preferably, the human voice feature components are first subjected to bidirectional consistency verification and preprocessing in the time and frequency domains to generate a standardized human voice feature sequence adapted for acoustic residual calculation. The human voice feature components are feature data separated from the initial acoustic characterization data table in the previous steps, carrying complete amplitude and phase information of the corresponding time-frequency grid units. However, in the initial separation process, there may be problems such as feature discontinuity at the time-frequency grid boundaries, loss of weak-energy human voice segments, and misalignment of time sequence indexes, which will directly affect the accuracy of subsequent signal reconstruction. Therefore, standardization preprocessing must be performed first. The preprocessing process involves first verifying whether the time index sequence of the human voice feature components matches the frame number of the original time-domain audio frame sequence. For feature segments with temporal misalignment, linear interpolation is performed to align them and eliminate the reference deviation in the time dimension. Then, dynamic range normalization is performed on the human voice feature amplitude within each time-frequency grid unit to map all amplitude data to a unified numerical range, eliminating the interference of amplitude differences between different audio segments on subsequent residual calculations. Simultaneously, the audio-visual synchronization metadata identifier and spatiotemporal alignment parameters corresponding to the human voice feature components are synchronously bound to the standardized feature sequence to ensure that the time reference of the feature sequence is completely consistent with the frame sequence of the original film. Finally, a standardized human voice feature sequence is generated, providing a standardized and unified input basis for subsequent acoustic residual calculations and nonlinear signal reconstruction.

[0059] Preferably, based on the standardized human voice feature sequence and the initial acoustic analysis representation data table, the extraction and calculation of the acoustic residual components of the human voice are performed to generate a residual feature sequence representing the lost information of the human voice features. The physical essence of the acoustic residual is the difference between the complete features of the real human voice in the original composite audio track data and the standardized human voice feature sequence obtained from the initial separation. It contains key information such as weak-energy human voice harmonics that were mistakenly suppressed during the initial separation process, features of the articulation transition segment, and low-frequency details of the voiced segment. This information is the core basis for completing the human voice features and improving the quality of human voice reconstruction. The specific process of extraction and calculation is as follows: First, the standardized human voice feature sequence is paired one-to-one with the corresponding original acoustic features in the initial acoustic characterization data table according to the time index and frequency index, ensuring that the features of each time-frequency grid unit can form a one-to-one pairing relationship and avoiding misalignment of the time-frequency dimension; then, for each paired time-frequency grid unit, the difference between the effective component related to human voice in the original acoustic features and the corresponding feature in the standardized human voice feature sequence is calculated to obtain the residual feature segment corresponding to the time-frequency grid unit; subsequently, the residual feature segments of all time-frequency grid units are arranged according to the ordered rules of time index and frequency index, and noise threshold screening is performed on the residual feature segments to remove noise-type residual components with energy below the preset threshold, retaining only the effective residual components related to human voice pronunciation features, and finally generating the residual feature sequence.

[0060] Preferably, based on the standardized human voice feature sequence and the residual feature sequence, a nonlinear signal reconstruction process based on acoustic residuals is performed to generate the reconstructed human voice feature components. This process differs from the traditional linear feature superposition method because the fundamental frequency and each harmonic of the human voice have nonlinear correlation characteristics. Simple linear superposition will lead to an imbalance in the harmonic ratio, causing distortion of the human voice quality. Therefore, a nonlinear reconstruction logic adapted to the pronunciation rules of the human voice is adopted. In the specific implementation, a nonlinear reconstruction mapping function adapted to the characteristics of human voice pronunciation is first constructed. This mapping function uses the fundamental frequency characteristics of the standardized human voice feature sequence as the core reference to establish a nonlinear correlation between the fundamental frequency and each harmonic. According to the dynamic changes of the fundamental frequency, the fusion weights of the residual components of each harmonic can be adjusted in real time to ensure that the proportion of human voice harmonics after reconstruction conforms to the natural pronunciation law. Then, the standardized human voice feature sequence is used as the basic body of signal reconstruction. The effective residual components of each harmonic in the residual feature sequence are fused into the corresponding time-frequency grid cells of the basic body according to the dynamic weights calculated by the nonlinear reconstruction mapping function, so as to complete the weak energy human voice harmonics and pronunciation transition segment features lost in the initial separation process. Subsequently, time-frequency domain continuity smoothing is performed on the fused feature sequence to eliminate the feature abrupt changes of the time-frequency grid boundary generated during the fusion process. At the same time, the phase information of the reconstructed feature sequence is verified to be consistent with the phase reference of the original composite audio track data, and finally, the reconstructed human voice feature components are generated.

[0061] Preferably, phase feature analysis and benchmark alignment preprocessing are performed on the reconstructed human voice feature components to generate a phase feature sequence adapted to the phase polarity alignment algorithm. Although the reconstructed human voice feature components complete the detailed features in the amplitude dimension, problems such as polarity reversal, phase jump, and phase shift may occur during the initial separation and reconstruction process. These problems will lead to defects such as sound quality distortion, noise, and phase discontinuity in the subsequent reconstruction of the human voice time domain waveform, which cannot meet the high requirements of film dubbing for human voice quality. Therefore, it is necessary to complete the phase feature analysis and benchmark establishment first. In the specific implementation, the complex features of each time-frequency grid unit of the reconstructed human voice feature components are first analyzed to separate the amplitude and phase information corresponding to each time-frequency grid unit. All phase information is arranged in the corresponding order of time index and frequency index to generate an initial phase distribution matrix. Then, based on the original audio phase information corresponding to the initial acoustic characterization data table, a unified phase polarity benchmark is established. This benchmark takes the phase polarity corresponding to the fundamental frequency of the human voice as the core reference standard to ensure that the phase polarity of the entire audio sequence has a unified and continuous reference system. Subsequently, the initial phase distribution matrix is ​​mapped to this unified phase polarity benchmark, and the temporal continuity of the phase sequence in each frequency dimension is checked to mark abnormal phase nodes with phase polarity reversal or phase jumps exceeding the preset range. Finally, a phase feature sequence carrying abnormal phase node markers is generated to provide an accurate processing target for subsequent phase feature compensation.

[0062] Preferably, a phase polarity alignment algorithm is used to perform spectral phase feature compensation processing on the reconstructed human voice feature components based on the phase feature sequence, so as to generate decoupled pure human voice stream data components that represent pure pronunciation features. The core logic of the phase polarity alignment algorithm is to conform to the continuous and smooth phase characteristics during human voice pronunciation, to perform targeted correction and compensation for abnormal phase nodes, and to preserve the natural characteristics of human voice phase, which is different from the mechanical sound quality and distortion problems that are easily caused by traditional hard phase correction methods. In the specific implementation, firstly, for the abnormal phase nodes marked in the phase feature sequence, a phase polarity alignment algorithm is used to perform polarity correction processing. For phase nodes with polarity reversal, a polarity reversal operation is performed to make the phase polarity of the node consistent with a unified phase polarity reference. Then, for abnormal nodes with phase jumps, linear interpolation smoothing processing is performed using the phase information of the adjacent normal phase nodes to eliminate the phase jump problem and ensure that the phase sequence of each frequency dimension has continuous and smooth characteristics in the time dimension. Subsequently, the phase information after correction and smoothing is recombined with the amplitude information in the reconstructed human voice feature components to restore the complete complex human voice features in each time-frequency grid cell. Then, the restored complex human voice features are subjected to inverse mapping processing from the frequency domain to the time domain to generate a human voice waveform sequence in the time domain. At the same time, the continuity check and noise suppression processing for speech segments are performed on the time domain waveform sequence to remove residual noise in speech segments. Finally, a decoupled clean human voice stream data component representing the clean pronunciation features is generated.

[0063] Preferably, spectral envelope extraction and feature standardization are performed on the environmental background feature components to generate a background sound standardized feature sequence adapted to energy feature aggregation. The environmental background feature components are feature data separated from the initial acoustic analysis representation data table in the previous steps. They contain the time-frequency features of non-human acoustic components such as background music and environmental sound effects in the original film. Only after dimensional alignment and spectral envelope extraction are completed can accurate aggregation processing be achieved with the decoupled pure human voice stream data components, further eliminating human voice residue in the background sound. In the specific implementation, the environmental background feature components are first subjected to time-frequency domain consistency verification to ensure that their time index range, frequency node division, and decoupled pure human voice stream data components are completely matched in terms of time and frequency dimensions. For the parts with mismatched dimensions, interpolation alignment processing is performed to eliminate the dimensional deviation between the time sequence and the frequency domain. Then, the verified environmental background feature components are subjected to spectral envelope extraction processing. The frequency domain energy distribution of each time slice is smoothly fitted along the time dimension to generate a spectral envelope curve that characterizes the change law of background sound energy with frequency. The spectral envelope curves of all time slices are arranged in chronological order to form a complete background sound spectral envelope sequence. Subsequently, the background sound spectral envelope sequence is subjected to amplitude normalization processing to map all envelope values ​​to a unified numerical range. At the same time, the audio-visual synchronization metadata identifier and spatiotemporal alignment parameters corresponding to the environmental background feature components are synchronously bound to the processed sequence, and finally a standardized background sound feature sequence is generated, which provides a standardized input basis for subsequent energy feature aggregation.

[0064] Preferably, based on the decoupled pure human voice stream data components and the background sound normalized feature sequence, energy feature aggregation processing based on spectral envelope is performed to generate a background sound environment energy feature parameter sequence characterizing the distribution of background ambient sound and music features. The core purpose of this energy feature aggregation processing is to utilize the accurate spectral features of the decoupled pure human voice stream data components to further eliminate residual human voice components in the background sound, while aggregating the core energy features of the background sound to form a complete background sound parameter sequence that can be directly integrated into the film dubbing process. In specific implementation, the decoupled pure human voice stream data components are first mapped from the time domain back to the frequency domain to generate a human voice spectral envelope sequence that completely corresponds to the human voice pronunciation process. This sequence can accurately characterize the energy distribution state of the human voice at different time and frequency nodes. Then, the human voice spectral envelope sequence and the background sound spectral envelope sequence in the background sound normalized feature sequence are paired one-to-one according to the time index and frequency index. For each time-frequency grid cell, the human voice residue suppression weight of the background sound feature within that grid cell is calculated based on the energy of the human voice spectral envelope. The higher the human voice energy of the grid cell, the greater the corresponding suppression weight. The process begins by removing residual human voice harmonic components from the background sound. Then, the background sound features, after human voice residue suppression, are aggregated based on spectral envelope energy features. This aggregates core feature parameters such as total background sound energy, peak energy, frequency band energy distribution, and spectral envelope shape for each time slice along the time dimension. Simultaneously, each time slice's feature parameters are bound to corresponding timestamp indexes, frame synchronization identifiers, scene type markers, and other related information. Finally, the aggregated feature parameters of all time slices are arranged in a structured manner according to time sequence, forming a sequence of background sound environment energy feature parameters with complete temporal and scene-related information.

[0065] Optionally, a role-identified voice track attribute mapping detail table is constructed to record the voice attribution relationship, including: The decoupled pure human voice stream data components are input into the pre-configured voiceprint feature clustering processing logic; The decoupled pure human voice stream data components are segmented by sliding window using voiceprint feature clustering processing logic to generate audio speech segment distribution blocks; Acoustic statistical modeling of audio speech segment distribution blocks is performed using a pre-built voiceprint feature extractor to extract and generate voiceprint embedding feature attributes of the corresponding speech subject; Based on the voiceprint embedding feature attributes, a role-identified voice track attribute mapping detail table is constructed to record the voice affiliation relationship.

[0066] Preferably, the decoupled pure human voice stream data components are first subjected to voiceprint processing adaptation preprocessing to generate a standardized human voice temporal sequence that conforms to the input specifications of voiceprint feature clustering processing logic. The decoupled pure human voice stream data components are the temporal human voice waveform data generated in the previous step, which fully carries the timestamp information aligned with the original video and the audio-visual synchronization metadata identifier. However, they contain invalid content such as silent segments without speech content, residual background noise, and speech gaps, which will interfere with the accuracy of subsequent segmentation and feature extraction. Therefore, adaptation preprocessing must be performed first. The preprocessing process involves first performing endpoint detection on the decoupled pure human voice stream data components to identify and mark the time boundaries between valid speech segments and invalid silence segments in the data. Invalid silence segments with energy below a preset threshold and residual background noise are removed, leaving only valid speech segments containing human voice content. Then, amplitude normalization is performed on the retained valid speech segments to map the waveform amplitude of all valid speech segments to a unified numerical range, eliminating the interference of volume differences between different characters and different dialogue segments on subsequent voiceprint feature extraction. Simultaneously, the timestamp information and audio-visual synchronization metadata identifier corresponding to the decoupled pure human voice stream data components are bound and associated with the valid speech segments at the corresponding positions to ensure that each valid speech segment can match the corresponding time and image content in the original film. Finally, a standardized human voice time sequence is generated, providing a standardized and effective input basis for the subsequent input voiceprint feature clustering processing logic.

[0067] Preferably, the standardized human voice temporal sequence is input into the pre-configured voiceprint feature clustering processing logic to complete the initialization of the processing logic and align the temporal reference of the input data. The voiceprint feature clustering processing logic is a pre-configured processing framework designed for the characteristics of multi-character dialogue in film dubbing scenarios. Its core is used to achieve accurate segmentation of human voice segments and clustering analysis of the identity features of the speaker. Unlike general voiceprint processing logic, which is only suitable for everyday short dialogue scenarios, it can better adapt to the acoustic characteristics of long lines of movie dialogue, multiple character alternations, and dynamic changes in emotions. The specific input and initialization process is as follows: First, the standardized human voice temporal sequence is written into the input buffer of the voiceprint feature clustering processing logic in chronological order. At the same time, the timestamp information bound to the effective speech segment and the audio-visual synchronization metadata identifier are written into the metadata management module of the logic to establish a one-to-one correspondence between the human voice data and the metadata. Then, using the linear time reference of the original film as a reference, the standardized human voice temporal sequence is subjected to time reference alignment verification to ensure that the time index of each sampling point in the sequence is completely matched with the playback timeline of the original film, eliminating the problem of time misalignment. Subsequently, the running parameters of the voiceprint feature clustering processing logic are initialized. Preset parameters adapted to the movie dubbing scenario, such as sliding window segmentation rules, voiceprint feature extraction dimensions, and clustering calculation rules, are loaded into the corresponding processing module of the logic to prepare for subsequent segmentation processing and feature extraction.

[0068] Preferably, the standardized human voice temporal sequence is segmented using a sliding window process based on voiceprint feature clustering logic to generate initial audio speech segment distribution blocks. The core design logic of the sliding window segmentation process is to match the continuous pronunciation characteristics of movie dialogues. By using a sliding window with overlapping intervals, the continuous standardized human voice temporal sequence is divided into multiple independent segments containing complete pronunciation content, avoiding segmentation errors across roles or sentences, and providing a foundation for the accurate extraction of voiceprint features for subsequent single roles. In the specific implementation, firstly, according to the segmentation rules after initialization, a fixed-length sliding window is set on the timeline of the standardized human voice temporal sequence. At the same time, an overlap interval is set between adjacent sliding windows. The length of the overlap interval can ensure the complete connection of the pronunciation content between adjacent windows, avoiding the problem of truncating the middle of the sentence. Then, the sliding window is controlled to slide along the timeline at a preset step size. The effective speech ratio is checked for each segment of human voice data covered by the sliding window. When the effective speech ratio in the window reaches a preset threshold, the human voice data covered by the window is marked as a candidate speech segment. Subsequently, boundary optimization processing is performed on all candidate speech segments, adjusting the start and end positions of the segments to the endpoints of the effective speech segments, ensuring that each candidate speech segment contains complete sentence pronunciation content and there is no problem of truncated pronunciation. Finally, all the boundary-optimized candidate speech segments are arranged in order according to the chronological order of the timeline to generate the initial audio speech segment distribution block.

[0069] Preferably, validity checks and redundancy processing are performed on the initial audio-speech segment distribution blocks to generate the final audio-speech segment distribution blocks. The initial audio-speech segment distribution blocks are generated by segmenting through a sliding window. These blocks may contain invalid segments that are too short, redundant segments with overlapping content, and mixed segments that cross the speech subject. These segments will directly affect the accuracy of subsequent voiceprint feature extraction, so targeted checks and processing are performed. In the specific implementation, the duration of each initial audio-speech segment distribution block is first checked, and invalid short segments with a duration below a preset threshold are removed. These segments are mostly interjections and brief noises between pronunciations, which do not have complete voiceprint features and cannot support effective recognition of the speaker. Then, the overlap of the remaining initial audio-speech segment distribution blocks is checked. For adjacent segments with an overlap ratio exceeding a preset threshold, the segments with more complete effective speech content are retained, and redundant overlapping segments are removed to avoid the same speech content being processed repeatedly. Subsequently, the pronunciation continuity of the remaining segments is checked, and mixed segments with obvious pronunciation pauses and energy abrupt changes are identified and split. These segments are likely to contain alternating pronunciations from multiple different characters. After splitting, it is ensured that each segment contains only continuous single-person pronunciation content. Finally, all the verified and processed segments are arranged in chronological order, and each segment is bound with corresponding timestamp information and audio-visual synchronization metadata identifiers to generate audio-speech segment distribution blocks.

[0070] Preferably, the pre-construction and parameter solidification of a voiceprint feature extractor for film dubbing scenarios are completed, providing a processing foundation for subsequent acoustic statistical modeling and voiceprint feature extraction. The voiceprint feature extractor is a feature encoding structure specifically designed for the characteristics of film characters' voices. Unlike general voiceprint extraction structures, it can effectively filter feature interference caused by changes in character emotions, dialogue content, and volume in film dialogue, accurately extracting identity attribute features strongly correlated with the speaker's physiological characteristics, and adapting to the voiceprint recognition needs of multiple characters and emotions in film dubbing scenarios. The pre-construction process involves first building a hierarchical processing structure for the voiceprint feature extractor. This structure comprises three core processing layers: a low-level acoustic feature extraction layer, a global statistical modeling layer, and an identity feature encoding layer. The low-level acoustic feature extraction layer extracts the basic acoustic dimensional features of the human voice. The global statistical modeling layer models the statistical distribution of the basic features and filters out dynamic features irrelevant to identity. The identity feature encoding layer maps the statistical features into a high-dimensional identity representation vector. Next, a large-scale movie character voice dataset is used to train the constructed structure. This dataset contains movie character voice samples of different genders, ages, and pronunciation styles, covering various scenarios involving changes in dialogue emotion, volume, and content. Through training, the voiceprint feature extractor learns the feature patterns in the human voice that are strongly correlated with the speaker's identity, enabling it to filter out irrelevant interference features. Finally, the parameters of the trained voiceprint feature extractor are solidified to ensure consistency between its processing and output feature dimensions, providing stable processing capabilities for subsequent batch feature extraction.

[0071] Preferably, a pre-built voiceprint feature extractor is used to perform acoustic statistical modeling on each audio speech segment distribution block to generate a global acoustic statistical representation of the corresponding segment. The core function of acoustic statistical modeling is to convert the time-domain waveform data of the audio speech segment distribution block into a statistical feature distribution that can characterize the speaker's identity attributes, eliminating the influence of irrelevant factors such as dialogue content and vocal emotion, and providing core statistical basis for subsequent voiceprint embedding feature extraction. In the specific implementation, low-level acoustic feature extraction is first performed on the temporal waveform data of each audio speech segment distribution block. The waveform data is processed into frames at fixed time intervals, and multiple dimensions of basic acoustic features are extracted from each frame of waveform data. These basic acoustic features can completely characterize the core acoustic attributes of the human voice in that frame, such as the spectral distribution, formant structure, and dynamic changes in pronunciation. Then, global statistical modeling is performed on the basic acoustic features of all frames within the entire audio speech segment distribution block. The statistical distribution parameters of each basic acoustic feature dimension in the entire segment are calculated, including the mean, variance, extreme value distribution, and dynamic change trend of the features. Through global statistical modeling, random fluctuations of single-frame features are eliminated, while dynamic feature components that change with the content of the dialogue and emotions are filtered out, and stable feature distributions related to the physiological structure of the speaker are retained. Finally, the statistical distribution parameters of all dimensions are combined in an orderly manner to generate a global acoustic statistical representation that uniquely corresponds to the audio speech segment distribution block. Each global acoustic statistical representation is bound to the timestamp and metadata information of the corresponding segment, providing input for subsequent identity feature encoding.

[0072] Preferably, based on the global acoustic statistical representation corresponding to each audio speech segment distribution block, feature mapping processing is performed using the identity feature encoding layer of the voiceprint feature extractor to extract and generate the voiceprint embedding feature attributes of the respective speaking subject. Voiceprint embedding feature attributes are high-dimensional feature vectors representing the speaker's identity. Voiceprint embedding feature attributes corresponding to different dialogue segments of the same speaking subject have high similarity, while voiceprint embedding feature attributes of different speaking subjects have significant distinguishability, serving as the core basis for subsequent role clustering and identity attribution labeling. In the specific implementation, the global acoustic statistical representation corresponding to each audio speech segment distribution block is first input into the identity feature encoding layer of the voiceprint feature extractor. Then, using the pre-trained multi-layer encoding structure in the identity feature encoding layer, a high-dimensional nonlinear mapping process is performed on the global acoustic statistical representation to gradually filter out residual feature components unrelated to the speaker's identity and strengthen identity feature components strongly correlated with the speaker's physiological characteristics. Finally, a high-dimensional feature vector with a fixed dimension is output, which is the voiceprint embedding feature attribute of the speaker to which the audio speech segment distribution block belongs. Subsequently, each voiceprint embedding feature attribute is bound to the timestamp information, audio-visual synchronization metadata identifier, and segment duration information of the corresponding audio speech segment distribution block, establishing a one-to-one correspondence between the voiceprint embedding feature attribute and the corresponding dialogue segment in the original film. Finally, feature space normalization processing is performed on all generated voiceprint embedding feature attributes to map all voiceprint embedding feature attributes to a unified high-dimensional feature space, ensuring that the voiceprint embedding feature attributes of different segments have a unified measurement standard, providing standardized identity feature data that can be directly compared for subsequent clustering processing.

[0073] Optionally, based on the voiceprint embedding feature attributes, a role-identified voice track attribute mapping detail table for recording voice attribution relationships is constructed, including: The voiceprint embedding feature attributes are mapped to their respective feature space to generate mapped voiceprint embedding feature vectors, and the geometric distance between each mapped voiceprint embedding feature vector is calculated to generate geometric distance distribution measurements. The cluster centers corresponding to each audio speech segment distribution block are determined based on the geometric distance distribution measurements. The cluster centers are introduced into the voiceprint comparison space and matched with the preset character identity feature library after feature denoising to determine the associated film character identity tags to which each audio speech segment distribution block belongs. Based on the associated film character identity tags, the decoupled pure voice stream data components are reorganized spatially based on the character affiliation attribute to construct a character-identified voice track attribute mapping detail table for recording voice affiliation relationships.

[0074] Preferably, all voiceprint embedding feature attributes generated in the preceding steps are subjected to feature space normalization mapping to generate mapped voiceprint embedding feature vectors. Voiceprint embedding feature attributes are high-dimensional identity features corresponding to each audio speech segment distribution block. Their generation process is based on a pre-built voiceprint feature extractor and needs to be uniformly mapped to the normalized voiceprint feature space fixed during the training of the voiceprint feature extractor. This ensures that the dimensional specifications and measurement benchmarks of all features are completely consistent, providing a reliable foundation for subsequent similarity comparison and clustering processing. In the specific implementation, the dimension of each voiceprint embedding feature attribute is first checked to see if it fully matches the preset voiceprint feature space dimension. For features with mismatched dimensions, linear interpolation or dimension pruning is performed to ensure that all features have a uniform dimensional specification. Then, norm normalization is performed on each voiceprint embedding feature attribute to uniformly map the modulus of all features to a fixed value, eliminating amplitude differences caused by changes in pronunciation volume and emotion in different segments, so that the geometric distance between different voiceprint embedding feature attributes can be directly used for similarity comparison of the speaking subjects. Finally, all features that have undergone dimension matching and normalization are arranged in order according to the time sequence of the corresponding audio speech segment distribution blocks to generate mapped voiceprint embedding feature vectors. Each mapped voiceprint embedding feature vector is bound to the timestamp information and audio-visual synchronization metadata identifier of the corresponding audio speech segment distribution block, providing standardized input under a unified metric for subsequent geometric distance calculation.

[0075] Preferably, based on all mapped voiceprint embedding feature vectors, geometric distance calculation is performed between each pair of vectors to generate geometric distance distribution measurements. The geometric distance calculation adopts the cosine distance calculation rule adapted to the voiceprint feature space. Cosine distance can accurately characterize the directional similarity of two high-dimensional feature vectors in the feature space, and is not affected by the absolute amplitude of the features. It can effectively adapt to the voiceprint feature comparison needs of the same character under different emotions, different lines, and different pronunciation volumes in movie scenes, and is different from the shortcomings of traditional Euclidean distance, which is easily affected by feature amplitude interference. In the specific implementation, firstly, pairwise feature vector combinations are constructed. Each mapped voiceprint embedding feature vector is paired with all other mapped voiceprint embedding feature vectors to form independent comparison pairs. Then, for each comparison pair, the cosine distance between the two feature vectors is calculated. The smaller the distance value, the higher the probability that the two feature vectors correspond to the same speaker; the larger the distance value, the higher the probability that the speakers are different speakers. Subsequently, the cosine distance values ​​corresponding to all comparison pairs are arranged in order according to the time index and pairing relationship of the corresponding feature vectors to form a two-dimensional distance distribution matrix. The rows and columns of this matrix correspond to the indices of the mapped voiceprint embedding feature vectors, and the value at the intersection of the row and column is the geometric distance between the two corresponding feature vectors. Finally, the distance distribution matrix is ​​statistically verified to remove abnormal distance values ​​that exceed the reasonable range, and missing pairing distance data is supplemented to generate complete geometric distance distribution measurements, providing core quantitative basis for subsequent cluster center determination.

[0076] Preferably, density-based unsupervised clustering is performed based on geometric distance distribution measurements to determine the cluster centers corresponding to each audio segment distribution block. This clustering process differs from traditional clustering methods that preset the number of clusters, automatically adapting to the number of characters appearing in the film without needing to know the total number of characters in advance, making it more suitable for batch processing needs of unknown film sources in film dubbing scenarios. In the specific implementation, firstly, based on the geometric distance distribution measurement values, a corresponding neighborhood range is set for each mapped voiceprint embedding feature vector. The radius of the neighborhood range is adaptively determined according to the statistical distribution results of the geometric distance distribution measurement values. The number of feature vectors contained within the neighborhood range is the density value of that feature vector. Then, core feature vectors with density values ​​higher than a preset density threshold are selected. These core feature vectors correspond to dense regions of voiceprint feature distribution in the feature space, representing the stable voiceprint feature distribution of a certain role. Subsequently, using the core feature vectors as the core, all feature vectors within the neighborhood range are divided into the same cluster. At the same time, different clusters with overlapping neighborhoods are merged to eliminate the problem of the same role being split into multiple clusters due to changes in pronunciation state. Finally, for each final cluster, the mean vector of all mapped voiceprint embedding feature vectors within the cluster is calculated. This mean vector is the cluster center corresponding to the cluster. Each cluster center uniquely corresponds to a pronunciation subject. At the same time, the timestamp and metadata information of the audio speech segment distribution blocks corresponding to all feature vectors within the cluster are bound to each cluster center, providing a core identity representation benchmark for subsequent role identity matching.

[0077] Preferably, feature denoising processing is performed on all generated cluster centers to generate clean cluster center features that meet the role matching requirements, while simultaneously calling and initializing the preset role identity feature library. The preset role identity feature library is a standardized voiceprint feature library pre-built for the target film. The library stores standard voiceprint feature vectors corresponding to each major character in the film, as well as basic character information corresponding to the standard voiceprint feature vectors, including character name, character number, dubbing reference information, etc. The standard voiceprint feature vectors in the library and the mapped voiceprint embedded feature vectors are in the same feature space and have completely consistent dimensions and measurement benchmarks. The specific process of feature denoising is as follows: First, for each cluster center, backtrack all mapped voiceprint embedding feature vectors within the corresponding cluster, calculate the geometric distance between each feature vector and the cluster center, and remove outlier feature vectors whose geometric distance exceeds a preset distance threshold. These outlier feature vectors are likely invalid features caused by pronunciation interjections or environmental noise. Then, for the clusters after removing outliers, recalculate the mean vector of the remaining feature vectors within the cluster to generate clean cluster center features that have been denoised and optimized. Finally, arrange all clean cluster center features in order according to the cluster number, and simultaneously load and initialize the role identity feature library, mapping all standard voiceprint feature vectors in the library to a unified comparison metric space, providing a noise-free comparison benchmark and matching reference for subsequent similarity metric matching.

[0078] Preferably, the pure cluster center features are introduced into the voiceprint comparison space and subjected to similarity metric matching with the standard voiceprint feature vectors in the preset role identity feature library to determine the associated movie role identity tags to which each audio speech segment distribution block belongs. The voiceprint comparison space is a high-dimensional metric space that is completely aligned with the voiceprint feature space, which can ensure that the metric benchmark of the comparison process is completely consistent and avoid the problem of decreased matching accuracy due to spatial misalignment. In specific implementation, each pure cluster center feature is first matched with all standard voiceprint feature vectors in the role identity feature library. For each matching pair, the cosine similarity between the two feature vectors is calculated. The higher the cosine similarity value, the higher the matching degree between the pure cluster center feature and the corresponding standard voiceprint feature vector, and the higher the probability of corresponding to the same role. Then, for each pure cluster center feature, the standard voiceprint feature vector with the highest cosine similarity value is selected. It is determined whether the highest cosine similarity value is higher than the preset matching threshold. If it is higher than the matching threshold, the role information corresponding to the standard voiceprint feature vector is taken as the pure cluster center feature. The cluster center feature corresponds to the associated film character identity tags of all audio speech segment distribution blocks within the cluster. If the highest cosine similarity value is lower than the matching threshold, an independent temporary character identity tag is generated for the cluster corresponding to the pure cluster center feature. This tag is used to mark the voice content of minor characters such as supporting actors and extras in the film who are not included in the character identity feature database. Finally, the corresponding associated film character identity tag is bound to each audio speech segment distribution block. At the same time, the tag information is bound and associated with the timestamp and audio-visual synchronization metadata identifier of the corresponding segment to ensure that each voice content has a clear character attribution mark, providing a clear attribution basis for subsequent voice stream data reassembly.

[0079] Preferably, based on the associated film character identity tags bound to each audio segment distribution block, spatial dimension reorganization processing based on character affiliation attributes is performed on the decoupled pure voice stream data components to generate multi-channel character-separated voice data streams categorized by character affiliation. The decoupled pure voice stream data components are temporal voice waveform data covering the entire duration of the film. Each audio segment distribution block corresponds to a unique time interval in the decoupled pure voice stream data components. The core of spatial dimension reorganization is to reorganize the voice data, which was originally arranged linearly in chronological order, into channels according to character affiliation attributes, thereby integrating all the voice content of the same character into the same channel. This differs from the traditional solution's processing logic, which can only output a single mixed voice track. In the specific implementation, an independent voice data channel is first assigned to each unique associated film character identity tag, with each channel corresponding to an independent timeline that is perfectly aligned with the linear time reference of the original film. Then, based on the time interval corresponding to each audio segment distribution block and the associated film character identity tag, the voice waveform data of the corresponding time interval in the decoupled pure voice stream data component is written to the same time position of the voice data channel corresponding to the tag. For time intervals without voice content, silence data is written to ensure the integrity of the channel timeline. Subsequently, a smooth transition processing is performed on the waveform data in each voice data channel, and fade-in and fade-out processing is performed at the connection positions between adjacent audio segments to eliminate waveform abrupt changes and noise caused by segment splicing. Finally, a multi-channel character-separated voice data stream is generated, categorized by character affiliation and perfectly aligned with the original film timeline. The voice data in each channel fully contains all the dialogue content of the corresponding character in the film, providing the core track-separated data foundation for the final character-identified voice track-separated attribute mapping detail table.

[0080] Preferably, based on multi-channel character-tracked voice data streams, associated film character identity tags, and corresponding audio-visual synchronization metadata identifiers, structured encapsulation and information integration processing are performed to construct a character-identified voice track attribute mapping detail table for recording voice attribution relationships. This detail table is a structured data file that can be directly accessed in the film dubbing process, fully recording the character attribution, time position, audio-visual synchronization information, track data index, and other core content of all voice segments in the film, which can directly support the accurate retrieval of the target character's voice trajectory during the dubbing process. The specific construction process is as follows: First, a structured field system for the detailed table is defined. This system includes eight core fields: role identifier, role identity information, voice segment start timestamp, voice segment end timestamp, audio-visual synchronization metadata identifier, track-by-track data channel index, voice segment duration, and voiceprint feature matching degree. Each field has a defined unified data format and filling rules. Next, a corresponding detailed table entry is generated for each audio segment distribution block. Following the defined field system, the associated film role identity tag, basic role information, timestamp information, and audio-visual synchronization metadata identifier for that segment are filled in sequentially. The system includes the index, segment duration, and voiceprint matching similarity values ​​for the corresponding voice data channels. Then, all entries in the detailed table are double-sorted according to the role identifier field and the start timestamp field. First, they are grouped by role identifier, and then all entries for the same role are arranged in chronological order to ensure that all voice segments for the same role are presented in the order of the movie playback. Finally, a data integrity check is performed on the complete detailed table to verify that the field information of each entry is complete and without missing information, that the timestamp information is perfectly aligned with the original movie time base, and that the role identifier corresponds one-to-one with the track-specific data channel index. This results in a verified, role-identified voice track-specific attribute mapping detailed table.

[0081] Optionally, based on the character-identified human voice track-splitting attribute mapping details table and the background sound environment energy feature parameter sequence, an automated film dubbing track-splitting audio control command stream data is generated to characterize the predicted optimal track-splitting state of the original film audio and can be parsed by the dubbing system, including: The character-identified human voice track attribute mapping details table and background sound environment energy feature parameter sequence are logically encapsulated into a data structure to generate a full-scene track-by-track acoustic timing sequence that includes timestamp index, character identification index and audio track physical feature distribution. Amplitude gain control based on a preset loudness standard is applied to the full-scene track-by-track acoustic timing sequence to generate a gain-controlled acoustic timing sequence. Based on the acoustic timing sequence after gain control, an automated film dubbing track splitting audio control command stream is generated to characterize the predicted optimal track splitting state of the original film audio and can be parsed by the dubbing system.

[0082] Preferably, a unified temporal reference alignment check and preprocessing are first performed on the character-identified voice track attribute mapping details table and the background sound environment energy feature parameter sequence to generate a standardized input dataset that meets the requirements of logical encapsulation. Both the character-identified voice track attribute mapping details table and the background sound environment energy feature parameter sequence are generated based on the same original film and share the linear time reference of the original film. However, during their respective generation processes, there may be problems such as temporal index misalignment, time interval mismatch, and missing metadata associations, which will directly affect the audio-visual synchronization accuracy of the subsequent encapsulated sequence and fail to meet the timeline accuracy requirements of the film dubbing process. Therefore, reference alignment and preprocessing must be completed first. In the specific implementation, the linear time base of the original film is used as the sole reference standard to verify whether the timestamp index of all voice segments in the character-identified voice track attribute mapping details table completely matches the time axis coverage of the background sound environment energy feature parameter sequence. For segments with temporal misalignment, linear interpolation is performed for alignment, and missing parts of the time interval are supplemented with corresponding silence markers and blank energy parameters to eliminate the benchmark deviation in the time dimension. Then, the metadata system of the two sets of data is unified and standardized. The character identification information and audio-visual synchronization metadata identifier in the character-identified voice track attribute mapping details table are associated and bound with the scene energy distribution information and frame synchronization markers in the background sound environment energy feature parameter sequence to ensure that the voice data and background sound data in the same time interval have a one-to-one metadata association relationship. Finally, the two sets of data that have been time-aligned and metadata-bound are arranged in an orderly manner according to a unified time axis order to generate a standardized input dataset, which provides a unified time base and complete metadata association for subsequent data structure logic encapsulation.

[0083] Preferably, the character-identified voice track attribute mapping detail table is subjected to structured parsing and multi-channel voice temporal sequence reconstruction to generate a character-tracked voice temporal data stream that is completely aligned with the original film timeline. The character-identified voice track attribute mapping detail table is structured entry-based data, with each entry corresponding to a single segment of voice content for a character. It includes core fields such as character identifier, timestamp, and audio-visual synchronization information. Reconstructing the discrete entry-based data into a continuous, character-channel-based temporal data stream is necessary to meet the requirements of subsequent full-scene sequence encapsulation. In the specific implementation, the character-identified voice track attribute mapping detail table is first fully parsed to extract all unique character identifier information. An independent voice data channel is assigned to each unique character identifier, with the timeline of each channel perfectly aligned with the linear time base of the original film, and the channel length perfectly matching the total duration of the original film. Then, based on the character identifier and timestamp information corresponding to each entry in the detail table, the decoupled clean voice stream data components within the corresponding time interval are written to the same time position of the corresponding voice data channel for that character. For time intervals without corresponding voice content, data conforming to industry standards is written. Accurate silence data is used to ensure the temporal continuity and timeline integrity of each channel. Subsequently, each voice data channel is bound with a corresponding role identifier index, voice segment timestamp index, and audio-visual synchronization metadata identifier. At the same time, the physical feature distribution of voice data in each channel is statistically analyzed, including core physical features such as temporal amplitude distribution, frequency domain energy distribution, and loudness change trend. Finally, the voice data channels corresponding to all roles are integrated to generate a multi-channel, temporally continuous, role-tracked voice temporal data stream with complete role identifiers and physical feature information, providing core data support for the voice dimension for subsequent full-scene sequence encapsulation.

[0084] Preferably, the background sound environmental energy feature parameter sequence is subjected to temporal normalization and dimensional alignment processing to generate a standardized background sound temporal data stream that is fully adapted to the character-tracked voice temporal data stream. The background sound environmental energy feature parameter sequence contains the full-time energy distribution information of background music and environmental sound effects in the original film. Only by completing temporal normalization, dimensional alignment, and metadata binding can it be seamlessly encapsulated with the character-tracked voice temporal data stream, avoiding problems such as temporal misalignment and dimensional mismatch between voice and background sound. In the specific implementation, the background sound environmental energy feature parameter sequence is first subjected to temporal continuity verification. According to the linear time base of the original film, the time sampling points in the sequence are interpolated and completed to eliminate the problems of uneven time sampling intervals and missing data points in the sequence, ensuring that the time sampling accuracy of the background sound sequence is completely consistent with the sampling accuracy of the character track voice temporal data stream. Then, the background sound environmental energy feature parameter sequence is subjected to dimension alignment processing, adjusting the time axis length and the number of sampling points of the sequence to completely match the character track voice temporal data stream, ensuring that the voice data of each time sampling point can correspond to a unique background sound energy parameter. Subsequently, the background sound environmental energy feature parameter sequence is bound with the corresponding timestamp index, scene type mark, and frame synchronization mark. At the same time, the physical feature distribution of the background sound sequence is statistically analyzed, including the energy change trend throughout the time period, frequency band energy distribution, peak level and other core physical features. Finally, a standardized background sound temporal data stream with temporal continuity, dimension matching and complete metadata and physical feature information is generated, providing core data support for the background sound dimension for subsequent full-scene sequence encapsulation.

[0085] Preferably, the character-specific audio timing data stream and the standardized background sound timing data stream are logically encapsulated using a data structure to generate a full-scene, track-specific acoustic timing sequence that includes timestamp indexes, character identifier indexes, and the distribution of physical characteristics of the audio tracks. This logical encapsulation differs from the traditional simple splicing of audio data. It adopts a structured encapsulation logic of "timing as the framework, track-specific as the objective, and metadata as the association," which allows the encapsulated sequence to simultaneously possess the temporal continuity of multi-track audio, the independent callability of character-specific tracks, and the metadata association of audio-visual synchronization, fully adapting to the needs of calling and editing track-specific audio in the film dubbing process. In the specific implementation, a unified data structure for the full-scene multi-track acoustic timing sequence is first defined. This structure uses time sampling points as the smallest unit. Within each time sampling point's structural unit, there is a timestamp index corresponding to that time point, amplitude data for all character voice channels, energy parameter data for the background sound channel, the identifier index of the corresponding character, audio-visual synchronization metadata identifier, and physical characteristic parameters of each audio track at that time point. Then, the character multi-track voice timing data stream and the standardized background sound timing data stream are written sequentially into the corresponding positions of the unified data structure according to the order of the time sampling points, ensuring that the voice data, background sound data, and metadata information at each time sampling point are completely corresponding without misalignment. No missing data was found. Subsequently, a full-time integrity check was performed on the packaged sequence. The timeline coverage of the check sequence was found to be completely consistent with the original film. The data continuity of each character channel met the requirements, and the metadata information was completely matched with the screen content at the corresponding time point. At the same time, the complete physical feature distribution data of each audio track in the entire sequence was extracted and added to the structure header of the sequence to form a fast-read audio track physical feature index. Finally, a full-scene track-by-track acoustic timing sequence with complete structure, continuous timing, clear track division, and complete metadata association was generated. This sequence can directly support the dubbing system to quickly retrieve the corresponding audio track data by character and time interval, providing a standardized processing object for subsequent amplitude gain control.

[0086] Preferably, the process involves determining a preset loudness benchmark for adaptation to film dubbing industry standards, and conducting full-track loudness statistical analysis of the acoustic timing sequence across all scenes, in order to generate a set of quantized loudness parameters required for gain control. Film dubbing has clear industry standards for audio loudness, and loudness standards differ across distribution channels and playback scenarios. Furthermore, maintaining a loudness balance between dialogue and background noise that conforms to listening habits is crucial. Therefore, determining the preset loudness standard for adaptation and completing loudness statistics for the entire sequence provides an accurate quantitative basis for subsequent amplitude gain control. In the specific implementation, firstly, based on the distribution channels and playback scenarios of the target film, the corresponding loudness standards of the film dubbing industry are retrieved to determine the core parameters of the preset loudness standards, including key indicators such as the overall target loudness, the target loudness of human dialogue, the loudness tolerance of background sound, the upper limit of peak level, and the dynamic range limit of loudness. Then, the full-scene track-by-track acoustic timing sequence is split into independent character voice tracks and background sound tracks. For each independent track, loudness statistics are calculated for the entire time period according to the preset time statistical window to obtain core statistical data such as the current average loudness, loudness dynamic change curve, peak level, and loudness distribution range of each track. Subsequently, for the full-scene track-by-track acoustic timing sequence, the statistical data of the overall mixed loudness and the loudness difference distribution between human dialogue and background sound are calculated to clarify the deviation between the current loudness level and the preset loudness standard. Finally, the loudness statistics of all tracks, the deviation from the preset standard, and the loudness dynamic change curve are integrated into a quantized loudness parameter set to provide an accurate quantitative calculation basis for subsequent amplitude gain control.

[0087] Preferably, based on a quantized loudness parameter set, independent target gain value calculations and dynamic gain curve generation are performed for each track to generate a gain control scheme adapted to each audio track. This process employs a logic of independent calculation for each track and overall balance constraints, which differs from the traditional method of uniform gain adjustment for the entire audio. It can ensure that each audio track meets the loudness standard while maintaining a reasonable loudness balance between the character's voice and the background sound, avoiding problems such as dialogue being masked by background sound and excessive loudness differences between different characters' voices that affect the translation effect. In the specific implementation, firstly, for each character's voice track, based on the deviation between the current average loudness of the track and the target loudness of the dialogue in the quantized loudness parameter set, the base fixed gain value of the track is calculated to ensure that the average loudness of the track meets the target loudness requirement of the dialogue. Then, for the background audio track, based on the current average loudness of the background sound and the background sound loudness tolerance in the quantized loudness parameter set, the base fixed gain value of the background audio track is calculated. Simultaneously, combined with the target loudness difference between the dialogue and the background sound, the base gain value of the background sound is constrained and adjusted to ensure that the loudness difference between the dialogue and the background sound conforms to the electrical requirements. The audiophile experience is assessed. Then, for each track's loudness dynamic change curve, dynamic gain adjustment calculations are performed. For time intervals where loudness fluctuations exceed a preset range, the dynamic gain value changes over time, smoothing loudness abrupt changes within the track. Simultaneously, it is ensured that the peak level after dynamic gain adjustment does not exceed the preset peak level upper limit to avoid clipping distortion. Finally, the basic fixed gain value of each track is combined with the time-varying dynamic gain value to generate a continuous gain control curve for each track, forming a complete track-by-track gain control scheme, providing accurate control basis for subsequent amplitude gain adjustments.

[0088] Preferably, based on a track-by-track gain control scheme, the full-scene track-by-track acoustic timing sequence is subjected to track-by-track amplitude gain adjustment and smooth transition processing to generate a gain-controlled acoustic timing sequence. In specific implementation, the full-scene track-by-track acoustic timing sequence is first split into independent character voice tracks and background sound tracks. For each track, the gain value of the corresponding gain control curve is applied to the audio amplitude data of that track at time-series sampling points, completing the full-time amplitude gain adjustment of that track. Then, for each track after gain adjustment, smoothing processing of the gain transition region is performed. For the connection points between fixed and dynamic gains, and the time intervals where dynamic gain changes, a smoothing window function is used to gradually transition the gain value, eliminating audio jumps and noise caused by sudden gain changes, ensuring a smooth and natural audio listening experience. Subsequently, all tracks that have undergone gain adjustment and smoothing processing... The system performs loudness and distortion checks on each track, verifying that the final loudness of each track meets the preset loudness standard, the peak level does not exceed the preset upper limit, and there are no issues such as clipping distortion or harmonic distortion. It also verifies that the loudness balance between human dialogue and background sound meets the preset requirements. Finally, all the checked tracks are logically repackaged according to the original data structure of the full-scene track-by-track acoustic timing sequence, supplementing and updating the physical feature distribution data and loudness parameter information of the tracks, while retaining the original timestamp index, character identification index, and audio-visual synchronization metadata identifier. The final result is an acoustic timing sequence with gain control that meets the loudness standard for film dubbing, has clear track division, continuous timing, and a balanced listening experience.

[0089] Optionally, based on the gain-controlled acoustic timing sequence, an automated film dubbing track splitting audio control command stream data is generated to characterize the predicted optimal track splitting state of the original film audio and can be parsed by the dubbing system, including: Determine the energy balance relationship between different character tracks based on the acoustic timing sequence after gain control; The gain-controlled acoustic timing sequence is encapsulated using a transmission protocol based on the movie dubbing standard to generate encapsulated acoustic transmission data packets. Based on the energy balance relationship, the encapsulated acoustic transmission data packet is physically resampled to generate resampled acoustic data. The resampled acoustic data is then encoded and converted to generate automated film dubbing track-splitting audio control command stream data that characterizes the predicted optimal track splitting state of the original audio and can be parsed by the dubbing system.

[0090] Preferably, the gain-controlled acoustic timing sequence is first subjected to track-by-track structured parsing and timing consistency verification to generate a standardized track-by-track audio dataset suitable for energy balance analysis. The gain-controlled acoustic timing sequence is a time-series audio data with a multi-channel track structure and complete metadata association. It contains multiple independent character voice tracks and a unified background sound track. Structured parsing and verification must be completed first to provide an accurate and compliant input basis for subsequent calculations of the energy balance relationship between character voice tracks. In the specific implementation, firstly, according to the original data structure of the full-scene segmented acoustic timing sequence, a full analysis is performed on the gain-controlled acoustic timing sequence to separate all independent character voice tracks and background sound tracks within the sequence, as well as character identifier indexes, timestamp indexes, audio-visual synchronization metadata identifiers, and audio track physical feature distribution data that are bound to each track. Then, using the linear time reference of the original film as the sole reference, timing consistency checks are performed on all the segmented audio tracks to verify that the time axis length, number of sampling points, and timestamp index of each track are completely matched, without timing misalignment, missing sampling points, or audio track disconnection issues. For audio tracks with abnormalities, interpolation is performed to complete and misalignment correction is performed. Subsequently, the audio sampling data of all tracks are normalized to unify the sampling bit depth and numerical encoding format of each track, eliminating energy calculation deviations caused by format differences. Finally, all the verified, corrected, and normalized audio track data, along with the corresponding character identifiers and metadata information, are integrated to generate a standardized segmented audio dataset, providing a core analysis object for determining the energy balance relationship between character audio tracks in the future.

[0091] Preferably, based on a standardized multi-track audio dataset, multi-dimensional audio track energy statistics and cross-analysis are performed to determine the energy balance relationship between different character audio tracks. This energy balance relationship is a quantitative constraint rule adapted to the auditory requirements of film dubbing scenarios. Its core purpose is to ensure that the loudness of different characters' dialogues conforms to narrative priorities, while maintaining the auditory clarity of overlapping dialogue scenes, as well as a reasonable loudness ratio between human voice dialogue and background sound. This differs from the simple logic of simply unifying the overall loudness in general audio processing, and fully meets the dubbing and mixing production needs of film dubbing. In the specific implementation, firstly, for each character's voice track in the standardized multi-track audio dataset, full-time energy statistics calculation is performed. According to a fixed time statistical window, the full-time energy distribution curve, average energy value, peak energy level, dialogue duty cycle, and other core energy statistical parameters of each character's voice track are obtained. Then, for the time intervals in the film where multiple characters' dialogues overlap, cross-energy analysis is performed to calculate the energy proportion and loudness difference of different character's voice tracks in the overlapping interval, clarifying the energy distribution relationship in overlapping dialogue scenes. Subsequently, combined with the narrative priority of the characters in the film, corresponding energy weight coefficients are set for different character's voice tracks. The core protagonist's voice track is given a higher priority weight, while the supporting characters' and extras' voice tracks are given appropriate auxiliary weights. At the same time, combined with the listening standards of the film dubbing industry, the baseline energy difference range between human voice dialogue and background sound tracks is determined. Finally, the energy statistical parameters of each character's voice track, the energy distribution rules of overlapping scenes, the character priority weight coefficients, and the baseline energy difference between human voice and background sound are integrated into a complete quantitative constraint system, forming an energy balance relationship between different character's voice tracks, providing a constraint basis for energy consistency in the subsequent waveform resampling and encoding conversion process.

[0092] Preferably, the adaptation of film translation standards and the pre-configuration of transmission protocols corresponding to the target distribution scenario are completed, providing a standardized encapsulation framework and parameter system for subsequent protocol encapsulation processing. The film translation process involves differentiated industry standards and transmission protocol specifications for different distribution channels and playback scenarios. These different specifications have specific requirements for audio track encapsulation structures, metadata field definitions, sampling format standards, and verification mechanisms. Standard adaptation and protocol configuration must be completed in advance to ensure that the encapsulated data packets can be directly and compatiblely parsed by the target translation and playback systems. In specific implementation, firstly, based on the target film's distribution channels and application scenarios, the corresponding film translation industry audio and video transmission standards are retrieved, clarifying the core requirements for multi-track audio encapsulation in the standards, including the upper limit of the number of audio track channels, character audio track identification rules, audio-visual synchronization timestamp accuracy, sampling format and sampling rate range, and required fields for metadata. Then, based on the adapted industry standards, the encapsulation framework of the transmission protocol is pre-configured, defining the layered encapsulation structure of the protocol, including five core layers: the basic audio frame layer, the audio track index layer, the character metadata layer, the translation control parameter layer, and the verification and error correction layer. Simultaneously, the field definitions, data lengths, encoding rules, and mapping relationships for each level are clearly defined. Subsequently, based on the pre-configured encapsulation framework, the operating parameters of the protocol encapsulation are set, including the encapsulation duration of audio frames, frame synchronization code rules, verification algorithm type, and metadata embedding method. At the same time, the energy balance relationship between different character audio tracks, which was previously determined, is incorporated into the predefined fields of the translation control parameter layer to ensure that the energy balance constraint rules can be transmitted synchronously with the encapsulated data packets. Finally, a complete transmission protocol encapsulation configuration scheme adapted to the target film translation standard is formed, providing a standardized execution basis for subsequent protocol encapsulation processing.

[0093] Preferably, based on the transmission protocol encapsulation configuration scheme, a layered protocol encapsulation process is performed on the acoustic timing sequence after gain control to generate encapsulated acoustic transmission data packets. This encapsulation process differs from traditional simple audio file encapsulation. It adopts a layered encapsulation logic of "audio data as the core, metadata as the association, and control parameters as constraints," which enables the encapsulated data packets to simultaneously possess lossless transmission capabilities for multi-track audio, accurate identification of character tracks, and editing control capabilities in the translation process, fully adapting to the business needs of the entire film translation workflow. In the specific implementation, the continuous audio data in the acoustic timing sequence after gain control is first segmented according to the pre-configured encapsulation frame duration to generate multiple continuous audio encapsulation frames. For each audio encapsulation frame, the corresponding frame synchronization code, timestamp index, and sampling format information are written to complete the encapsulation of the basic audio frame layer. Then, according to the role identifier index, the audio encapsulation frames of different role tracks are classified and collected to construct an audio track index layer. A unique track ID is assigned to each independent role track and background sound track, and the physical feature distribution data and basic role information of the corresponding track are written to establish a one-to-one mapping relationship between track IDs and audio encapsulation frames. Subsequently, in the role element... In the data layer, the character identifier, audio-visual synchronization metadata identifier, and start and end time information of the vocal segments corresponding to each audio track are written. In the translation control parameter layer, the energy balance relationship between different character audio tracks, loudness control parameters, and track editing permission information are written. Finally, at the end of the encapsulated data packet, the verification and error correction code calculated based on the full data is written to complete the encapsulation of the verification and error correction layer. At the same time, the protocol compliance verification is performed on the encapsulated full data packet to verify that the structure, fields, and format of the data packet fully comply with the requirements of the target film translation standard. Finally, a structurally complete, compliant, and compatible encapsulated acoustic transmission data packet with complete control attributes is generated.

[0094] Preferably, based on the energy balance relationship between different character audio tracks, parameter adaptation calculations for physical waveform resampling are performed to generate a set of quantization control parameters for resampling processing. The core objective of physical waveform resampling is to adapt the audio data to the sampling rate standards of the target translation and playback systems. However, problems such as audio energy changes, detail loss, and phase distortion can easily occur during resampling, which can disrupt the previously determined energy balance relationship between character audio tracks. Therefore, it is necessary to perform resampling parameter adaptation calculations in advance based on the energy balance relationship to ensure that the energy balance relationship of each audio track remains stable during the resampling process. In the specific implementation, the basic parameters such as the target sampling rate, target sampling bit depth, and core indicators of anti-aliasing filtering are first determined according to the target film translation standards and the technical requirements of the target system. Then, based on the energy balance relationship between different character audio tracks, the corresponding resampling gain compensation coefficient is calculated for each character audio track. For supporting character audio tracks and extras audio tracks with lower average energy, appropriate gain compensation coefficients are set to avoid the suppression of weak dialogue details during the resampling process. A phase protection coefficient is set for the main character audio track to ensure the phase consistency and auditory clarity of the main character's dialogue. Subsequently, for the background sound track, the filter threshold and gain constraint coefficient for background sound resampling are calculated based on the reference energy difference between human voice and background sound to avoid the energy of the background sound exceeding the preset tolerance range after resampling, thus disrupting the loudness balance between human voice and background sound. Finally, the basic parameters of resampling, the gain compensation coefficients of each audio track, the phase protection coefficients, and the filter constraint parameters are integrated into a set of quantized control parameters to provide accurate control basis for subsequent physical waveform resampling processing, ensuring that the energy balance relationship between each audio track is not disrupted after resampling.

[0095] Preferably, based on the resampling quantization control parameter set, physical waveform resampling processing is performed on the encapsulated acoustic transmission data packets to generate waveform-resampled acoustic data. In the specific implementation, the encapsulated acoustic transmission data packets are first decapsulated to extract the original waveform sampling data, corresponding character identifier index, and timestamp information of each audio track. Simultaneously, the energy balance parameters within the data packets are extracted and cross-validated with the resampling quantization control parameter set to ensure parameter consistency. Then, for the original waveform sampling data of each audio track, anti-aliasing filtering is performed according to the basic parameters in the quantization control parameter set. High-frequency components exceeding the Nyquist frequency of the target sampling rate are filtered out to prevent aliasing distortion after resampling. During the filtering process, the filtering threshold and phase protection coefficient corresponding to each audio track are followed to ensure that dialogue details are not lost and phase shifts are not observed. Finally, the data after anti-aliasing filtering... The waveform data is processed by sampling rate conversion and bit depth matching, converting the original sampling rate and sampling bit depth to the specifications corresponding to the target standard. At the same time, according to the gain compensation coefficients corresponding to each audio track, amplitude gain compensation is performed on the resampled waveform data to correct the energy deviation that occurred during the resampling process. Finally, energy consistency verification is performed on all resampled audio track data to verify the energy distribution between each character's audio track and the energy difference between the human voice and the background sound, ensuring that it fully meets the previously determined energy balance requirements. At the same time, the waveform data is verified to be free of clipping distortion, phase distortion, noise, and dropout issues, ultimately generating acoustic data after waveform resampling that meets the target sampling standard and has a stable energy balance.

[0096] Preferably, the acoustic data after waveform resampling undergoes encoding conversion preprocessing adapted to the film dubbing scenario to generate an adaptation configuration scheme for encoding conversion. The film dubbing process has differentiated requirements for audio encoding. The dubbing editing stage uses lossless encoding formats to achieve lossless sound quality after multiple edits, while the final distribution stage adapts to lossy compression encoding formats for corresponding channels. Simultaneously, it requires the complete preservation of character track information, metadata, and control parameters during the encoding process, unlike general audio encoding which only focuses on compression ratio and sound quality. In the specific implementation, the encoding format type is first determined based on the target application scenario. For translation and editing scenarios, a lossless compression encoding format is selected, while for distribution and playback scenarios, a lossy compression encoding format corresponding to the channel standard is chosen. Simultaneously, the core encoding parameters are clarified, including compression bitrate, channel configuration, frame length settings, and encoding complexity level. Next, considering the characteristics of multi-character track separation, independent inter-track encoding rules are configured to ensure that each character's audio track and background audio track uses an independent encoding stream. This supports the translation system's independent retrieval, editing, and replacement of individual audio tracks, avoiding the loss of track information caused by multi-track mixed encoding. Subsequently, the embedding rules for metadata and control parameters are determined. Character identifier indexes, timestamp indexes, audio-visual synchronization metadata, energy balance relationships between different character audio tracks, and loudness control parameters are embedded into the corresponding auxiliary fields of the encoding stream, ensuring that the encoded data stream completely retains all translation control-related information. Finally, a complete encoding conversion and adaptation configuration scheme is formed, providing a standardized execution basis for subsequent encoding conversion processing.

[0097] Preferably, based on the encoding conversion and adaptation configuration scheme, the acoustic data after waveform resampling is encoded and finally structured and encapsulated to generate automated film dubbing track-segmentation audio control command stream data that characterizes the predicted optimal track-segmentation state of the original film audio and can be parsed by the dubbing system. In specific implementation, firstly, according to the independent inter-track encoding rules in the encoding conversion and adaptation configuration scheme, each independent character audio track and background audio track in the waveform resampling acoustic data is encoded independently to generate an independent encoded sub-stream corresponding to each audio track. During the encoding process, the configured encoding format and parameters are followed to ensure that the sound quality meets the standard requirements for film dubbing. Then, all encoded sub-streams are arranged in order according to the character identifier index, and the character metadata, audio-visual synchronization information, energy balance relationship, loudness control parameters, and track editing control commands are embedded into the auxiliary data fields of the encoded stream according to preset rules to establish a one-to-one mapping relationship between the encoded sub-streams and the metadata and control commands. Subsequently, the encoding conversion and adaptation configuration scheme is further refined. The integrated encoded stream undergoes final structured encapsulation, adding file header identifiers, track index tables, metadata index tables, and checksums to form a complete instruction stream file structure. Finally, the generated automated film dubbing track-by-track audio control instruction stream data undergoes full compliance and usability verification. This verifies that the instruction stream data can be normally parsed, retrieved, and edited by mainstream film dubbing systems; that the separation, sound quality, and loudness of each character's audio track meet preset requirements; that the character attribution information and audio-visual synchronization information are complete and accurate; and that the energy balance between different character audio tracks is stable and controllable. Ultimately, this forms automated control instruction stream data that meets the needs of the entire film dubbing process and can directly drive track-by-track audio processing in the dubbing stage.

[0098] like Figure 2 As shown in the figure, an automated track-splitting device for original audio in film dubbing is provided in an embodiment of this application, comprising: The first program unit is used to retrieve the original composite audio track data carried by the original film and perform temporal segmentation on the original composite audio track data based on the overlap sampling rule to generate a temporal audio frame sequence; and to construct an acoustic analysis initial characterization data table for low-level audio feature analysis based on the temporal audio frame sequence. The second program unit is used to inject the acoustic analytical initial representation data table as input excitation into the pre-trained acoustic source deep decoupling model for processing, so as to generate the corresponding time-frequency domain acoustic mask weight sequence; based on the time-frequency domain acoustic mask weight sequence and the acoustic analytical initial representation data table, a background sound environment energy feature parameter sequence representing the distribution of background ambient sound and music features is generated. The third program unit is used to construct a role-identified human track attribute mapping detail table for recording the human voice attribution relationship; The fourth program unit is used to generate automated film dubbing track splitting audio control command stream data based on the character-identified human voice track splitting attribute mapping details table and the background sound environment energy characteristic parameter sequence, which is used to characterize the predicted optimal track splitting state of the original film audio and can be parsed by the dubbing system.

[0099] like Figure 3 As shown in the figure, an electronic device according to an embodiment of this application includes: a memory and a processor. The memory stores a computer program, and the processor is used to run the computer program to implement an automated movie dubbing method with quality assessment and iterative feedback as described in any of the above claims.

[0100] Figures 2-3 For an exemplary description, please refer to the above. Figure 1 The illustrated embodiment.

Claims

1. A method for automated track splitting of original audio for film dubbing, characterized in that, include: Step 1: Retrieve the original composite audio track data carried by the original video and perform temporal segmentation on the original composite audio track data based on the overlap sampling rule to generate a temporal audio frame sequence; construct an acoustic analysis initial characterization data table based on the temporal audio frame sequence; Step 2: The initial acoustic representation data table is injected as input excitation into the pre-trained acoustic source deep decoupling model for processing to generate the corresponding time-frequency domain acoustic mask weight sequence. Based on the time-frequency domain acoustic mask weight sequence and the acoustic analytical initial characterization data table, a background sound environment energy feature parameter sequence is generated. Step 3: Construct a role-identified voice track attribute mapping detail table to record the voice attribution relationship; Step 4: Based on the character-identified human voice track-splitting attribute mapping details table and the background sound environment energy feature parameter sequence, generate automated film dubbing track-splitting audio control command stream data to characterize the predicted optimal track-splitting state of the original film audio and to be parsed by the dubbing system.

2. The method for automated track splitting of original audio for film dubbing according to claim 1, characterized in that, An initial acoustic characterization data table is constructed based on the temporal audio frame sequence, including: The time-domain audio frame sequence is mapped from the time axis to the frequency axis to generate complex distribution features characterizing each frame at different frequency nodes. The complex distribution features correspond to multiple time-frequency grid cells in the time-frequency domain, and each time-frequency grid cell corresponds to a complex value at the intersection of a time index and a frequency index. Based on the aforementioned complex distribution characteristics, power spectrum energy distribution data records reflecting the evolution of the original audio energy distribution on the time and frequency axes are generated; Metadata carrying image synchronization timestamps is extracted from the continuous frame sequence of the original video to generate the extracted image synchronization metadata identifier; Based on the power spectrum energy distribution data record mapped to the extracted image synchronization metadata identifier, an acoustic analysis initial characterization data table is generated.

3. The method for automated track splitting of original audio for film dubbing according to claim 2, characterized in that, Based on the power spectrum energy distribution data records mapped to the extracted image synchronization metadata identifier, an acoustic analysis initial characterization data table is generated, including: The power spectrum energy distribution data is mapped to the linear time base corresponding to the extracted image synchronization metadata identifier to generate a spatiotemporally aligned energy distribution mapping table, and the spatiotemporal alignment relationship between the audio energy fluctuation and the continuous image frame sequence is determined accordingly. The audio energy fluctuation is used to characterize the dynamic change trend of the power spectrum energy distribution data on the time axis. The spatiotemporally aligned energy distribution mapping table is subjected to feature space projection dimensionality reduction based on principal component analysis to generate a feature-dimensionally reduced power spectrum energy distribution table. The power spectrum energy distribution table after feature dimensionality reduction is subjected to feature dimension compression to generate an initial characterization data table for acoustic analysis.

4. The method for automated track splitting of original audio for film dubbing according to claim 1, characterized in that, The acoustic analytical initial representation data table is injected as input stimulus into a pre-trained acoustic source deep decoupling model for processing to generate a corresponding time-frequency domain acoustic mask weight sequence. This includes: injecting the acoustic analytical initial representation data table as input stimulus into a semantic feature extraction network in the pre-trained acoustic source deep decoupling model to extract global acoustic semantic features; and using the mask estimation branch in the acoustic source deep decoupling model to calculate the human voice survival probability weight value based on the global acoustic semantic features to generate a corresponding time-frequency domain acoustic mask weight sequence.

5. The method for automated track splitting of original audio for film dubbing according to claim 4, characterized in that, The generation of a background acoustic environment energy feature parameter sequence based on the time-frequency domain acoustic mask weight sequence and the acoustic analytical initial characterization data table includes: Based on the time-frequency domain acoustic mask weight sequence and the acoustic analytical initial characterization data table, a time-frequency grid energy modulation distribution data table reflecting the time-frequency energy distribution weight is generated; Based on the time-frequency grid energy modulation distribution data table, the human voice feature components are separated from the acoustic analysis initial characterization data table; Based on the time-frequency grid energy modulation distribution data table, the environmental background feature components are separated from the acoustic analytical initial characterization data table; Based on the human voice feature components and the environmental background feature components, a sequence of background sound environment energy feature parameters representing the distribution of background ambient sound and music features is generated.

6. The method for automated track splitting of original audio for film dubbing according to claim 5, characterized in that, The step of generating a sequence of background sound environment energy feature parameters characterizing the distribution of background ambient sound and music features based on the human voice feature components and the environmental background feature components includes: The human voice feature components are reconstructed using a nonlinear signal based on acoustic residuals to generate reconstructed human voice feature components. The reconstructed human voice feature components are subjected to feature compensation for the spectral phase using a phase polarity alignment algorithm to generate decoupled pure human voice stream data components that characterize pure pronunciation features. Based on the decoupled pure human voice stream data components and using the environmental background feature components, energy feature aggregation based on spectral envelope is performed to generate a sequence of background sound environment energy feature parameters that characterize the distribution of background ambient sound and music features.

7. The method for automated track splitting of original audio for film dubbing according to claim 1, characterized in that, Construct a role-identified voice track attribute mapping detail table to record voice attribution relationships, including: The decoupled pure human voice stream data components are input into the pre-configured voiceprint feature clustering processing logic; The decoupled pure human voice stream data components are segmented by a sliding window using the voiceprint feature clustering processing logic to generate audio speech segment distribution blocks; The audio speech segment distribution block is acoustically statistically modeled using a pre-built voiceprint feature extractor in order to extract and generate the voiceprint embedding feature attributes of the corresponding speech subject; Based on the voiceprint embedding feature attributes, a role-identified voice track attribute mapping detail table is constructed to record the voice affiliation relationship.

8. The method for automated track splitting of original audio for film dubbing according to claim 7, characterized in that, The step of constructing a role-identified voice track attribute mapping detail table for recording voice attribution relationships based on the voiceprint embedding feature attributes includes: The voiceprint embedding feature attributes are mapped to their respective feature spaces to generate mapped voiceprint embedding feature vectors, and the geometric distance between each mapped voiceprint embedding feature vector is calculated to generate geometric distance distribution measurements. Based on the geometric distance distribution measurements, the cluster centers corresponding to each audio speech segment distribution block are determined; The cluster centers are introduced into the voiceprint comparison space and matched with the preset role identity feature library after feature denoising to determine the associated movie role identity tags to which each audio speech segment distribution block belongs. Based on the associated film character identity tags, the decoupled pure voice stream data components are reorganized spatially based on the character affiliation attribute to construct a character-identified voice track attribute mapping detail table for recording voice affiliation relationships.

9. The method for automated track splitting of original audio for film dubbing according to claim 1, characterized in that, The automated film dubbing track-splitting audio control command stream data, generated based on the character-identified human voice track-splitting attribute mapping detail table and the background sound environment energy feature parameter sequence, is used to characterize the predicted optimal track-splitting state of the original film audio and can be parsed by the dubbing system. This includes: The character-identified human voice track attribute mapping details table and the background sound environment energy feature parameter sequence are logically encapsulated into a data structure to generate a full-scene track-based acoustic timing sequence that includes timestamp index, character identification index and audio track physical feature distribution. The full-scene track-by-track acoustic timing sequence is subjected to amplitude gain control based on a preset loudness standard to generate a gain-controlled acoustic timing sequence. Based on the acoustic timing sequence after gain control, an automated film dubbing track splitting audio control command stream is generated to characterize the predicted optimal track splitting state of the original film audio and can be parsed by the dubbing system.

10. The method for automated track splitting of original audio for film dubbing according to claim 9, characterized in that, Based on the gain-controlled acoustic timing sequence, an automated film dubbing track-splitting audio control command stream is generated to characterize the predicted optimal track splitting state of the original film's audio and can be parsed by the dubbing system. This includes: The energy balance relationship between different character audio tracks is determined based on the acoustic timing sequence after gain control. The gain-controlled acoustic timing sequence is encapsulated using a transmission protocol based on the movie dubbing standard to generate encapsulated acoustic transmission data packets. Based on the energy balance relationship, the encapsulated acoustic transmission data packet is physically resampled to generate waveform-resampled acoustic data. The waveform-resampled acoustic data is then encoded and converted to generate automated film dubbing track-splitting audio control command stream data that characterizes the predicted optimal track splitting state of the original film audio and can be parsed by the dubbing system.

11. An automated track-splitting device for original film audio in film dubbing, characterized in that, include: The first program unit is used to retrieve the original composite audio track data carried by the original film and perform temporal segmentation on the original composite audio track data based on the overlap sampling rule to generate a temporal audio frame sequence; and to construct an acoustic analysis initial characterization data table for low-level audio feature analysis based on the temporal audio frame sequence. The second program unit is used to inject the acoustic analytical initial characterization data table as input excitation into the pre-trained acoustic source deep decoupling model for processing, so as to generate the corresponding time-frequency domain acoustic mask weight sequence. Based on the time-frequency domain acoustic mask weight sequence and the acoustic analysis initial characterization data table, a background sound environment energy feature parameter sequence representing the distribution of background ambient sound and music features is generated. The third program unit is used to construct a role-identified human track attribute mapping detail table for recording the human voice attribution relationship; The fourth program unit is used to generate automated film dubbing track splitting audio control command stream data based on the character-identified human voice track splitting attribute mapping details table and the background sound environment energy characteristic parameter sequence, which is used to characterize the predicted optimal track splitting state of the original film audio and can be parsed by the dubbing system.