Audio and video real-time intelligent shunting translation method and system combined with buffer strategy

By dynamically adjusting the size of the audio and video buffers, combined with semantic analysis and network quality priority processing, the problems of data backlog and latency in existing real-time audio and video translation technologies are solved, achieving more efficient and accurate translation results.

CN121615662BActive Publication Date: 2026-05-12JIANGSU ZHIMENG INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU ZHIMENG INTELLIGENT TECH CO LTD
Filing Date
2026-01-29
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

现有音视频实时翻译技术中,缺乏动态缓冲机制和分流机制,导致网络条件差时数据积压或丢失,翻译延迟增加,音视频数据未按语义重要性优先级处理,影响翻译的实时性和准确性。

Method used

A real-time intelligent audio and video splitting translation method combining buffering strategies is adopted. Audio and video data are split through a preset data splitting model, the size of audio and video buffers is dynamically adjusted, translation is performed according to network transmission quality and semantic analysis priority, and a hierarchical buffer is constructed to prioritize the processing of key information.

Benefits of technology

It improves the accuracy and real-time performance of translation, reduces data waiting time, enhances translation efficiency, and ensures smooth transmission and contextual coherence of audio and video data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121615662B_ABST
    Figure CN121615662B_ABST
Patent Text Reader

Abstract

The application relates to the field of audio and video real-time translation technology and discloses an audio and video real-time intelligent shunting translation method and system combined with a buffering strategy, which comprises the following steps: analyzing real-time acquired audio and video data, shunting the audio and video data through a preset data shunting model to obtain audio data and video data; configuring a dynamic buffering mechanism, performing semantic analysis and data division on the audio data of an audio buffering area to obtain an audio data block set; analyzing the semantic importance of the data to calculate a corresponding priority, performing key frame division on the video data to obtain a video key frame priority and an audio block priority; respectively translating the video key frame and the audio data block through a preset translation model to obtain a first translation result and a second translation result; and dynamically adjusting the display time of the translation result to obtain a real-time translation result; through dynamic adjustment of the buffering area size and the hierarchical structure, the application can reduce the waiting time and improve the accuracy and real-time performance of the translation result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of real-time audio and video translation technology, and more specifically to a real-time intelligent audio and video splitting translation method and system that incorporates buffering strategies. Background Technology

[0002] Currently, the demand for real-time audio and video translation technology is growing in fields such as international conferences, online education, live streaming, and telemedicine. Existing technologies typically employ streaming processing, parsing and translating the audio and video streams as a whole. However, this method processes audio and video data uniformly, lacks a splitting mechanism, suffers from low processing efficiency, and uses a fixed-size buffer, making it unable to adapt to dynamic changes in network transmission quality. This affects the real-time nature and accuracy of the translation, and the asynchronous translation results impact the accuracy and fluency of real-time communication.

[0003] Existing technologies suffer from the following problems: they use fixed-size buffers, which cannot be dynamically adjusted according to network transmission quality, leading to data backlog or loss under poor network conditions and increased translation latency; they treat audio and video data as a whole without considering the semantic relationships between them, resulting in asynchronous translation results and poor coherence; and they lack priority assessment of semantic importance, causing key information to be delayed, which affects the real-time performance and accuracy of the translation. To address at least one of the above problems, this application proposes a real-time intelligent audio and video splitting translation method and system that incorporates a buffering strategy. Summary of the Invention

[0004] To address the shortcomings of existing technologies, the purpose of this application is to provide a real-time intelligent audio and video splitting and translation method and system that incorporates buffering strategies, effectively solving the problems in the background technology. The specific technical solution of this application is as follows:

[0005] A real-time intelligent audio and video splitting translation method incorporating buffering strategies includes:

[0006] The system parses the real-time acquired audio and video data and splits the data using a preset data splitting model to obtain audio and video data.

[0007] Configure a dynamic buffering mechanism to dynamically adjust the size of the audio buffer based on real-time monitored network transmission quality parameters, perform semantic analysis and data partitioning on the audio data in the audio buffer, and obtain a set of audio data blocks;

[0008] By combining video data and audio data block sets, the semantic importance of the data is analyzed to calculate the corresponding priority. The video data is divided into keyframes to obtain the video keyframe priority and audio block priority.

[0009] According to the video keyframe priority and audio block priority, the video keyframe and audio data block are translated by a preset translation model to obtain the first translation result and the second translation result.

[0010] By analyzing the mapping relationship between the first and second translation results in chronological order, and dynamically adjusting the display time of the translation results, real-time translation results are obtained.

[0011] Specifically, the process of parsing the real-time acquired audio and video data and splitting the audio and video data using a preset data splitting model to obtain audio data and video data includes:

[0012] The real-time acquired audio and video data is parsed to determine the data track encoding;

[0013] Based on the data track encoding, the audio and video data are split using a preset data splitting model to obtain audio data and video data.

[0014] Specifically, the dynamic buffering mechanism includes:

[0015] Based on real-time monitored network transmission quality parameters, analyze network status to determine dynamic buffer configuration parameters, construct hierarchical buffers, and dynamically adjust the size of audio buffers.

[0016] Semantic analysis and data partitioning are performed on the audio data in the audio buffer to obtain the first audio unit set. The silence interval duration between adjacent audio units is analyzed, and the first audio unit set is merged and optimized to obtain the audio data block set.

[0017] Specifically, the step of analyzing network status based on real-time monitored network transmission quality parameters to determine dynamic buffer configuration parameters, constructing a hierarchical buffer, and dynamically adjusting the size of the audio buffer includes:

[0018] Based on real-time monitored network transmission quality parameters, the network status is analyzed using a preset network quality analysis model to determine dynamic buffer configuration parameters.

[0019] According to the dynamic buffer configuration parameters, the audio buffer is divided into a first buffer layer, a second buffer layer and a third buffer layer. The first buffer layer stores the audio data to be processed. The second buffer layer preloads audio data according to real-time audio semantics. The third buffer layer expands when the network state changes, resulting in a layered buffer.

[0020] By combining the dynamic buffer configuration parameters and real-time audio semantic analysis results, the boundaries of the buffer layer are dynamically adjusted, and the size of the audio buffer is dynamically adjusted as well.

[0021] Specifically, by combining the dynamic buffer configuration parameters and real-time audio semantic analysis results, the boundaries of the buffer layer are dynamically adjusted, and the size of the audio buffer is dynamically adjusted, including:

[0022] By combining dynamic buffer configuration parameters and real-time audio semantic analysis results, the buffer duration corresponding to audio data in the first buffer layer is analyzed, and the boundary of the first buffer layer is optimized to obtain the first optimized buffer layer.

[0023] Using a pre-defined semantic prediction model, semantic prediction is performed on the real-time audio semantic analysis results, the buffer duration corresponding to the preloaded audio data is calculated, and the boundary of the second buffer layer is optimized to obtain the second optimized buffer layer.

[0024] Based on the dynamic buffer configuration parameters and the real-time network status, the corresponding expansion requirements under the real-time network status are calculated, and the boundary of the third buffer layer is optimized to obtain the third optimized buffer layer.

[0025] The size of the audio buffer is dynamically adjusted by combining the first optimized buffer layer, the second optimized buffer layer, and the third optimized buffer layer.

[0026] Specifically, the audio data in the audio buffer undergoes semantic analysis and data partitioning to obtain a first audio unit set. The silence interval duration between adjacent audio units is analyzed, and the first audio unit set is merged and optimized to obtain an audio data block set, including:

[0027] Acoustic features are extracted from the audio data of the first buffer layer in the audio buffer, and semantic analysis and data partitioning are performed through a preset speech recognition model to obtain the first audio unit set.

[0028] Based on the first audio unit set, the silence interval duration between adjacent audio units is analyzed, and adjacent audio units with silence interval duration less than a preset interval threshold are merged to obtain the second audio unit set.

[0029] The semantic integrity of the audio units in the second audio unit set is analyzed, and the semantic boundaries are expanded and optimized to obtain the audio data block set.

[0030] Specifically, the analysis of the semantic integrity of audio units in the second audio unit set, and the expansion and optimization of semantic boundaries, yields an audio data block set, including:

[0031] Using a pre-defined semantic analysis model, the semantic integrity of the audio units in the second audio unit set is analyzed, and the corresponding semantic integrity score is calculated.

[0032] For audio units whose semantic integrity score is less than a preset semantic threshold, extend them forward and backward by a preset semantic length each until the semantic integrity score is greater than or equal to the preset semantic threshold, thus obtaining the first extended unit;

[0033] For the first extension unit whose semantic length is greater than a preset length threshold, select semantic segmentation points that satisfy the semantic integrity score greater than or equal to the preset semantic threshold, and perform semantic segmentation on the first extension unit to obtain the second extension unit;

[0034] The audio units with semantic integrity scores greater than or equal to a preset semantic threshold in the second extended unit and the second audio unit set are combined to obtain an audio data block set.

[0035] Specifically, the step of combining video and audio data block sets, analyzing the semantic importance of the data to calculate the corresponding priority, and dividing the video data into keyframes to obtain video keyframe priorities and audio block priorities includes:

[0036] By combining video and audio data blocks, the semantic importance of the data is analyzed. A pre-defined priority analysis model is used to analyze semantic density, emotional intensity, and contextual relevance, and the corresponding priorities are obtained by fusion.

[0037] Based on priority, the video data is divided into keyframes to obtain the video keyframe priority and audio block priority.

[0038] Specifically, the step of dividing the video data into keyframes according to priority to obtain video keyframe priorities and audio block priorities includes:

[0039] The video data is divided into keyframes according to priority, resulting in a keyframe set.

[0040] Semantic consistency analysis is performed on each keyframe and audio block to construct a semantic correlation matrix between keyframes and audio blocks. The priority of keyframes and audio blocks is analyzed through a preset dynamic priority analysis model to obtain the video keyframe priority and audio block priority.

[0041] A real-time intelligent audio-video splitting and translation system incorporating a buffering strategy is used to implement the aforementioned real-time intelligent audio-video splitting and translation method incorporating a buffering strategy, including:

[0042] The data splitting module parses the real-time acquired audio and video data and splits the audio and video data through a preset data splitting model to obtain audio data and video data.

[0043] The audio data block partitioning module is configured with a dynamic buffering mechanism. It dynamically adjusts the size of the audio buffer based on real-time monitored network transmission quality parameters, performs semantic analysis and data partitioning on the audio data in the audio buffer, and obtains a set of audio data blocks.

[0044] The priority analysis module combines video data and audio data block sets to analyze the semantic importance of the data, calculate the corresponding priority, divide the video data into keyframes, and obtain the video keyframe priority and audio block priority.

[0045] The split translation module translates the video keyframes and audio data blocks according to the video keyframe priority and audio block priority, respectively, using a preset translation model to obtain a first translation result and a second translation result.

[0046] The real-time translation module analyzes the mapping relationship between the first and second translation results in chronological order and dynamically adjusts the display time of the translation results to obtain real-time translation results.

[0047] The beneficial effects of this application are as follows: Data is split into audio and video streams. By monitoring network transmission quality parameters in real time, the size of the audio buffer is dynamically adjusted, and a hierarchical buffer is constructed. Buffer boundaries are optimized based on network status and semantic prediction, enabling smooth data transmission. Combining video and audio data block sets, semantic importance is analyzed. The priority of video keyframes and audio blocks is calculated using a priority analysis model, prioritizing key information. Semantic analysis and silence interval detection are performed on the audio buffer data, short-interval audio units are merged, and semantic integrity scores are analyzed for boundary expansion and segmentation, resulting in an optimized audio data block set, thus improving translation accuracy. Dynamic buffering and priority processing, dynamically adjusting the buffer size and hierarchical structure, reduce data waiting time and improve translation efficiency. Analyzing semantic integrity for audio data block segmentation and priority calculation allows for accurate translation of corresponding key information, reducing semantic translation errors and improving the accuracy and real-time performance of translation results. Attached Figure Description

[0048] Figure 1 This is a flowchart illustrating the real-time intelligent audio and video splitting translation method incorporating a buffering strategy as described in this application embodiment.

[0049] Figure 2 This is a schematic diagram of the audio buffer in an embodiment of this application;

[0050] Figure 3 A flowchart illustrating the construction of the audio data block set in the embodiments of this application;

[0051] Figure 4This is a schematic diagram of the structure of the real-time intelligent audio and video splitting translation system with buffering strategy in an embodiment of this application. Detailed Implementation

[0052] The present application will be further described in detail below with reference to the accompanying drawings and embodiments.

[0053] In the embodiments of this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design that is described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0054] Hereinafter, the terms "first," "second," and other generic terms are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of the embodiments of this application, unless otherwise stated, "multiple" means two or more.

[0055] refer to Figure 1 As shown, the specific implementation of the real-time intelligent audio and video splitting translation method combining buffering strategy in this application includes:

[0056] S101. Parse the real-time acquired audio and video data, and split the audio and video data through a preset data splitting model to obtain audio data and video data.

[0057] S102. Configure a dynamic buffering mechanism to dynamically adjust the size of the audio buffer based on real-time monitored network transmission quality parameters, perform semantic analysis and data partitioning on the audio data in the audio buffer, and obtain a set of audio data blocks.

[0058] S103. Combining the video data and audio data block sets, analyze the semantic importance of the data, calculate the corresponding priority, divide the video data into key frames, and obtain the video key frame priority and audio block priority.

[0059] S104. According to the video keyframe priority and audio block priority, the video keyframe and audio data block are translated by a preset translation model to obtain the first translation result and the second translation result.

[0060] S105. Analyze the mapping relationship between the first and second translation results in chronological order, dynamically adjust the display time of the translation results, and obtain real-time translation results.

[0061] Existing real-time audio and video translation technologies typically employ fixed buffering strategies and simple streaming processing, translating audio and video data directly as a whole or after coarse-grained separation. This approach lacks adaptability to dynamically changing network environments, easily leading to data loss, translation delays, or stuttering when network transmission quality is unstable. Furthermore, the failure to intelligently prioritize data based on semantic importance means critical information cannot be processed in a timely manner, and the lack of an effective synchronization mechanism results in disjointed translations in terms of timing and context, impacting the accuracy and fluency of real-time communication.

[0062] In this embodiment, the real-time acquired audio and video data is parsed to obtain data track encoding. The audio and video data are then split using a preset data splitting model to obtain audio data and video data. By effectively separating the audio and video data at the data input point, differentiated processing can be performed based on the respective data characteristics of audio and video. This provides a data foundation for independent dynamic buffering and priority analysis calculations, improves the parallelism of data processing, reduces processing latency, and enhances translation efficiency and real-time performance.

[0063] Specifically, a dynamic buffering mechanism is configured to dynamically adjust the size of the audio buffer based on real-time monitored network transmission quality parameters. These parameters include, but are not limited to, packet loss rate, latency jitter, and available bandwidth calculated by the receiver. Semantic analysis and data partitioning are performed on the audio data in the audio buffer to obtain a set of audio data blocks. Through the dynamic buffering mechanism, combined with the network environment and content semantics, dynamic and elastic buffering settings can effectively smooth out data flow instability caused by network jitter, prevent data underloading or overloading, and partition data blocks based on semantic integrity. This ensures that each audio unit is a semantically consistent whole, avoiding mistranslation and ambiguity caused by incomplete input information, improving the accuracy and readability of translation results, and enhancing the robustness of the processing flow in the face of network fluctuations.

[0064] Furthermore, by combining video and audio data blocks, the semantic importance of the data is analyzed, and semantic density, emotional intensity, and contextual relevance are extracted and fused to obtain the corresponding priorities. The video data is then divided into keyframes to obtain video keyframe priorities and audio block priorities. By analyzing the semantic importance of audio and video data, key content with high information content, strong emotions, or that connects different parts of the text can be prioritized for processing. This ensures that important information is processed and translated first when system computing resources are limited, optimizing the user's perceptual experience. By establishing semantic connections between audio and video, related visual and auditory information can be processed collaboratively when displaying translation results, improving the synchronicity and contextual coherence of the translation results.

[0065] Specifically, video keyframes and audio blocks are sorted from highest to lowest priority, with higher-priority data being translated first. For video keyframes, the translation model includes, but is not limited to, a neural machine translation model based on the Transformer architecture. The neural machine translation model is trained using a large amount of historical video keyframe data to obtain a pre-trained neural machine translation model. The video keyframes are then input into the pre-trained neural machine translation model, which translates the video keyframes to obtain the first translation result. For audio blocks, the translation model includes, but is not limited to, an end-to-end speech translation model. The speech translation model is trained using a large amount of historical audio block data to obtain a pre-trained speech translation model. The audio blocks are then input into the pre-trained speech translation model to obtain the second translation result. The two models can be computed in parallel.

[0066] It should be noted that prioritizing data translation allows for the translation of high-value information segments based on computing resources. This ensures the efficient transmission of critical information in the face of sudden traffic surges or complex scenarios. Separating audio and video translation paths allows for the selection of appropriate translation models for image text and voice dialogue, thereby improving the accuracy and efficiency of the translation process.

[0067] Specifically, the mapping relationship between the first and second translation results is analyzed in chronological order. The first and second translation results are aligned according to their corresponding original audio and video times. The display time of the translation results is dynamically adjusted based on the chronological order. For example, for video subtitles of lower importance, the display can be slightly delayed to ensure that the currently playing important audio translation is presented to the user completely and without interference. When audio silence is detected and there is key video text that needs emphasis, the video translation result is displayed first, resulting in a real-time translation. Through dynamic time adjustment, the translated text information can be perceptually synchronized with the viewed image and heard sound, avoiding translation misalignment and improving the real-time performance and accuracy of the translation results.

[0068] This application performs data splitting on audio and video data. By monitoring network transmission quality parameters in real time, it dynamically adjusts the size of the audio buffer and constructs a hierarchical buffer. The buffer boundaries are optimized based on network status and semantic prediction to ensure smooth data transmission. Combining video and audio data block sets, semantic importance is analyzed. A priority analysis model is used to calculate the priority of video keyframes and audio blocks, prioritizing the processing of key information. Semantic analysis and silence interval detection are performed on the audio buffer data, merging short-interval audio units. Semantic integrity scores are analyzed for boundary expansion and segmentation, resulting in an optimized audio data block set, thus improving translation accuracy. Through dynamic buffering and priority processing, dynamically adjusting the buffer size and hierarchical structure reduces data waiting time and improves translation efficiency. Analyzing semantic integrity for audio data block segmentation and priority calculation accurately translates corresponding key information, reduces semantic translation errors, and improves the accuracy and real-time performance of the translation results.

[0069] Furthermore, the real-time acquired audio and video data is parsed, and the audio and video data are split using a preset data splitting model to obtain audio data and video data, including:

[0070] S201. Analyze the real-time acquired audio and video data to determine the data track encoding;

[0071] S202. Based on the data track encoding, the audio and video data are split using a preset data splitting model to obtain audio data and video data.

[0072] In this embodiment, for real-time acquired audio and video data, the audio and video data file header or initial information packet is located and read. The file header or initial information packet includes metadata defining the data structure. The metadata is parsed to determine the data track encoding, identifying the number of tracks in the data and the encoding standard of each track. For example, video tracks use H.264 or H.265 encoding, and audio tracks use AAC or MP3 encoding. By determining the data track encoding, accurate data support is provided for data splitting, improving the accuracy of data splitting.

[0073] Specifically, based on the data track encoding, audio and video data are split using a pre-defined data splitting model. This model includes, but is not limited to, a random forest model pre-trained using a large amount of historical track encoding data. The model analyzes the data track encoding, matches the data, and divides it into audio and video data. By splitting the audio and video data, they can be analyzed and processed independently. Simultaneous processing of audio and video data improves the efficiency and accuracy of data processing.

[0074] Furthermore, dynamic buffering mechanisms include:

[0075] S301. Based on the real-time monitored network transmission quality parameters, analyze the network status to determine the dynamic buffer configuration parameters, construct a hierarchical buffer, and dynamically adjust the size of the audio buffer.

[0076] S302. Perform semantic analysis and data partitioning on the audio data in the audio buffer to obtain the first audio unit set. Analyze the silence interval duration between adjacent audio units and merge and optimize the first audio unit set to obtain the audio data block set.

[0077] In this embodiment, based on real-time monitored network transmission quality parameters, including but not limited to network latency, jitter, packet loss rate, and available bandwidth, the network status is analyzed to determine dynamic buffer configuration parameters, and a hierarchical buffer is constructed. The boundaries of each buffer layer are dynamically adjusted according to the dynamic buffer configuration parameters, and the size of the audio buffer is dynamically adjusted as well. By constructing a hierarchical buffer, unstable network data streams can be transformed into relatively stable data streams, enhancing the system's ability to resist network jitter, avoiding data loss or processing delays caused by temporary network degradation, improving the system's robustness in real-world complex network environments, and adapting to different data buffering needs through hierarchical design, effectively balancing resource efficiency and performance assurance.

[0078] Specifically, semantic analysis and data segmentation are performed on the audio data in the audio buffer to identify the start and end points of speech segments, resulting in the first set of audio units. The silence interval duration between adjacent audio units is analyzed. If the silence interval duration is less than a preset interval threshold, the two units are determined to be semantically closely related. The interval threshold can be set according to the language and scenario requirements to avoid misjudging a natural pause in a sentence as the end of the sentence. The two units are then merged and optimized to obtain a set of audio data blocks.

[0079] It should be noted that semantically guided audio segmentation can avoid the semantic fragmentation caused by relying solely on silence detection, transforming audio from a signal stream into a sequence of units with a preliminary grammatical structure. By merging excessively short silence intervals, it can effectively restore the natural pauses in the speaker's sentences, avoiding the mechanical cutting of a complete semantic unit into multiple small fragments that cannot stand alone. This provides higher-quality, more contextually complete input data for the translation process, reducing the risk of incoherent translation results, loss of context, or mistranslation due to fragmented input information, and improving the accuracy and fluency of the translation results.

[0080] Furthermore, based on real-time monitored network transmission quality parameters, the network status is analyzed to determine dynamic buffer configuration parameters, a hierarchical buffer is constructed, and the size of the audio buffer is dynamically adjusted, including:

[0081] S401. Based on the real-time monitored network transmission quality parameters, analyze the network status and determine the dynamic buffer configuration parameters through a preset network quality analysis model.

[0082] S402. According to the dynamic buffer configuration parameters, the audio buffer is divided into a first buffer layer, a second buffer layer and a third buffer layer. The first buffer layer stores the audio data to be processed. The second buffer layer preloads audio data according to real-time audio semantics. The third buffer layer expands when the network state changes, thus obtaining a layered buffer.

[0083] S403. Combining the dynamic buffer configuration parameters and real-time audio semantic analysis results, dynamically adjust the boundary of the buffer layer and dynamically adjust the size of the audio buffer.

[0084] In this embodiment, based on real-time monitored network transmission quality parameters, including but not limited to packet round-trip delay, latency jitter, and packet loss rate, the network status is analyzed through a preset network quality analysis model to determine dynamic buffer configuration parameters. The network quality analysis model includes, but is not limited to, a random forest model pre-trained using a large amount of historical network transmission quality parameter data. The model learns the nonlinear relationship between network parameter combinations and the optimal buffer strategy from historical data and outputs dynamic buffer configuration parameters, including but not limited to the baseline size of the buffer and the initial capacity allocation ratio of each layer of buffers.

[0085] It should be noted that, unlike the traditional passive and slow response to network fluctuations, this embodiment can proactively and predictively adjust resource allocation strategies based on accurate diagnosis of network status by dynamically configuring buffers through a network quality analysis model, thereby improving the system's ability to judge the network environment and the accuracy of buffer configuration.

[0086] like Figure 2 As shown, according to the dynamic buffer configuration parameters, the audio buffer is divided into a first buffer layer, a second buffer layer, and a third buffer layer. The first buffer layer includes a small but highest priority queue used to store the audio data to be processed, ensuring that the speech recognition engine always has data to process. The second buffer layer includes an area with intelligent pre-reading function, which preloads audio data based on real-time audio semantics. When the system analyzes that the current audio is in an incomplete sentence or detects specific keywords, it will trigger the pre-loading mechanism to store the audio data corresponding to the subsequent semantics in advance into this layer to counteract fixed network transmission delays and ensure semantic coherence. The third buffer layer expands when the network state changes, which is used to enhance the system's resilience.

[0087] It is important to emphasize that by setting up a layered buffer structure, a fully functional and collaborative buffer structure can be constructed. The first buffer layer is used to ensure the real-time performance of processing, the second buffer layer is used to improve the continuity of semantics, and the third buffer layer is used to maintain the robustness of the system. The layered architecture can avoid the contradictions of a single buffer when dealing with multi-objective optimization, and improve the efficiency and real-time performance of the system translation process.

[0088] Specifically, by combining dynamic buffer configuration parameters and real-time audio semantic analysis results, the boundaries of the buffer layer are dynamically adjusted. This dynamic adjustment process includes, but is not limited to, when semantic analysis indicates an upcoming lengthy speech, the system will appropriately expand the boundaries of the second buffer layer within the limits allowed by the configuration parameters to accommodate more preloaded data, thus dynamically adjusting the size of the audio buffer. By optimizing the audio buffer boundaries based on actual content load and network trends, fixed resource waste or temporary shortages can be avoided, improving resource utilization efficiency. This allows the system to adapt to dynamic changes in network conditions and fluctuations in audio semantics, reducing system resource overhead and improving system translation performance and the accuracy of translation results.

[0089] Furthermore, combining the dynamic buffer configuration parameters and real-time audio semantic analysis results, the boundaries of the buffer layer are dynamically adjusted, and the size of the audio buffer is dynamically adjusted, including:

[0090] S501. Combining dynamic buffer configuration parameters and real-time audio semantic analysis results, analyze the buffer duration corresponding to audio data in the first buffer layer, optimize the boundary of the first buffer layer, and obtain the first optimized buffer layer.

[0091] S502. Using a preset semantic prediction model, perform semantic prediction on the real-time audio semantic analysis results, calculate the buffer duration corresponding to the preloaded audio data, optimize the boundary of the second buffer layer, and obtain the second optimized buffer layer.

[0092] S503. Based on the dynamic buffer configuration parameters and the real-time network status, calculate the corresponding expansion requirements under the real-time network status, optimize the boundary of the third buffer layer, and obtain the third optimized buffer layer.

[0093] S504: Combines the first optimized buffer layer, the second optimized buffer layer and the third optimized buffer layer to dynamically adjust the size of the audio buffer.

[0094] In this embodiment, the system continuously monitors the total amount of audio data currently stored and to be processed in the first buffer layer, calculates the buffering time required to clear the current buffer content by combining parameters such as audio sampling rate, compares the buffering time with the target buffering time range set for the first buffer layer in the dynamic buffer configuration parameters, and analyzes the data throughput required to translate the current semantics by combining real-time audio semantic analysis results. For example, when the current audio is identified as a segment with an extremely fast speaking speed, it reflects an increase in the data throughput demand per unit time. The data throughput and buffering time are weighted and fused to calculate the maximum data volume of the first buffer layer, and the boundary of the first buffer layer is optimized to obtain the first optimized buffer layer.

[0095] It should be noted that by dynamically optimizing the boundary of the first buffer layer, it can be ensured that the speech recognition engine can obtain a data stream that is neither too small, causing frequent waiting, nor too large, introducing excessively high latency. This ensures real-time translation of the data within the first buffer layer, avoids quality fluctuations caused by unstable data supply, and improves real-time translation quality.

[0096] Specifically, a pre-defined semantic prediction model is used to predict the semantics of real-time audio semantic analysis results. This model includes, but is not limited to, a recurrent neural network model pre-trained using a large amount of historical audio data. The model predicts one or more semantic units that will appear next in the current context. Based on the predicted text length and corresponding pronunciation duration, the buffering time required for preloading audio data is calculated. The boundary of the second buffer layer is optimized according to the required buffering time, and the upper limit of the capacity of the second buffer layer is adjusted to obtain the second optimized buffer layer. By predicting the buffering time through semantic analysis, inherent network transmission delays can be avoided, the necessary data can be preloaded, and the waiting time in the speech data translation process can be reduced. This effectively ensures the fluency of the dialogue and the coherence of the semantics, improving the user experience in real-time translation interaction.

[0097] Specifically, the system continuously monitors real-time network status parameters, including but not limited to the upward trend of packet loss rate and the aggravation of latency jitter. Combining dynamic buffer configuration parameters and real-time network status parameters, it calculates the required expansion demand under real-time network conditions. It performs weighted fusion calculation on the degree of deterioration of network indicators to calculate the impact capacity of network indicator deterioration on the calculation process, obtains the expansion demand, optimizes the boundary of the third buffer layer, and directly expands the maximum capacity limit of the third buffer layer. This provides sufficient space to accumulate unprocessed data packets during sudden network congestion, preventing loss, and thus obtaining the optimized third buffer layer.

[0098] It should be noted that by dynamically expanding the boundary of the third buffer layer, data generated during peak network congestion periods can be absorbed, reducing data loss caused by network problems. This enhances the system's environmental adaptability under harsh network conditions and ensures that core voice data can still be completely preserved and processed even when the network is unstable, thereby improving the system's reliability and the accuracy of the translation results.

[0099] Specifically, by combining the first, second, and third optimized buffer layers, the size of the audio buffer is dynamically adjusted based on the audio buffer memory budget in the dynamic buffer configuration parameters. When a layer needs significant expansion due to network degradation, the pre-allocated quota of another layer can be moderately compressed without severely impacting the core functions of other layers, ensuring that total memory does not overflow. When all layers have redundancy, the entire layer is shrunk to release resources. By dynamically adjusting the audio buffer, limited memory resources can be dynamically allocated according to real-time changes in processing needs, semantic context, and network conditions. Global coordination can avoid resource conflicts or overall inefficiency caused by isolated optimization of each buffer layer, thereby improving overall resource utilization efficiency.

[0100] Furthermore, semantic analysis and data partitioning are performed on the audio data in the audio buffer to obtain the first audio unit set. The silence interval duration between adjacent audio units is analyzed, and the first audio unit set is merged and optimized to obtain the audio data block set, including:

[0101] S601. The acoustic features of the audio data of the first buffer layer in the audio buffer are extracted, and the semantic analysis and data partitioning are performed through the preset speech recognition model to obtain the first audio unit set.

[0102] S602. Based on the first audio unit set, analyze the silence interval duration between adjacent audio units, and merge adjacent audio units whose silence interval duration is less than a preset interval threshold to obtain a second audio unit set.

[0103] S603. Analyze the semantic integrity of the audio units in the second audio unit set, expand and optimize the semantic boundaries to obtain the audio data block set.

[0104] In this embodiment, acoustic features are extracted from the audio data of the first buffer layer in the audio buffer, including but not limited to Mel frequency cepstral coefficients and filter bank features. Semantic analysis and data segmentation are performed using a preset speech recognition model. The speech recognition model includes but is not limited to an RNN model pre-trained using a large number of acoustic features. The model automatically identifies the start and end boundaries of speech segments and segments the audio data to obtain a first set of audio units. Through semantic recognition segmentation, accurate semantic base units are provided for the translation process, avoiding the semantic fragmentation problem caused by fixed-length slicing and improving the accuracy of the translation results.

[0105] Specifically, based on the first set of audio units, the difference between the end timestamp of each audio unit and the start timestamp of the next audio unit is calculated to obtain the silence interval duration between adjacent audio units. An interval threshold is set through statistical analysis of a large number of natural dialogues. The silence interval duration is compared with the preset threshold to distinguish between natural short pauses in the speech and semantic boundaries. When the silence interval duration is less than the preset threshold, it reflects that the two audio units are semantically closely related, and these two adjacent audio units are merged. After merging all audio units that meet the conditions, a second set of audio units is obtained. By merging audio units separated by brief natural pauses, the complete semantics of the speaker's speech flow can be restored, avoiding the mechanical cutting of a coherent expression into multiple fragmented parts. This improves the semantic coherence and contextual integrity of the translation process, reduces unnatural translation results, lost context, or logical confusion caused by fragmented input information, and improves the accuracy and coherence of the translation results.

[0106] Specifically, the semantic integrity of the audio units in the second audio unit set is analyzed, and the corresponding semantic integrity score is calculated. For units with semantic integrity scores lower than a preset semantic threshold, the semantic boundaries are expanded and optimized to obtain an audio data block set. By analyzing semantic integrity and optimizing boundaries, it is ensured that each data block is a semantically complete whole with sufficient contextual information and capable of independent meaning, thus improving the quality of the input data for the translation model. Translation based on complete context can improve the accuracy, fluency, and naturalness of the final translation result.

[0107] like Figure 3 As shown, the semantic integrity of the audio units in the second audio unit set is analyzed, and the semantic boundaries are expanded and optimized to obtain the audio data block set, including:

[0108] S701. Analyze the semantic integrity of the audio units in the second audio unit set using a preset semantic analysis model, and calculate the corresponding semantic integrity score.

[0109] S702. For audio units whose semantic integrity score is less than a preset semantic threshold, extend them forward and backward by a preset semantic length until the semantic integrity score is greater than or equal to the preset semantic threshold, thus obtaining the first extended unit.

[0110] S703. For the first extension unit whose semantic length is greater than the preset length threshold, select the semantic segmentation points that satisfy the semantic integrity score greater than or equal to the preset semantic threshold, and perform semantic segmentation on the first extension unit to obtain the second extension unit.

[0111] S704. Combine the audio units in the second extended unit and the second audio unit set whose semantic integrity scores are greater than or equal to a preset semantic threshold to obtain an audio data block set.

[0112] In this embodiment, the semantic integrity of audio units in the second audio unit set is analyzed by a preset semantic analysis model. The semantic analysis model includes, but is not limited to, a neural network model pre-trained using a large amount of historical audio unit data. The model analyzes the semantic integrity of the second audio unit and calculates the corresponding semantic integrity score. By calculating the semantic integrity score, accurate data support is provided for boundary optimization.

[0113] Specifically, all audio units are traversed, and units with semantic integrity scores less than a preset semantic threshold are selected. The semantic threshold can be set according to the translation accuracy requirements. For each unit, a preset semantic length is extended backward and forward respectively. Backward means towards historical data, and forward means towards future data. The preset semantic length can be set according to the translation accuracy requirements. The semantic integrity score of the extended audio unit is recalculated until the semantic integrity score of the new unit is greater than or equal to the preset semantic threshold. The iteration stops, and the extended audio unit is taken as the first extended unit.

[0114] It should be noted that boundary expansion optimization by analyzing semantic integrity can effectively repair the semantic fragmentation problem caused by inaccurate initial segmentation of speech recognition or improper pauses by the speaker. This ensures that each unit used for translation contains the minimum complete context necessary to generate an accurate translation, avoiding problems such as mistranslation, ambiguity, or unnatural output caused by incomplete input information, and improving the understandability and accuracy of the translation results.

[0115] Specifically, for the first extended unit whose semantic length exceeds a preset length threshold (which can be set according to translation accuracy requirements), within this unit, starting from the unit's starting point, the evaluation window is gradually slidable. The size of the evaluation window can be set according to translation accuracy requirements. The semantic integrity score of the text within the evaluation window is calculated, and positions within the evaluation window whose semantic integrity score is greater than or equal to the preset semantic threshold are selected as semantic segmentation points. The first extended unit is semantically segmented according to these segmentation points, dividing the long unit into multiple shorter, but semantically complete, sub-units, which are then used as the second extended unit. Audio units with semantic integrity scores greater than or equal to the preset semantic threshold from the second extended unit and the second audio unit set are combined to obtain an audio data block set.

[0116] It is important to emphasize that by controlling the length of audio units and segmenting their internal semantics, we can effectively avoid excessively long sentences caused by pursuing semantic completeness. By breaking down overly long extended units containing multiple semantics into a series of semantically complete and appropriately long sub-units, we can ensure translation accuracy while improving translation efficiency and real-time performance.

[0117] Furthermore, by combining the video and audio data block sets, the semantic importance of the data is analyzed to calculate the corresponding priorities. Keyframes are then divided into video data to obtain video keyframe priorities and audio block priorities, including:

[0118] S801. Combining video and audio data blocks, analyze the semantic importance of the data, analyze semantic density, emotional intensity and contextual relevance through a preset priority analysis model, and fuse them to obtain the corresponding priority.

[0119] S802. According to priority, the video data is divided into keyframes to obtain the video keyframe priority and audio block priority.

[0120] In this embodiment, the semantic importance of video and audio data blocks is analyzed by combining them. A preset priority analysis model is used to analyze semantic density, emotional intensity, and contextual relevance. For audio data blocks, the model analyzes the translated text to calculate semantic density and calculates emotional intensity by analyzing the acoustic features of the audio and the text content. For video data, the model calculates the update rate of visual information by detecting scene changes, assesses emotional intensity by recognizing facial expressions, and evaluates the contextual relevance by analyzing the position of the data block in the time series. The semantic density, emotional intensity, and contextual relevance are then fused to obtain the corresponding priority. By calculating the priority, it is ensured that, under the condition of limited computing resources, resources can be allocated to data segments with high information content, strong emotional expression, or that are crucial to understanding the context. This optimizes the overall efficiency of key information transmission and improves the user's perception speed and understanding depth of core content in real-time interaction.

[0121] Specifically, video data is divided into keyframes according to priority, resulting in video keyframe priorities and audio block priorities. By analyzing the semantic relationships between audio and video and calibrating the priorities, it can be ensured that audio and video information that is closely related in time and content can be identified as the same high-priority group during the translation process, and thus receive collaborative processing and synchronous output. This effectively avoids the problem of audio-visual asynchrony, improves translation quality, and enhances the real-time nature of translation results.

[0122] Furthermore, the video data is divided into keyframes according to priority, resulting in video keyframe priorities and audio block priorities, including:

[0123] S901. Divide the video data into keyframes according to priority to obtain a keyframe set;

[0124] S902. Perform semantic consistency analysis on each keyframe and audio block, construct a semantic correlation matrix between keyframes and audio blocks, and analyze the priorities corresponding to keyframes and audio blocks through a preset dynamic priority analysis model to obtain the video keyframe priority and audio block priority.

[0125] In this embodiment, video data is divided into keyframes according to priority to obtain a keyframe set. The video stream is analyzed according to priority, and the camera boundaries are located by scene change detection. The boundaries are filtered and weighted according to priority scores. Through differential sampling, keyframes are extracted according to the corresponding time intervals to obtain the keyframe set. By extracting keyframes, it is ensured that key visual context is not lost due to insufficient sampling during high-value information periods, reducing the amount of data processed and improving translation efficiency.

[0126] Specifically, semantic consistency analysis is performed on each keyframe and audio block. An encoder-decoder structure based on an attention mechanism generates one or more descriptive texts for each keyframe. Each audio block already has its corresponding translated text. Semantic similarity is determined by calculating the cosine similarity between the semantic space vectors of each pair of keyframe descriptive texts and audio block translated texts. The similarity calculation results of all keyframes and all audio blocks are arranged into a two-dimensional table to construct a semantic association matrix between keyframes and audio blocks, reflecting the inherent connections in content between different modal data segments. The semantic association matrix is ​​input into a pre-defined dynamic priority analysis model. This model includes, but is not limited to, a graph optimization model pre-trained using a large amount of historical data. The model treats keyframes and audio blocks as nodes in a graph, with association as edge weights. The final priority of each node is calculated by propagating and aggregating the importance of nodes, resulting in the video keyframe priority and audio block priority. Through semantic association analysis, the priority of audio-video pairs that semantically support each other and jointly convey core information can be identified and increased, improving the temporal consistency and accuracy of the translation results.

[0127] like Figure 4 As shown, a real-time intelligent audio-video splitting and translation system incorporating a buffering strategy is used to implement a real-time intelligent audio-video splitting and translation method incorporating a buffering strategy, including:

[0128] The data splitting module parses the real-time acquired audio and video data and splits the audio and video data through a preset data splitting model to obtain audio data and video data.

[0129] The audio data block partitioning module is configured with a dynamic buffering mechanism. It dynamically adjusts the size of the audio buffer based on real-time monitored network transmission quality parameters, performs semantic analysis and data partitioning on the audio data in the audio buffer, and obtains a set of audio data blocks.

[0130] The priority analysis module combines video data and audio data block sets to analyze the semantic importance of the data, calculate the corresponding priority, divide the video data into keyframes, and obtain the video keyframe priority and audio block priority.

[0131] The split translation module translates the video keyframes and audio data blocks according to the video keyframe priority and audio block priority, respectively, using a preset translation model to obtain a first translation result and a second translation result.

[0132] The real-time translation module analyzes the mapping relationship between the first and second translation results in chronological order and dynamically adjusts the display time of the translation results to obtain real-time translation results.

[0133] In this embodiment, the data splitting module identifies and demultiplexes audio and video tracks, generating independent audio and video data streams. This splitting of audio and video data provides a data foundation for parallel data processing and translation, improving translation efficiency. The audio data block partitioning module dynamically constructs and adjusts hierarchical audio buffers based on real-time network quality parameters, effectively smoothing network jitter. It performs acoustic feature extraction, speech recognition, silence interval analysis, and semantic integrity optimization on the audio within the buffers, obtaining a set of semantically coherent audio data blocks. Combining adaptive network buffering and semantic-driven intelligent data segmentation avoids interference from network instability and semantic fragmentation on audio processing, improving the accuracy of translation results.

[0134] Specifically, the priority analysis module calculates a comprehensive priority by fusing the translated text of video content and audio data blocks, integrating semantic density, emotional intensity, and contextual relevance. It then extracts keyframes from the video data to obtain co-optimized video keyframe and audio block priorities. This ensures that high-priority data is processed first when computing resources are limited, improving translation efficiency. The split-stream translation module routes high-priority video keyframes and audio data blocks to dedicated translation engines. The video path performs optical character recognition and text translation, while the audio path performs end-to-end speech translation or speech recognition followed by text translation. First and second translation results are generated in parallel. Through priority task scheduling and parallel translation, computing resources are intelligently allocated and efficiently utilized, further improving translation efficiency.

[0135] Specifically, the real-time translation module analyzes the mapping relationship and semantic context between the first and second translation results on the timeline, dynamically adjusting their timing and duration in the final display sequence to achieve synchronized and smooth presentation of audio-visual translations. Through its timing management and context-aware display decisions, it avoids the problems of asynchronous and mutually interfering outputs of multimodal translation results, obtaining a real-time translation result stream that is consistent in both time and content, thus improving the accuracy and real-time performance of the translation results.

[0136] The above description is merely a preferred embodiment of this application. The scope of protection of this application is not limited to the above embodiments. All technical solutions falling within the scope of this application's concept are within the scope of protection of this application. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of this application should also be considered within the scope of protection of this application.

Claims

1. A real-time intelligent audio and video splitting translation method combining buffering strategies, characterized in that, include: The system parses the real-time acquired audio and video data and splits the data using a preset data splitting model to obtain audio and video data. Based on real-time monitored network transmission quality parameters, the network status is analyzed using a preset network quality analysis model to determine dynamic buffer configuration parameters. According to the dynamic buffer configuration parameters, the audio buffer is divided into a first buffer layer, a second buffer layer and a third buffer layer. The first buffer layer stores the audio data to be processed. The second buffer layer preloads audio data according to real-time audio semantics. The third buffer layer expands when the network state changes, resulting in a layered buffer. By combining the dynamic buffer configuration parameters and real-time audio semantic analysis results, the boundary of the buffer layer is dynamically adjusted, and the size of the audio buffer is dynamically adjusted. Semantic analysis and data partitioning are performed on the audio data in the audio buffer to obtain the first audio unit set. The silence interval duration between adjacent audio units is analyzed, and the first audio unit set is merged and optimized to obtain the audio data block set. By combining video data and audio data block sets, the semantic importance of the data is analyzed to calculate the corresponding priority. The video data is divided into keyframes to obtain the video keyframe priority and audio block priority. According to the video keyframe priority and audio block priority, the video keyframe and audio data block are translated by a preset translation model to obtain the first translation result and the second translation result. By analyzing the mapping relationship between the first and second translation results in chronological order, and dynamically adjusting the display time of the translation results, real-time translation results are obtained.

2. The real-time intelligent audio and video splitting translation method combining buffering strategy according to claim 1, characterized in that, The process involves parsing the real-time acquired audio and video data, and then splitting the audio and video data using a preset data splitting model to obtain audio data and video data, including: The real-time acquired audio and video data is parsed to determine the data track encoding; Based on the data track encoding, the audio and video data are split using a preset data splitting model to obtain audio data and video data.

3. The real-time intelligent audio and video splitting translation method combining buffering strategy according to claim 1, characterized in that, Combining the dynamic buffer configuration parameters and real-time audio semantic analysis results, the boundaries of the buffer layer are dynamically adjusted, and the size of the audio buffer is dynamically adjusted, including: By combining dynamic buffer configuration parameters and real-time audio semantic analysis results, the buffer duration corresponding to audio data in the first buffer layer is analyzed, and the boundary of the first buffer layer is optimized to obtain the first optimized buffer layer. Using a pre-defined semantic prediction model, semantic prediction is performed on the real-time audio semantic analysis results, the buffer duration corresponding to the preloaded audio data is calculated, and the boundary of the second buffer layer is optimized to obtain the second optimized buffer layer. Based on the dynamic buffer configuration parameters and the real-time network status, the corresponding expansion requirements under the real-time network status are calculated, and the boundary of the third buffer layer is optimized to obtain the third optimized buffer layer. The size of the audio buffer is dynamically adjusted by combining the first optimized buffer layer, the second optimized buffer layer, and the third optimized buffer layer.

4. The real-time intelligent audio and video splitting translation method combining buffering strategy according to claim 3, characterized in that, The process involves semantic analysis and data partitioning of the audio data in the audio buffer to obtain a first audio unit set. The silence interval duration between adjacent audio units is analyzed, and the first audio unit set is merged and optimized to obtain an audio data block set, including: Acoustic features are extracted from the audio data of the first buffer layer in the audio buffer, and semantic analysis and data partitioning are performed through a preset speech recognition model to obtain the first audio unit set. Based on the first audio unit set, the silence interval duration between adjacent audio units is analyzed, and adjacent audio units with silence interval duration less than a preset interval threshold are merged to obtain the second audio unit set. The semantic integrity of the audio units in the second audio unit set is analyzed, and the semantic boundaries are expanded and optimized to obtain the audio data block set.

5. The real-time intelligent audio and video splitting translation method combining buffering strategy according to claim 4, characterized in that, The analysis of the semantic integrity of the audio units in the second audio unit set, and the expansion and optimization of the semantic boundaries, yields a set of audio data blocks, including: Using a pre-defined semantic analysis model, the semantic integrity of the audio units in the second audio unit set is analyzed, and the corresponding semantic integrity score is calculated. For audio units whose semantic integrity score is less than a preset semantic threshold, extend them forward and backward by a preset semantic length each until the semantic integrity score is greater than or equal to the preset semantic threshold, thus obtaining the first extended unit; For the first extension unit whose semantic length is greater than a preset length threshold, select semantic segmentation points that satisfy the semantic integrity score greater than or equal to the preset semantic threshold, and perform semantic segmentation on the first extension unit to obtain the second extension unit; The audio units with semantic integrity scores greater than or equal to a preset semantic threshold in the second extended unit and the second audio unit set are combined to obtain an audio data block set.

6. The real-time intelligent audio and video splitting translation method combining buffering strategy according to claim 1, characterized in that, The process involves combining video and audio data blocks, analyzing the semantic importance of the data to calculate corresponding priorities, dividing the video data into keyframes, and obtaining video keyframe priorities and audio block priorities, including: By combining video and audio data blocks, the semantic importance of the data is analyzed. A pre-defined priority analysis model is used to analyze semantic density, emotional intensity, and contextual relevance, and the corresponding priorities are obtained by fusion. Based on priority, the video data is divided into keyframes to obtain the video keyframe priority and audio block priority.

7. The real-time intelligent audio and video splitting translation method combining buffering strategy according to claim 6, characterized in that, The step of dividing video data into keyframes according to priority to obtain video keyframe priorities and audio block priorities includes: The video data is divided into keyframes according to priority, resulting in a keyframe set. Semantic consistency analysis is performed on each keyframe and audio block to construct a semantic correlation matrix between keyframes and audio blocks. The priority of keyframes and audio blocks is analyzed through a preset dynamic priority analysis model to obtain the video keyframe priority and audio block priority.

8. A real-time intelligent audio and video streaming translation system incorporating buffering strategies, characterized in that: The method for implementing the real-time intelligent audio and video splitting translation method with buffering strategy as described in any one of claims 1 to 7 includes: The data splitting module parses the real-time acquired audio and video data and splits the audio and video data through a preset data splitting model to obtain audio data and video data. The audio data block partitioning module is configured with a dynamic buffering mechanism. It dynamically adjusts the size of the audio buffer based on real-time monitored network transmission quality parameters, performs semantic analysis and data partitioning on the audio data in the audio buffer, and obtains a set of audio data blocks. The priority analysis module combines video data and audio data block sets to analyze the semantic importance of the data, calculate the corresponding priority, divide the video data into keyframes, and obtain the video keyframe priority and audio block priority. The split translation module translates the video keyframes and audio data blocks according to the video keyframe priority and audio block priority, respectively, using a preset translation model to obtain a first translation result and a second translation result. The real-time translation module analyzes the mapping relationship between the first and second translation results in chronological order and dynamically adjusts the display time of the translation results to obtain real-time translation results.