A short video-based multilingual intelligent translation method and system

By recognizing voice data and scene combinations in short videos, matching country names and voice types, the problem of inaccurate translation in existing technologies is solved, enabling accurate translation and updating of multiple voice videos.

CN121078267BActive Publication Date: 2026-07-31QUEEN BEE NETWORK TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
QUEEN BEE NETWORK TECH (SHENZHEN) CO LTD
Filing Date
2025-08-26
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

In existing technologies, short video translation methods cannot be compatible with multiple voice data and the types of voices to be translated, resulting in inaccurate translations and the inability to update multiple voice videos.

Method used

By identifying multiple voice data points through voice detection, recognizing voice content and scenarios, constructing scenario combinations, matching country names and voice types, and achieving intelligent translation, the system can also decompose the voice and video into multiple scenario regions for translation updates when necessary.

Benefits of technology

It achieves accurate translation of multiple audio and video data, taking into account the overall consideration of multiple audio data and audio types, thus ensuring the accuracy of translation and the accuracy of updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121078267B_ABST
    Figure CN121078267B_ABST
Patent Text Reader

Abstract

This invention discloses a multilingual intelligent translation method and system based on short videos. The invention relates to the technical field of intelligent translation methods. It determines an intelligent translation mode for the audio content based on multiple audio data and the types of audio to be translated. This intelligent translation mode triggers the intelligent translation of each audio segment and outputs the corresponding translated audio segments. The translated audio segments are labeled with the corresponding country names, improving the accuracy of the translated audio segments. Therefore, multiple audio videos are determined based on the playback nodes of multiple translated audio segments, the corresponding audio scenes, and the video content of the short video. If the multiple audio videos need to be changed, they are decomposed into multiple audio scene regions. The audio content to be translated is adjusted according to the change of audio scene regions, and new audio segments are determined based on the changed audio scenes and the audio content to be translated to update the multiple audio videos, ensuring the accuracy of the new audio segments and achieving the updating of the multiple audio videos.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of intelligent translation methods, and more particularly to a multilingual intelligent translation method and system based on short videos. Background Technology

[0002] With the development of technology, short videos are gradually being applied to people's lives. Short videos are watched by users and presented with corresponding video content. They have corresponding video content in different video scenarios. In the current technology, the translation of short videos requires a holistic translation, which translates the video and audio of the short video in a unified way, such as into English, Japanese or Mandarin. This overall translation requires human intervention and does not take into account the overall consideration of multiple audio data and the types of audio to be translated, which affects the accuracy of the translated audio segments and makes it impossible to update multiple audio and video. Summary of the Invention

[0003] The purpose of this invention is to overcome the shortcomings of the prior art. This invention provides a multilingual intelligent translation method and system based on short videos.

[0004] This invention provides a multilingual intelligent translation method based on short videos, including:

[0005] Multiple voice data are determined based on the voice detection of short videos, and the corresponding voice content and voice scene are determined based on the recognition of multiple voice data.

[0006] Scene combinations are constructed based on various speech scenarios. The corresponding country name is determined based on the recognition of the scene combination, and the corresponding speech type is matched based on the country name.

[0007] The intelligent translation mode is determined based on multiple voice data and the types of voices to be translated. This intelligent translation mode triggers the intelligent translation of each voice content and outputs the corresponding translated voice segments, which are marked with the corresponding country names.

[0008] Multiple translated audio segments are collected, and playback nodes for the multiple translated audio segments are determined based on the matching of the multiple translated audio segments and short videos. Multiple audio-video segments are determined based on the playback nodes of the multiple translated audio segments, the corresponding audio scenes, and the video content of the short videos.

[0009] If multiple audio and video segments need to be replaced, they are broken down into multiple audio scene regions. The audio content to be translated is adjusted according to the change of the audio scene region, and a new audio segment is determined based on the changed audio scene and the audio content to be translated, so as to update the multiple audio and video segments.

[0010] This invention provides a multilingual intelligent translation system based on short videos. The system is applied to the aforementioned multilingual intelligent translation method based on short videos. The multilingual intelligent translation system based on short videos includes:

[0011] The voice module is used to determine multiple voice data based on the voice detection of short videos, and to determine the corresponding voice content and voice scene based on the recognition of multiple voice data.

[0012] The speech category module is used to construct scene combinations based on various speech scenarios, determine the corresponding country name based on the recognition of the scene combination, and match the corresponding speech category based on the country name.

[0013] The country name module is used to determine the intelligent translation mode of the speech content based on multiple speech data and the speech type to be translated. This intelligent translation mode triggers the intelligent translation of each speech content and outputs the corresponding translated speech segment, which is marked with the corresponding country name.

[0014] The multi-audio-video module is used to acquire multiple translated audio segments, determine the playback nodes of multiple translated audio segments based on the matching of multiple translated audio segments and short videos, and determine the multi-audio-video based on the playback nodes of multiple translated audio segments, the corresponding audio scenes, and the video content of the short videos.

[0015] The update module is used to break down multiple audio and video segments into multiple audio scene regions if multiple audio and video segments need to be replaced. It adjusts the audio content to be translated according to the change of audio scene regions, and determines new audio segments based on the changed audio scene and the audio content to be translated, so as to update the multiple audio and video segments.

[0016] Compared with the prior art, the beneficial effects of the present invention are:

[0017] In this embodiment of the invention, the method involves determining multiple voice data points based on voice detection in a short video, identifying the corresponding voice content and voice scene based on the recognition of these multiple voice data points, constructing scene combinations based on each voice scene, determining the corresponding country name based on the recognition of the scene combinations, and matching the corresponding voice type based on the country name. An intelligent translation mode for the voice content is determined based on the multiple voice data points and the voice type to be translated. This intelligent translation mode triggers the intelligent translation of each voice content point and outputs the corresponding translated voice segment. The translated voice segment is marked with the corresponding country name. By introducing voice types and considering the overall compatibility of multiple voice data points and the voice types to be translated, the accuracy of the intelligent translation mode for the voice content is ensured, and the accuracy of the translated voice segment is improved.

[0018] Therefore, multiple translated audio segments are collected, and playback nodes for these segments are determined based on the matching of the translated audio segments with short videos. Multiple audio-video segments are then defined based on their playback nodes, corresponding audio scenes, and the video content of the short videos. If the multiple audio-video segments need to be replaced, they are decomposed into multiple audio scene regions. The audio content to be translated is adjusted according to the change in the audio scene region, and new audio segments are determined based on the changed audio scene and the audio content to be translated to update the multiple audio-video segments. This introduction of audio scene region replacement allows for a holistic consideration of the changed audio scene and the audio content to be translated, ensuring the accuracy of the new audio segments and enabling the updating of the multiple audio-video segments. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the multilingual intelligent translation method based on short videos in an embodiment of the present invention.

[0020] Figure 2 This is a flowchart illustrating step S11 in the multilingual intelligent translation method based on short videos according to an embodiment of the present invention.

[0021] Figure 3 This is a flowchart illustrating step S12 in the multilingual intelligent translation method based on short videos according to an embodiment of the present invention.

[0022] Figure 4 This is a flowchart illustrating step S13 in the multilingual intelligent translation method based on short videos in this embodiment of the invention.

[0023] Figure 5 This is a flowchart illustrating step S14 of the multilingual intelligent translation method based on short videos in an embodiment of the present invention.

[0024] Figure 6 This is a flowchart illustrating step S15 of the multilingual intelligent translation method based on short videos in an embodiment of the present invention.

[0025] Figure 7 This is a schematic diagram of the structural composition of a multilingual intelligent translation system based on short videos in an embodiment of the present invention. Detailed Implementation

[0026] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0027] Please see Figures 1 to 7 A multilingual intelligent translation method based on short videos is proposed and applied to multilingual intelligent translation scenarios. The short video-based multilingual intelligent translation method includes:

[0028] Step S11: Determine multiple voice data based on the voice detection of the short video, and determine the corresponding voice content and voice scene based on the recognition of the multiple voice data;

[0029] Step S12: Construct scene combinations based on each speech scene, determine the corresponding country name based on the recognition of the scene combination, and match the corresponding speech type based on the country name;

[0030] Step S13: Based on multiple voice data and the types of voices to be translated, determine the intelligent translation mode of the voice content. This intelligent translation mode triggers the intelligent translation of each voice content and outputs the corresponding translated voice segments, which are marked with the corresponding country names.

[0031] Step S14: Collect multiple translated speech segments, determine the playback nodes of multiple translated speech segments based on the matching of multiple translated speech segments and short videos, and determine multiple speech and video segments based on the playback nodes of multiple translated speech segments, the corresponding speech scenes, and the video content of the short videos.

[0032] Step S15: If multiple audio and video need to be replaced, the multiple audio and video are decomposed into multiple audio scene regions. The audio content to be translated is adjusted according to the change of audio scene regions. New audio segments are determined according to the changed audio scene and the audio content to be translated in order to update the multiple audio and video.

[0033] refer to Figure 2 In step S11, the specific steps are as follows:

[0034] S111: On the short video platform, collect the user's selected favorite videos, and determine the corresponding short video based on the detection of the favorite video. Based on the short video, the corresponding favorite tags and the user's viewing history on the short video platform, determine multiple key voice data. At this time, multiple key voice data are presented in the same short video and the corresponding key factors are marked.

[0035] S112: Determine the recognition mode of the speech data based on the data location and key factors of multiple key speech data, trigger the autonomous recognition of multiple speech data based on the recognition mode, collect multiple speech features, and determine the corresponding speech content and speech scene based on the selection of multiple speech features.

[0036] In the embodiments of this application, the user's selected favorite videos are collected on the short video platform, and the corresponding short video is determined based on the detection of the favorite video. Multiple key voice data are determined based on the short video, the corresponding favorite tags, and the user's viewing history on the short video platform. At this time, multiple key voice data are presented in the same short video and the corresponding key factors are marked. This takes into account the overall consideration of the short video, the corresponding favorite tags, and the user's viewing history on the short video platform, and ensures the accuracy of multiple key voice data.

[0037] At this point, the system first obtains the "favorite videos" actively selected or marked by users from the short video platform. These videos are content that users have liked, collected, subscribed to, or explicitly marked as interesting. The system records the basic information of these videos, including metadata such as video ID, duration, and publisher. It then obtains user behavior data through the platform API, builds user interest profiles, and identifies video content that users have explicitly expressed their preferences for.

[0038] The system analyzes the videos selected by users to extract their content features (such as theme, style, language type, etc.), and then filters out short videos with similar features from the platform database. This step realizes the transition from "user's explicit preferences" to "related content expansion". Using content analysis algorithms (such as visual feature extraction and audio feature analysis) and collaborative filtering technology, other short videos with similar content to the videos selected by users are found.

[0039] The system comprehensively analyzes information from three dimensions: the audio content of the short video itself, video-related interest tags (such as "technology," "food," "travel," etc.), and the user's historical viewing records (viewing duration, interaction frequency, etc.). Through this information, the system identifies the most important audio segments in the video, which are usually the core parts of the content or the parts that users are most interested in. Using multimodal fusion analysis, combining audio signal processing, natural language processing, and user behavior analysis, the system determines key audio data through weight calculation. For each identified key audio data, the system labels corresponding key factors, including: the temporal position of the audio in the video, the importance level of the audio content, the scene category to which the audio belongs, and the linguistic features of the audio, establishing a multi-dimensional labeling system and assigning structured label information to each key audio data.

[0040] Furthermore, the recognition pattern of the voice data is determined based on the data location and key factors of multiple key voice data. The recognition pattern triggers the autonomous recognition of multiple voice data and collects multiple voice features. The corresponding voice content and voice scenario are determined based on the selection of multiple voice features. This approach takes into account the overall consideration of selecting multiple voice features and ensures the accuracy of the corresponding voice content and voice scenario.

[0041] At this point, the system analyzes the key speech data acquired in S111, mainly based on two dimensions: data location: timestamp information of the speech in the video, such as start time, end time, duration, etc.; key factors: including importance level, scene category, language features, and other labeling information. Based on this information, the system selects the most suitable recognition mode for different types of speech data. For example: high importance + complex language features → high-precision deep learning recognition mode; medium importance + clear speech → standard speech recognition mode; low importance + background noise → noise reduction and enhancement recognition mode. Through pattern matching algorithms, the speech features are matched with the preset recognition mode library, and the optimal recognition strategy is automatically selected.

[0042] Based on the determined recognition pattern, the system activates the corresponding speech recognition engine to process key speech data. This process is "autonomous," meaning that the system can: automatically call the appropriate recognition model, adaptively adjust recognition parameters, process multiple speech data segments in parallel, monitor recognition quality in real time and adjust strategies when necessary, adopt multi-threaded processing technology, run different speech recognition engines simultaneously, and evaluate the recognition effect in real time through the quality monitoring module.

[0043] During speech recognition, the system extracts various speech features, including: acoustic features such as pitch, volume, intensity, and speech rate; linguistic features such as language type, dialect, accent, and technical terms; emotional features such as the speaker's emotional state and tone; and environmental features such as background noise, echo, and reverberation. These features are stored in structured data form to provide a foundation for subsequent analysis. Signal processing techniques and machine learning algorithms are used to extract multidimensional feature vectors from the original audio.

[0044] The system filters and analyzes the collected speech features to ultimately determine: Speech content: the recognized speech is converted into text and semantic information is annotated; Speech scene: the type of scene in which the dialogue occurs is determined based on the speech features, such as teaching, dialogue, explanation, etc.; The filtering process includes: removing low-confidence features, merging similar features, identifying the correlation between features, and optimizing understanding based on context; a feature selection algorithm and a scene classification model are used, combined with context analysis technology, to improve the accuracy of content recognition and scene judgment.

[0045] refer to Figure 3 In step S12, the specific steps are as follows:

[0046] S121: Collect various speech scenes, determine the corresponding scene elements based on the detection of each speech scene, determine the first combination of speech scenes based on multiple scene elements and the presentation position of each speech scene, and determine the second combination of speech scenes based on the scene content and individual scene elements of each speech scene.

[0047] S122: Construct scene combinations based on the first and second combinations of voice scenes, and determine multiple complete sub-scene combinations according to the division of scene combinations. Each sub-scene combination presents the corresponding national attractions or national marker features.

[0048] S123: Multiple country feature elements are determined based on the recognition of multiple complete sub-scene combinations. The corresponding country name is determined based on the multiple country feature elements, the shooting IP of the short video, and the video information. The corresponding voice type is determined based on the matching of the country name and the voice database. Each country has a corresponding voice type.

[0049] In the embodiments of this application, various voice scenes are collected, and corresponding scene elements are determined based on the detection of each voice scene. A first combination of voice scenes is determined based on multiple scene elements and the presentation position of each voice scene. A second combination of voice scenes is determined based on the scene content and individual scene elements of each voice scene. This approach takes into account the overall consideration of the scene content and individual scene elements of each voice scene, ensuring the accuracy of the second combination of voice scenes.

[0050] At this point, the system first collects all the speech scenes identified in step S112, and then performs in-depth analysis on each scene to extract key scene elements. Scene elements include: environmental elements: such as indoor / outdoor, city / country, specific place type, etc.; theme elements: such as tourism, food, education, entertainment, etc.; person elements: such as number of speakers, identity characteristics, interaction methods, etc.; object elements: key items or landmarks appearing in the scene; language elements: language type, accent characteristics, professional terminology, etc. The system analyzes video footage using computer vision technology, analyzes environmental sounds using audio processing technology, and analyzes speech content using natural language processing technology to comprehensively extract multi-dimensional features.

[0051] The system analyzes the distribution patterns of scene elements along the timeline, grouping spatially or temporally adjacent and element-similar speech scenes together. The first grouping focuses on: temporal continuity (adjacency of scenes along the video timeline, i.e., scenes that are close in time belong to the same group); spatial consistency (whether the physical spaces where scenes occur are consistent, such as the same indoor location or the same outdoor area); and element similarity (whether different scenes share key environmental or object elements). A temporal clustering algorithm is used to group scenes based on timestamps and element similarity matrices, setting time and element similarity thresholds as grouping conditions. The first grouping categorizes speech scenes from the perspectives of time, space, and element similarity, emphasizing the continuity and consistency of scenes in the physical dimension. Its core is to group temporally adjacent, spatially consistent, and element-similar speech scenes together using a temporal clustering algorithm.

[0052] The system analyzes speech scenarios from both semantic and content perspectives, grouping scenarios that are thematically related and content-coherent. The second grouping focuses on: thematic consistency (whether scenarios revolve around the same theme or topic); content coherence (whether there is logical or narrative continuity between scenarios); character association (whether the same characters or roles appear in scenarios); and language style consistency (whether the language expression styles in scenarios are similar). Using topic models and semantic similarity calculations, combined with narrative structure analysis, the system identifies content-related scenario combinations. The second grouping analyzes speech scenarios from both semantic and content perspectives, emphasizing the coherence of theme, content, characters, and language style. Its core is to group thematically related and content-coherent scenarios together using topic models and semantic similarity calculations.

[0053] Specifically, suppose there is a short video about an "Asian culinary journey," 10 minutes long, containing multiple audio scenarios; the system first identifies the following audio scenarios and their elements:

[0054] Scene A (0:00-1:30): Environmental elements: indoor, kitchen environment, bright lighting; Theme elements: food preparation, teaching; Character elements: a chef (Asian appearance), narrator; Object elements: wok, knives, various vegetables, seasonings; Language elements: Mandarin Chinese, cooking terminology;

[0055] Scene B (1:31-2:45): Environmental elements: indoor, restaurant environment, warm lighting; Thematic elements: food display and tasting; Character elements: same chef, several diners; Object elements: finished dishes, tableware, table; Language elements: Mandarin Chinese, food evaluation vocabulary;

[0056] Scene C (2:46-4:20): Environmental elements: outdoors, street market, natural light; Theme elements: food procurement, market exploration; Character elements: the same chef, market vendors; Object elements: various fresh ingredients, stalls, shopping bags;

[0057] Language elements: Standard Chinese, market trading terminology;

[0058] Scene D (4:21-6:15): Environmental elements: indoor, another kitchen environment, different layout; Theme elements: food preparation, teaching; Character elements: another chef (Japanese face), narrator; Object elements: sushi tools, seafood, rice; Language elements: Japanese, cooking terminology;

[0059] Scene E (6:16-7:50): Environmental elements: indoor setting, Japanese restaurant, traditional decorations; Theme elements: food display and tasting; Character elements: the same Japanese chef and several diners; Object elements: finished sushi, Japanese tableware, dining table; Language elements: Japanese, food evaluation vocabulary;

[0060] Scene F (7:51-10:00): Environmental elements: outdoor, Japanese street scene, night view; Theme elements: cultural introduction, street food; Character elements: host, street food vendor; Object elements: various Japanese street food stalls; Language elements: Japanese, Mandarin Chinese (narration).

[0061] The system analyzes scene elements and their presentation locations to form the first combination:

[0062] First combination 1: Scene A + Scene B + Scene C; Temporal continuity: The three scenes are consecutive on the timeline (0:00-4:20); Spatial consistency: Although Scene A and B are different indoor spaces, they are both located in the same building; Although Scene C is outdoors, it is geographically adjacent to Scene B; Element similarity: All three scenes involve Chinese cooking, have the same characters (chefs), and have a coherent theme (from production to display to procurement).

[0063] First combination 2: Scene D + Scene E + Scene F; Temporal continuity: The three scenes are consecutive on the timeline (4:21-10:00); Spatial consistency: Scenes D and E are located in different areas of the same building; Although Scene F is outdoors, it is geographically adjacent to Scene E; Element similarity: All three scenes involve Japanese food culture, have the same characters (Japanese chefs), and have a coherent theme (from production to display to cultural introduction).

[0064] The system analyzes the scene content and scene elements to form a second combination:

[0065] The second group 1: Scene A + Scene D; Theme consistency: Both are food preparation teaching scenes; Content coherence: Showcasing cooking techniques from different cultural backgrounds; Character relevance: Both are taught by professional chefs; Language style consistency: Both use instructional language and include professional terminology.

[0066] Group 2: Scene B + Scene E; Thematic consistency: Both are scenes of food display and tasting; Content coherence: Both showcase food experiences from different cultural backgrounds; Character connection: Both involve interaction between chefs and diners; Language style consistency: Both use descriptive language and include sensory evaluation vocabulary.

[0067] The second group 3: Scene C + Scene F; Theme consistency: Both scenes explore local ingredients and culture; Content coherence: Both showcase the ways of obtaining ingredients in different cultural contexts; Character connection: Both involve interaction between the host and locals; Language style consistency: Both use exploratory language and include elements of cultural exchange.

[0068] Through this series of processes, the system successfully organized the original six speech scenes into two different combinations: the first combination is based on spatiotemporal continuity and physical spatial relationships, and the second combination is based on content themes and semantic associations. These two combination methods provide multi-angle scene association information for the subsequent scene combination construction (S122), enabling the system to more comprehensively understand the structure and logic of the video content, laying the foundation for accurate country identification and speech type matching.

[0069] Furthermore, scene combinations are constructed based on the first and second combinations of voice scenes, and multiple complete sub-scene combinations are determined according to the division of scene combinations. Each sub-scene combination presents the corresponding national attractions or national marker features, which is compatible with the overall consideration of scene combination division and ensures the accuracy of multiple complete sub-scene combinations.

[0070] At this point, the first combination (based on spatiotemporal continuity) and the second combination (based on content theme consistency) generated in S121 are merged to construct a higher-level "scene combination". This combination is not a simple superposition, but a comprehensive consideration of spatiotemporal proximity and semantic relevance to form a more complete and semantically coherent "scene combination". At the same time, a graph structure or clustering algorithm is used to take the first combination and the second combination as input nodes and perform association analysis through similarity calculation (such as cosine similarity, Jaccard similarity). A combination threshold is set, and only when both combination methods support classifying certain scenes into one category are they confirmed as a "scene combination". Each scene combination must meet three conditions: spatiotemporal proximity, semantic coherence, and theme consistency.

[0071] After constructing the initial scene combination, the system will "divide" it, aiming to break down the large combination into multiple "sub-scene combinations". Each sub-scene combination should be a semantically complete and structurally independent content unit, and should be able to clearly present the iconic features of a certain country or region (such as famous attractions, cultural symbols, architectural styles, language habits, etc.). At the same time, temporal segmentation algorithms (such as sliding window, change point detection) are used to segment the scene combination. Each sub-scene combination must meet the following conditions: have clear start and end time points; contain multiple speech scenes with consistent themes; and contain identifiable national or regional marker features (such as visual features like the Eiffel Tower, Mount Fuji, Taj Mahal, etc., or linguistic features like Japanese, French, Italian, etc.).

[0072] Each sub-scene combination must not only be complete in content but also possess "country marker features," which clearly point to a specific country or region through multimodal information such as visual, auditory, and linguistic features. These features can include: visual features such as landmarks, natural landscapes, clothing styles, and streetscapes; auditory features such as background music, dialect accents, and environmental sounds (such as market vendors' calls and subway announcements); and linguistic features such as language types, vocabulary habits, and cultural terms. Simultaneously, a multimodal fusion model (such as CLIP) is used to jointly embed visual and linguistic features and calculate the similarity with the country marker features. A country marker feature database is established, containing typical visual, auditory, and linguistic feature vectors from various countries, used to match sub-scene combinations. Feature extraction and matching are performed on each sub-scene combination to determine its corresponding country or region.

[0073] Specifically, suppose we have a short video about an "Asian Food Tour" containing the following audio scenarios: Scenario A: Introducing a ramen shop on the streets of Tokyo (Japanese, with Tokyo Tower in the background); Scenario B: Explaining matcha culture inside a temple in Kyoto (Japanese, with Kinkaku-ji Temple in the background); Scenario C: Introducing Korean BBQ on the streets of Seoul (Korean, with Namsan Tower in the background); Scenario D: Introducing seafood at a haenyeo (female divers) village in Jeju Island (Korean, with a haenyeo statue in the background); Scenario E: Introducing Pad Thai at a night market in Bangkok (Thai, with the Grand Palace in the background); Scenario F: Explaining traditional dance at a temple in Chiang Mai (Thai, with Wat Phra That Doi Suthep in the background).

[0074] First combination (based on spatiotemporal continuity): Combination 1: Scene A + Scene B (both in Japan, continuous in time); Combination 2: Scene C + Scene D (both in South Korea, continuous in time); Combination 3: Scene E + Scene F (both in Thailand, continuous in time); Second combination (based on the consistency of content theme): Combination 1: Scene A + Scene C + Scene E (all street food introductions); Combination 2: Scene B + Scene D + Scene F (all cultural explanation scenes).

[0075] The system found that the first and second combinations highly overlapped in the three regions of "Japan", "Korea", and "Thailand", and therefore constructed three scene combinations: Japan combination: Scene A + Scene B; Korea combination: Scene C + Scene D; Thailand combination: Scene E + Scene F. The system performed an integrity check on each scene combination and found that each combination already met the requirements of semantic integrity and structural independence, so it was directly used as a sub-scene combination: Sub-scene combination 1: Japanese food and culture (Scene A + Scene B); Sub-scene combination 2: Korean food and culture (Scene C + Scene D); Sub-scene combination 3: Thai food and culture (Scene E + Scene F).

[0076] Sub-scene combination 1 (Japan): Visual features: Tokyo Tower, Kinkaku-ji (Golden Pavilion), traditional Japanese architecture; Auditory features: Japanese narration, traditional Japanese background music; Linguistic features: Japanese vocabulary (such as "ramen" and "matcha"); Country marker feature matching: Japan (95% matching rate);

[0077] Sub-scene combination 2 (South Korea): Visual features: Namsan Tower, Haenyeo statue, Korean street scene; Auditory features: Korean narration, Korean background music; Linguistic features: Korean vocabulary (such as "grilled meat" and "haenyeo"); Country marker feature matching: South Korea (match rate 92%);

[0078] Sub-scene combination 3 (Thailand): Visual features: Grand Palace, Wat Phra That Doi Suthep, Thai night market; Auditory features: Thai narration, traditional Thai music; Linguistic features: Thai vocabulary (such as "stir-fried rice noodles", "traditional dance"); Country marker feature matching: Thailand (match rate 96%).

[0079] The system successfully organized the original multiple speech scenarios into three sub-scenarios with clear country markers. Each sub-scenarios is not only complete in content and independent in structure, but also clearly points to a specific country through multimodal features. This refined scenario organization method provides high-quality input data for subsequent country identification (S123) and speech type matching, ensuring the accuracy and cultural adaptability of multilingual intelligent translation.

[0080] Therefore, multiple country feature elements are determined based on the recognition of multiple complete sub-scene combinations. The corresponding country name is determined based on multiple country feature elements, the shooting IP of the short video, and video information. The corresponding voice type is determined based on the matching of the country name and the voice database. Each country has a corresponding voice type, which takes into account the overall consideration of matching country name and voice database, and ensures the accuracy of the corresponding voice type.

[0081] At this point, the system will conduct an in-depth analysis of the various sub-scene combinations generated by S122, extracting country-related feature elements. These elements may include: visual features: such as famous scenic spots, architectural styles, clothing, natural landscapes, etc.; auditory features: such as background music, language tone, ambient sounds, etc.; text features: such as subtitles, signs, interface text, etc.; behavioral features: such as etiquette, customs, actions, etc.; and object features: such as specialty foods, handicrafts, vehicles, etc. Simultaneously, computer vision models (such as YOLO, ResNet) are used to identify visual elements in the scene; audio classification models (such as VGGish) are used to analyze background music and ambient sounds; text information is extracted using OCR technology; and a multimodal fusion model is used to comprehensively analyze the above features to form a country feature vector.

[0082] The system cross-validates the extracted country feature elements with the short video's metadata (such as shooting IP address, upload location, video title, description, tags, etc.) to ultimately determine the country name corresponding to each sub-scene combination; it compares the country feature elements with a predefined country feature database to calculate the matching degree; it combines IP geographic location information (such as IP location query API) to obtain the shooting location; it uses natural language processing (NLP) technology to analyze geographic keywords in the video title, description, and tags; and through a weighted voting mechanism, it integrates information from multiple sources to output the country name.

[0083] Once the country name is determined, the system queries the built-in voice database to match the commonly used speech types (including languages, dialects, accents, etc.) for that country, providing a linguistic foundation for subsequent translation or dubbing. The voice database contains the major languages, dialects, accents, and usage scenarios of each country. The matching logic considers not only official languages ​​but also regional languages, minority languages, and commonly used tourist languages. For multilingual countries (such as Switzerland and India), the language selection will be further refined based on the context. The output results include information such as the main language, secondary languages, and recommended accents.

[0084] Specifically, suppose the short video contains the following three sub-scene combinations (from S122): Sub-scene combination 1: includes Tokyo Tower, Kinkaku-ji Temple, Japanese garden, and Japanese narration; Sub-scene combination 2: includes Eiffel Tower in Paris, French narration, and French cafe; Sub-scene combination 3: includes Statue of Liberty in New York, English narration, and American street scene.

[0085] Sub-scene combination 1: Visual features: Tokyo Tower, Kinkaku-ji Temple, Japanese garden; Auditory features: Japanese narration, traditional Japanese music; Textual features: Japanese subtitles, "Tokyo" sign; Behavioral features: Bowing etiquette; Item features: Kimono, matcha bowl.

[0086] Sub-scene combination 2: Visual features: Eiffel Tower, French architecture, coffee shop; Auditory features: French narration, French chanson music; Textual features: French subtitles, "Paris" logo; Behavioral features: French cheek kiss; Object features: baguette, coffee cup.

[0087] Sub-scene combination 3: Visual features: Statue of Liberty, Manhattan skyline; Auditory features: American English narration, jazz music; Textual features: English subtitles, "NYC" sign; Behavioral features: American waving greeting; Object features: hot dog, taxi.

[0088] Sub-scene combination 1: Country feature element matching: Japan (98% match); Filming IP: Tokyo (Japan); Video tags: #JapanTravel #Tokyo; Country name confirmed: Japan; Sub-scene combination 2: Country feature element matching: France (96% match); Filming IP: Paris (France); Video title: Romantic Trip to Paris; Country name confirmed: France; Sub-scene combination 3: Country feature element matching: USA (97% match); Filming IP: New York (USA); Video description: Explore iconic landmarks in New York; Country name confirmed: USA; Japan: Primary language: Japanese (standard Tokyo accent); Secondary language: English (common for travel); Recommended accent: Standard Japanese (NHK style); France: Primary language: French (Parisian accent); Secondary language: English (common for travel); Recommended accent: Standard French (Parisian accent); USA: Primary language: American English (general accent); Secondary language: Spanish (some regions); Recommended accent: American English (neutral accent).

[0089] The system successfully identified the three sub-scenes as Japan, France, and the United States, and determined the corresponding speech type for each country. This country identification method based on multimodal features and metadata cross-validation greatly improves the accuracy of identification and provides a reliable linguistic foundation for subsequent multilingual translation or dubbing. At the same time, the system can also flexibly handle the language selection of multilingual countries to ensure that the translation results conform to local language habits and cultural background.

[0090] refer to Figure 4 In step S13, the specific steps are as follows:

[0091] S131: Collect multiple voice data. In each voice data, determine the basic voice type of the voice data based on the recognition of the voice data. Determine the corresponding matching coefficient based on the matching of the basic voice type of the voice data and the voice type to be translated. Determine the intelligent translation mode of the voice content based on the matching coefficient, the voice type to be translated and the corresponding voice data.

[0092] S132: In the intelligent translation mode of speech content, multiple translation paths are determined based on the recognition of the intelligent translation mode of speech content, and the optimal translation path is determined based on the multiple translation paths, the speech content of the speech data, and the speech scene.

[0093] S133: Trigger autonomous translation of speech data based on the optimal translation path, and perform intelligent translation of each speech content to output the corresponding translated speech segments. The translated speech segments are marked with the corresponding country name and the corresponding speech scene.

[0094] In the embodiments of this application, multiple voice data are collected. In each voice data, the basic voice type of the voice data is determined based on the recognition of the voice data. The corresponding matching coefficient is determined based on the matching of the basic voice type of the voice data and the voice type to be translated. The intelligent translation mode of the voice content is determined based on the matching coefficient, the voice type to be translated and the corresponding voice data. This approach takes into account the overall consideration of the matching coefficient, the voice type to be translated and the corresponding voice data, ensuring the accuracy of the intelligent translation mode of the voice content.

[0095] At this point, the system obtains the processed key voice data from steps S11 and S112. This data already contains information such as voice content, timestamps, and key factors. The acquisition process includes: converting voice data of different formats into a unified format; ensuring that the voice data clarity meets the identifiable standard; performing preliminary classification according to dimensions such as scene, duration, and importance; and associating the voice data with corresponding video clips, subtitle information, etc.

[0096] For each speech data point, language recognition is performed to determine its original language type. This process includes: extracting acoustic features such as the speech spectrum, pitch, and rhythm; comparing the extracted features with known language models; scoring the credibility of the recognition results; segmenting and recognizing speech containing multiple languages; and further identifying dialect or accent features within the same language.

[0097] The matching coefficient is calculated based on the degree of matching between the basic speech type and the target translation language. The calculation is based on factors such as: language similarity (based on factors such as language family, grammatical structure, and vocabulary similarity); translation history data (analysis of historical translation quality data, such as accuracy and fluency); user preferences (considering user preferences for specific language translations); scenario adaptability (the importance weight of language matching in different scenarios); cultural relevance (the degree of cultural background relevance between languages); and the matching coefficient range (usually set between 0 and 1, with higher values ​​indicating better matching).

[0098] Based on the matching coefficient, target language, and voice data content, the most suitable translation mode is selected. Optional modes include: Literal translation mode: maintains the original text structure, suitable for technical and professional content; Free translation mode: focuses on conveying meaning, suitable for literary and expressive content; Cultural adaptation mode: considers cultural differences and performs localization; Hybrid mode: combines the flexible application of multiple translation strategies; Real-time optimization mode: dynamically adjusts strategies based on feedback during the translation process.

[0099] Specifically, suppose we have a short video about making Italian food, containing the following audio data: Audio data 1: timestamp (00:15-00:23), content "Now we're adding fresh basil leaves," key factor (ingredient introduction); Audio data 2: timestamp (00:45-00:58), content "Cook the pasta to aldente, this is the essence of pasta," key factor (cooking technique); Audio data 3: timestamp (01:20-01:35), content "Finally, sprinkle on Parmesan cheese, Buonappetito!", key factor (completion step); Audio data 1: the basic speech type is Chinese, confidence level 0.98; Audio data 2: the basic speech type is Chinese, confidence level 0.97; Audio data 3: the language is mixed, with the Chinese part having a confidence level of 0.96 and the Italian part "Buonappetito" having a confidence level of 0.99.

[0100] Assuming the target translation language is English, the matching coefficients are calculated as follows: Voice Data 1 (Chinese → English): Language Similarity: 0.7 (different language families but with extensive translation experience); Translation History Quality: 0.85 (high quality Chinese-English translation); User Preference: 0.9 (users prefer Chinese-English translation); Scene Adaptability: 0.8 (clear translation needs for cooking scenarios); Final Matching Coefficient: (0.7 + 0.85 + 0.9 + 0.8) / 4 = 0.81; Voice Data 2 (Chinese → English): Contains the Italian term "aldente," and terminology retention needs to be considered in the matching coefficient calculation; Final Matching Coefficient: 0.78 (slightly lower than pure Chinese content); Voice Data 3 (Mixed Language → English): Chinese Part Matching Coefficient: 0.81; Italian Part Matching Coefficient: 0.65 (limited historical data for Italian-English translation); Final Matching Coefficient: 0.73 (weighted average).

[0101] Based on the matching coefficient and content characteristics, the following translation modes were selected: Voice data 1 (matching coefficient 0.81): "Cultural Adaptation Mode" was selected; Reason: Cooking terminology requires cultural adaptation to ensure understanding by English-speaking audiences; Translation strategy: Retain the technical term "basil leaf," but adjust the expression to better suit the speaking habits of English cooking programs; Voice data 2 (matching coefficient 0.78): "Hybrid Mode" was selected; Reason: Technical terms require literal translation, but the expression requires interpretive translation; Translation strategy: Retain the original "aldente" and add explanatory translations, such as "cooked to aldente (chewy) state"; Voice data 3 (matching coefficient 0.73): "Cultural Adaptation Mode + Real-time Optimization Mode" was selected; Reason: Greetings require cultural adaptation, and mixed languages ​​require dynamic adjustment; Translation strategy: "Finally sprinkle with Parmesan cheese, enjoy your meal (Buonappetito)!", retain the original Italian text and add culturally equivalent translations.

[0102] Furthermore, in the intelligent translation mode of speech content, multiple translation paths are determined based on the recognition of the intelligent translation mode of speech content. The optimal translation path is determined based on the multiple translation paths, the speech content of the speech data, and the speech scene. This takes into account the overall consideration of multiple translation paths, the speech content of the speech data, and the speech scene, ensuring the accuracy of the optimal translation path.

[0103] At this point, the system has already determined an "intelligent translation mode" (such as "cultural adaptation mode," "literal translation mode," "hybrid mode," etc.) for each audio content. Based on this, step S132 further identifies multiple translation paths existing in each mode. Translation paths refer to different translation strategies, expressions, or technical implementation methods from the source language to the target language. Each translation mode corresponds to multiple paths. For example, the paths in the literal translation mode include: Path A: word-for-word translation, retaining the original sentence structure; Path B: lexical-level literal translation, but adjusting the word order to conform to the target language's habits; the paths in the cultural adaptation mode include: Path C: complete localization expression, replacing culturally specific elements; Path D: retaining the original cultural elements, adding explanatory annotations; the paths in the hybrid mode include: Path E: partial literal translation + partial free translation; Path F: dynamically switching between literal and free translation based on the content. The system will automatically generate all translation paths based on the translation mode library, historical translation data, and scenario rules for subsequent evaluation and selection.

[0104] After identifying multiple translation paths, the system needs to evaluate the applicability of each path by considering the following three key factors, and finally select the best path: Speech content features: including language complexity, density of technical terms, emotional tone, and expression style; Speech scene features: including scene type (e.g., teaching, entertainment, tourism), cultural background, and target audience; Translation path features: including translation accuracy, fluency, cultural suitability, and processing speed. The system will use a multi-dimensional scoring model to comprehensively score each path and select the path with the highest score as the "best translation path." Examples of evaluation indicators: Language accuracy (weight 40%); Cultural suitability (weight 30%); Naturalness of expression (weight 20%); Processing efficiency (weight 10%).

[0105] Therefore, the system triggers autonomous translation of speech data based on the optimal translation path and performs intelligent translation of each speech content to output the corresponding translated speech segments. The translated speech segments are marked with the corresponding country name and speech scene. This system introduces speech types and takes into account multiple speech data and speech types to be translated, ensuring the accuracy of the intelligent translation mode of speech content and improving the accuracy of the translated speech segments.

[0106] At this point, the optimal translation path for each segment of speech content has been determined in S132; the first step in S133 is to call the corresponding translation engine or model to automatically translate the speech content according to the selected path. This process includes: calling the corresponding translation model or API according to the translation path type (such as literal translation, cultural adaptation, hybrid mode, etc.); combining speech scene, context information, cultural background, etc. to help the translation model understand the context; the system automatically performs the translation and generates the target language text; and preliminarily checks the grammatical correctness, fluency, and semantic consistency of the translation results.

[0107] After translation is completed, the system organizes the translation results of each audio segment into "translated audio segments". This process includes: outputting the text content of the target language; if audio output is required, calling the TTS (Text-to-Speech) module to synthesize the translated text into audio; aligning the translated audio segments with the original video timeline to ensure synchronized playback; and encapsulating the translation results into structured data for easy subsequent processing or embedding in the video.

[0108] To facilitate subsequent management and display, each translated audio segment is marked with two key pieces of information: Country Name: indicating the target country or language region for which the translation is intended (e.g., "United States", "Japan", "France"); Audio Scene: indicating the scene type to which the audio segment belongs (e.g., "Tourism Introduction", "Teaching Explanation", "Food Sharing"). These tags can be stored as metadata or directly embedded into video subtitles, audio tracks, or player control information.

[0109] Specifically, suppose we have a short video of an Italian chef demonstrating how to make "Carbonara" on the streets of Rome; the target language is Chinese, and the target country is China; the optimal translation path is the cultural adaptation model (path C); the original audio content is: "PerlaCarbonaraautentica,usiamosologuanciale,uova,pecorinoepepenero.Nientep anna,nienteaglio!"; the system calls the "cultural adaptation translation model," which excels at transforming elements of Italian culinary culture into concepts familiar to Chinese audiences; the translation result is: "Authentic Carbonara uses only pork cheek meat (guanciale), eggs, Pecorino cheese, and black pepper; no cream, and no garlic!"

[0110] The system outputs the translated audio segment, generating the aforementioned Chinese text; it then calls the Chinese TTS module to generate standard Mandarin audio with natural intonation and moderate rhythm; the translated audio is synchronized with the lip movements and actions of the chef in the original video; the output consists of a subtitle file (.srt) with a timeline and a separate audio file (.wav); country name tag: Country: China; audio scene tag: Scene: Food Tutorial; the system completes the entire process from translation path to final translation result, including translation execution, speech synthesis, time alignment, and content tagging. This process not only ensures the accuracy and naturalness of the translated content but also provides important support for subsequent video editing, multilingual version management, and personalized recommendations through structured tagging; in practical applications, this mechanism can significantly improve the efficiency of multilingual adaptation of short videos and the user experience.

[0111] refer to Figure 5 In step S14, the specific steps are as follows:

[0112] S141: Collect multiple translated speech segments, determine the corresponding data volume based on the detection of translated speech segments, and mark two replacement nodes of the translated speech segments. At the same time, collect short videos and determine the corresponding replacement data segments based on the matching of short videos and translated speech segments.

[0113] S142: Determine the corresponding replacement area based on the replacement data segment, the two replacement nodes of the translated speech segment, and the data volume of the translated speech segment, and determine the corresponding playback node based on the detection of the replacement area, so as to determine the playback node of multiple translated speech segments;

[0114] S143: Determine the first sub-speech video based on the playback nodes of multiple translated speech segments and the corresponding speech scenes; determine the second sub-speech video based on the playback nodes of multiple translated speech segments and the video content of the short video; and determine multiple speech videos based on the first sub-speech video and the second sub-speech video.

[0115] In the embodiments of this application, multiple translated speech segments are collected, the corresponding data volume is determined based on the detection of the translated speech segments, and two replacement nodes of the translated speech segments are marked. At the same time, short videos are collected, and the corresponding replacement data segments are determined based on the matching of the short videos and the translated speech segments. This approach takes into account the overall consideration of matching short videos and translated speech segments, ensuring the accuracy of the corresponding replacement data segments.

[0116] At this point, the system retrieves multiple translated audio segments from step S133. These audio segments typically include: target language text (such as English, Chinese, Japanese, etc.); synthesized audio; timestamps (start and end times); country markers (such as "United Kingdom" and "Japan"); and audio scene markers (such as "travel introduction" and "food tutorial"). Each audio segment also has a unique ID for easy tracking and matching later. Audio segments are usually stored as audio files (such as .wav or .mp3) or audio streams. The timestamps of the audio segments are used to align with the playback timeline of the short video.

[0117] The system performs data volume detection on each translated audio segment to determine key information such as its size, duration, and text length. These data volume indicators directly affect subsequent replacement node marking and video synchronization. Text length: counts the number of characters or words corresponding to the audio segment; Audio duration: calculates the actual playback time (seconds) of the audio segment; Data size: measures the storage size of the audio file (MB); Time span: records the start and end times of the audio segment.

[0118] For each translated audio segment, two replacement nodes are marked. These two nodes determine the insertion position and duration of the audio segment in the video. The start replacement node is the earliest time point suitable for inserting new translation content, usually aligned with the start time of the original video audio. The end replacement node is the latest time point suitable for ending the replacement content, usually aligned with the end time of the original video audio. The marking of replacement nodes is based on timestamps and video content analysis to ensure that the audio and video are synchronized.

[0119] The system collects raw short videos, analyzes their content structure (such as camera transitions and changes in scene content), and matches them with translated audio segments to determine which video clips are suitable for audio replacement. Simultaneously, it retrieves raw video files from platforms or user uploads; uses computer vision technology to analyze video camera transitions, scene changes, and character actions; compares the timestamps of the translated audio segments with the video's timeline to find the best-matching video clips; and marks the video clips that need audio replacement, which typically have the same time span as the original audio.

[0120] Furthermore, the corresponding replacement area is determined based on the replacement data segment, the two replacement nodes of the translated speech segment, and the data volume of the translated speech segment. The corresponding playback node is then determined based on the detection of the replacement area, thus identifying the playback nodes for multiple translated speech segments. This approach incorporates the overall consideration of replacement area detection and ensures the accuracy of the corresponding playback nodes.

[0121] At this point, based on the replacement data segment determined in S141, the two replacement nodes (start node and end node) of the translated audio segment, and the data volume of the audio segment, the system accurately calculates the replacement area in the video. This process ensures that the translated audio can accurately replace the corresponding part in the original video. Simultaneously, the replacement area = replacement data segment + replacement node adjustment + data volume compensation. Start replacement node: the time point in the original video where replacement needs to begin; End replacement node: the time point in the original video where replacement needs to end; Data volume compensation: fine-tuning the replacement area based on the difference between the duration of the translated audio segment and the duration of the original audio segment; Optionally, mapping the timestamp of the audio segment to the video timeline; Comparing the duration difference between the original audio segment and the translated audio segment; Fine-tuning the boundary of the replacement area forward or backward based on the duration difference.

[0122] The system detects the identified replacement areas and identifies the most suitable playback node for playing the translated audio segment. The playback node is the precise time point at which the translated audio segment should begin playing during video playback. Playback node = start time of replacement area + synchronization offset. Synchronization offset: an adjustment value determined based on factors such as changes in video content and scene transitions. Node verification: ensuring that the playback node is consistent with the points of change in video content (such as scene transitions, character actions, etc.). At the same time, computer vision technology is used to identify scene change points in the video, analyze the correspondence between video content and audio content, and fine-tune the playback node based on the content synchronization analysis results to obtain the best viewing experience.

[0123] The system repeats the above process for all translated audio segments to determine the precise playback node for each segment in the video. This step ensures that multiple audio segments can be seamlessly connected in the video, forming a smooth multilingual viewing experience. The system arranges the playback nodes of all audio segments in chronological order and checks whether the time intervals between adjacent playback nodes are reasonable, ensuring that all playback nodes are synchronized with the overall video content. Optionally, the playback node sequence is optimized as a whole to avoid conflicts or overlaps, and appropriate transition processing is added between adjacent audio segments to ensure smooth switching. Simulated playback verifies the accuracy and rationality of all playback nodes. The system implements the conversion process from changing data segments to precise playback nodes. This step ensures that the translated audio segments can be accurately located and played in the video, solving the time synchronization problem in multilingual videos. In practical applications, this mechanism can significantly improve the multilingual adaptation quality of short videos, ensuring that global users receive a smooth and natural viewing experience.

[0124] Therefore, a first sub-audio video is determined based on the playback nodes of multiple translated audio segments and the corresponding audio scenarios, and a second sub-audio video is determined based on the playback nodes of multiple translated audio segments and the video content of the short video. Multiple audio videos are determined based on the first and second sub-audio videos, which takes into account the overall consideration of the first and second sub-audio videos and ensures the accuracy of multiple audio videos.

[0125] At this point, based on the playback nodes of the multiple translated audio segments determined in S142, and the corresponding audio scenes (such as "food introduction," "cooking techniques," "cultural background," etc.) for each audio segment, the system constructs a sub-video dominated by audio scenes, called the "first sub-audio video." The first sub-audio video is a sub-video constructed dominated by audio scenes. Its core is to organize the translated audio segments according to their playback nodes and corresponding audio scenes (such as "food introduction," "cooking techniques," "cultural background," etc.), ensuring accurate mapping of audio content to the video timeline, semantic alignment of audio and visual content, and adding transition effects at scene transitions to ensure video smoothness. At the same time, the audio scenes are semantically aligned with the visual content in the video to ensure consistency between audio content and visual content; playback nodes are mapped to the video timeline to ensure audio segments are inserted at the correct time points; and transition effects are added at scene transitions to ensure video smoothness.

[0126] The system uses the same playback nodes, but this time it combines the video content of the short video itself (such as scene changes, camera transitions, and subtitle appearances) to construct a sub-video dominated by the video content, called the "second sub-audio-video." The second sub-audio-video is a sub-video constructed primarily based on video content. Its core lies in matching the playback nodes of the audio segment with key points in the video content, ensuring that the audio segment and video content change synchronously. Simultaneously, it detects keyframes and adjusts the playback nodes to guarantee the continuity between audio and video content. The system analyzes key points in the short video such as scene changes, camera transitions, and subtitle appearances; matches the playback nodes of the audio segment with key points in the video content; ensures that the audio segment and video content change synchronously; simultaneously detects keyframes in the video (such as camera transitions and subtitle appearances); aligns the playback nodes of the audio segment with keyframes; checks the continuity between audio and video content, and adjusts the playback nodes as necessary.

[0127] The system merges the first sub-audio-video (scene-driven) and the second sub-audio-video (content-driven) to generate the final multi-audio-video. This process ensures the best match between the audio scene and the video content. The system also aligns the two sub-videos on the timeline and merges their content. It handles conflicts between the two sub-videos (such as inconsistent playback nodes) and optimizes the quality of the merged video to ensure smooth playback.

[0128] Specifically, suppose we have a short video titled "Italian Food Preparation" that is 2 minutes long (00:00:00-00:02:00); the system has generated three translated audio segments (A, B, C) and determined their playback nodes.

[0129] Audio Segment A: Playback Node: 00:00:10; Audio Scene: "Food Introduction"; Corresponding Video Content: Showcasing the finished pasta dish; Audio Segment B: Playback Node: 00:00:41; Audio Scene: "Cooking Techniques"; Corresponding Video Content: Chef demonstrating pasta cooking techniques; Audio Segment C: Playback Node: 00:01:31; Audio Scene: "Cultural Background"; Corresponding Video Content: Showcasing Italian food culture; Construction process of the first sub-audio-video (scene-driven): Insert audio segment A at 00:00:10 to match the "Food Introduction" scene; Insert audio segment B at 00:00:41 to match the "Cooking Techniques" scene; Insert audio segment C at 00:01:31 to match the "Cultural Background" scene; Ensure natural scene transitions, such as switching the visuals from the finished product display to the cooking process when transitioning from "Food Introduction" to "Cooking Techniques".

[0130] Audio segment A: Playback point: 00:00:10; Key point in video content: At 00:00:10, the camera switches to the finished product display; Audio segment B: Playback point: 00:00:41; Key point in video content: At 00:00:41, the camera switches to the noodle cooking process; Audio segment C: Playback point: 00:01:31; Key point in video content: At 00:01:31, the camera switches to the cultural display; Construction process of the second sub-audio-video: Insert audio segment A at 00:00:10, synchronizing with the camera switching to the finished product display; Insert audio segment B at 00:00:41, synchronizing with the camera switching to the noodle cooking process; Insert audio segment C at 00:01:31, synchronizing with the camera switching to the cultural display; Ensure that the audio segments and video content changes are completely synchronized.

[0131] Comparing the playback points of the first and second sub-audio videos, they were found to be completely consistent. Checking the matching degree between the scene and content confirmed that "food introduction" perfectly matched the finished product display, "cooking techniques" perfectly matched the pasta cooking process, and "cultural background" perfectly matched the cultural presentation. There were no conflicts, so the two sub-videos were directly merged. Multiple audio videos: 00:00:10: Play the "food introduction" audio segment, showcasing the finished pasta; 00:00:41: Play the "cooking techniques" audio segment, showcasing the pasta cooking process; 00:01:31: Play the "cultural background" audio segment, showcasing Italian food culture.

[0132] The system realizes the complete construction process from playback node to final multi-audio video. This step ensures the best matching of audio and video through dual verification of scene-driven and content-driven approaches. In practical applications, this mechanism can significantly improve the multilingual adaptation quality of short videos, ensuring that global users have a smooth and natural viewing experience.

[0133] refer to Figure 6 In step S15, the specific steps are as follows:

[0134] S151: Collect multiple voice and video recordings, determine voice replacement events based on the multiple voice and video recordings and the user's voice requirements information, and determine multiple voice replacement items based on the parsing of the voice replacement events. Each voice replacement item contains corresponding sub-voice requirements.

[0135] S152: In each voice replacement project, determine the corresponding video decomposition mode based on the sub-voice requirements and multiple voice and video corresponding to the voice replacement project, and trigger the decomposition of multiple voice and video to output multiple voice scene areas.

[0136] S153: Real-time monitoring of changes in multiple speech scene areas, marking speech content to be translated for further adjustment of the speech content to be translated, determining new speech segments based on the changed speech scene and the speech content to be translated, and updating multiple speech and video.

[0137] In the embodiments of this application, multiple voice and video recordings are collected, and voice replacement events are determined based on the multiple voice and video recordings and the user's voice demand information. Multiple voice replacement items are determined based on the parsing of the voice replacement events. Each voice replacement item includes corresponding sub-voice demands, which is compatible with the overall consideration of the parsing of voice replacement events and ensures the accuracy of multiple voice replacement items.

[0138] At this point, the system retrieves the generated multi-speech video from S143. Such videos typically contain the following characteristics: multiple language speech segments (such as Chinese, English, Japanese, French, etc.), each speech segment is labeled with the corresponding country and scene; the video and speech content have been initially synchronized; the video format is common formats such as MP4, AVI, and MOV; the speech encoding is high-quality audio encoding such as AAC and PCM; and the metadata includes speech segment timestamps, language tags, scene tags, etc.

[0139] The system receives and analyzes the user's voice request information, compares it with the current multi-video audio content, identifies the audio parts that need to be replaced, and generates a "voice replacement event." The user's voice request information includes: language preference (e.g., wanting to change certain segments from English to Spanish); scene adaptation (e.g., wanting to change the audio from a "formal occasion" style to a "relaxed everyday" style); and content adjustment (e.g., wanting to add extra narration for a certain scene or delete redundant content). The system uses Natural Language Processing (NLP) to perform semantic understanding of the user input, compares the differences between the current video audio and the user's requests, and generates a voice replacement event based on the comparison results. This event includes: the time range for replacement; the type of language to be replaced; and the style or content requirements for replacement.

[0140] The voice replacement event is broken down into multiple specific "voice replacement projects". Each project contains clear sub-voice requirements to facilitate subsequent processing. Project decomposition: decomposed by time sequence, decomposed by language type, and decomposed by scene style. Sub-voice requirements include: target language, voice style (formal, conversational, lively, etc.), scene type (teaching, entertainment, advertising, etc.), and content requirements (addition, deletion, rewriting, etc.).

[0141] The system achieves precise conversion from user needs to specific replacement items. This step, through needs analysis, event recognition, and project decomposition, ensures that subsequent video updates can be carried out accurately and efficiently. In practical applications, this mechanism can significantly improve the personalized adaptation capability of multilingual videos, enabling content creators to quickly adjust the audio content of videos based on user feedback and provide a better user experience.

[0142] Furthermore, in each voice replacement project, the corresponding video decomposition mode is determined based on the sub-voice requirements and multiple voice and video segments corresponding to the voice replacement project, and the decomposition of multiple voice and video segments is triggered to output multiple voice scene regions. This approach is compatible with the overall consideration of the sub-voice requirements and multiple voice and video segments corresponding to the voice replacement project, ensuring the accuracy of the corresponding video decomposition mode.

[0143] At this point, the system analyzes the sub-speech requirements of each of the multiple voice replacement projects generated in S151, and determines the video decomposition mode corresponding to each project by combining the structural characteristics of multiple voice and video. The video decomposition modes include: time slice mode: the video is cut according to time points, which is suitable for accurately replacing the voice content within a certain period of time; scene slice mode: the video is cut according to scene content, which is suitable for adjusting the voice according to semantics or picture content; and hybrid slice mode: the cutting method combines time and scene, which is suitable for complex video structures.

[0144] The system decomposes multi-audio video according to a defined decomposition pattern, generating multiple audio scene regions. Each region contains: a video segment (image + audio), corresponding audio scene tags, time range markers, and the original audio content (for subsequent replacement). Simultaneously, the video decomposition tools include FFmpeg, OpenCV, and a custom video processing engine. Decomposition accuracy is either frame-level or second-level, depending on the requirements. Output format: each region is saved as an independent video segment or marked as a virtual region (without actual file splitting). The system achieves precise conversion from audio replacement projects to specific video decomposition. This step, through the selection of decomposition patterns and precise video segmentation, ensures that subsequent audio replacement work is performed in the correct regions. In practical applications, this mechanism can significantly improve the editing accuracy and efficiency of multilingual videos, enabling content creators to flexibly respond to various user needs and provide a higher quality personalized experience.

[0145] Therefore, by monitoring the changes in multiple speech scene regions in real time and marking the speech content to be translated, further adjustments can be made to the speech content to be translated. Based on the changed speech scene and the speech content to be translated, new speech segments are determined to update multiple speech and video. This approach takes into account both the changed speech scene and the speech content to be translated, ensuring the accuracy of the new speech segments. At the same time, the introduction of speech scene region changes enables the overall consideration of the changed speech scene and the speech content to be translated, ensuring the accuracy of the new speech segments and achieving the updating of multiple speech and video.

[0146] At this point, the system monitors multiple voice scene regions decomposed in S152 in real time to detect whether a voice replacement requirement has occurred. When a user or the system triggers a replacement event, the system marks the voice content that needs to be translated. Real-time monitoring is performed based on timestamps or scene changes, marking the voice segments that need to be translated in the voice scene region, including: time range, original voice content, target language, and voice style requirements. Monitoring trigger conditions: user-initiated trigger (such as selecting a new language or style), system-initiated trigger (such as detecting substandard voice quality or content mismatch).

[0147] The system further adjusts the marked audio content to ensure that the translation results meet user needs and video scene requirements. Adjustments include: language style adjustment (e.g., formal, colloquial, humorous), speech rate adjustment (e.g., speeding up, slowing down), speech emotion adjustment (e.g., enthusiastic, calm, excited), and technical terminology processing (e.g., ensuring accurate translation of technical terms). Technical details include: using NLP technology for style transfer, using TTS technology for speech synthesis, and using speech emotion recognition technology for emotion matching.

[0148] The system generates new audio segments based on the adjusted translation content and replaces them in the corresponding audio scene area, ultimately updating multiple audio and video. It uses TTS technology to generate new audio segments, ensuring synchronization between the new audio segments and the video footage. The system also performs quality assessments on the new audio segments to ensure clarity and naturalness. The update process includes: generating new audio segments, performing time synchronization processing, replacing the original audio segments, performing quality checks, and updating multiple audio and video.

[0149] The system achieves a complete closed loop from voice monitoring to video updates. This step, through real-time monitoring, content adjustment, and voice updates, ensures that multilingual videos can dynamically respond to user needs and maintain a high-quality viewing experience. In practical applications, this mechanism can significantly improve the flexibility and real-time nature of multilingual adaptation for short videos, enabling content creators to quickly respond to user feedback and provide a better personalized experience.

[0150] Please see Figure 7 , Figure 7 This is a schematic diagram of the structural composition of a short video-based multilingual intelligent translation system according to an embodiment of the present invention; the short video-based multilingual intelligent translation system includes:

[0151] The voice module 21 is used to determine multiple voice data based on the voice detection of the short video, and to determine the corresponding voice content and voice scene based on the recognition of the multiple voice data.

[0152] The speech category module 22 is used to construct scene combinations based on various speech scenarios, determine the corresponding country name based on the recognition of the scene combination, and match the corresponding speech category based on the country name.

[0153] The country name module 23 is used to determine the intelligent translation mode of the speech content based on multiple speech data and the speech type to be translated. The intelligent translation mode triggers the intelligent translation of each speech content and outputs the corresponding translated speech segment, which is marked with the corresponding country name.

[0154] The multi-speech and video module 24 is used to acquire multiple translated speech segments, determine the playback nodes of multiple translated speech segments based on the matching of multiple translated speech segments and short videos, and determine the multi-speech and video based on the playback nodes of multiple translated speech segments, the corresponding speech scenes and the video content of the short videos.

[0155] The update module 25 is used to decompose multiple audio and video into multiple audio scene regions if multiple audio and video need to be replaced. The audio content to be translated is adjusted according to the change of audio scene regions, and a new audio segment is determined according to the changed audio scene and the audio content to be translated, so as to update the multiple audio and video.

[0156] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

Claims

1. A short video-based multilingual intelligent translation method, characterized in that, include: Multiple voice data are determined based on the voice detection of short videos, and the corresponding voice content and voice scene are determined based on the recognition of multiple voice data. The process involves constructing scene combinations based on various speech scenarios, determining the corresponding country name based on the recognition of these scene combinations, and matching the corresponding speech type based on the country name. This includes: collecting various speech scenarios; determining corresponding scene elements based on the detection of each speech scenario; determining a first combination of speech scenarios based on multiple scene elements and the presentation position of each speech scenario; determining a second combination of speech scenarios based on the scene content and scene elements of each speech scenario; constructing scene combinations based on the first and second combinations of speech scenarios; determining multiple complete sub-scene combinations based on the division of scene combinations, with each sub-scene combination presenting corresponding national attractions or national marker features; determining multiple national feature elements based on the recognition of multiple complete sub-scene combinations; determining the corresponding country name based on multiple national feature elements, the shooting IP of the short video, and video information; and determining the corresponding speech type based on the matching of the country name and the speech database, with each country having a corresponding speech type. The intelligent translation mode is determined based on multiple voice data and the types of voices to be translated. This intelligent translation mode triggers the intelligent translation of each voice content and outputs the corresponding translated voice segments, which are marked with the corresponding country names. The process involves: collecting multiple translated speech segments; determining playback nodes for the translated speech segments based on matching them with short videos; determining multiple speech-video segments based on the playback nodes, corresponding speech scenes, and video content of the short videos; collecting multiple translated speech segments; determining the corresponding data volume based on the detection of the translated speech segments; marking two replacement nodes for each translated speech segment, designated as the starting replacement node and the initial replacement node; simultaneously determining corresponding replacement data segments based on matching the short videos with the translated speech segments; determining corresponding replacement regions based on the replacement data segments, the two replacement nodes of the translated speech segments, and the data volume of the translated speech segments; determining corresponding playback nodes based on the detection of these replacement regions; determining the playback nodes for multiple translated speech segments; determining a first sub-speech-video segment based on the playback nodes and corresponding speech scenes; determining a second sub-speech-video segment based on the playback nodes and video content of the short videos; and determining multiple speech-video segments based on the first and second sub-speech-video segments. The first sub-audio video is a sub-video constructed primarily based on the audio scene; the second sub-audio video is a sub-video constructed primarily based on the video content; the multiple audio videos contain the following features: audio segments, video format, audio encoding, and metadata; comparing the playback nodes of the first and second sub-audio videos, they are found to be completely consistent; checking the matching degree between the scene and the content, there is no conflict, so the first and second sub-audio videos are directly merged. If multiple audio and video segments need to be replaced, they are broken down into multiple audio scene regions. The audio content to be translated is adjusted according to the change of the audio scene region, and a new audio segment is determined based on the changed audio scene and the audio content to be translated, so as to update the multiple audio and video segments. 2.The short video based multilingual intelligent translation method according to claim 1, characterized in that, The process of determining multiple voice data points based on voice detection in short videos, and determining the corresponding voice content and voice scene based on the recognition of these multiple voice data points, includes: In short video platforms, the user's selected favorite videos are collected, and the corresponding short videos are determined based on the detection of these favorite videos. Based on the short videos, the corresponding favorite tags, and the user's viewing history on the short video platform, multiple key voice data are determined. At this time, multiple key voice data are presented in the same short video, and the corresponding key factors are marked. The recognition pattern of the voice data is determined based on the data location and key factors of multiple key voice data. The autonomous recognition of multiple voice data is triggered based on the recognition pattern, and multiple voice features are collected. The corresponding voice content and voice scene are determined based on the selection of multiple voice features. 3.The short video based multilingual intelligent translation method according to claim 1, characterized in that, The intelligent translation mode, which determines the speech content based on multiple speech data and the type of speech to be translated, triggers intelligent translation of each speech content and outputs the corresponding translated speech segments. These translated speech segments are labeled with the corresponding country names, including: Multiple speech data points are collected. Within each speech data point, the basic speech type is determined based on speech recognition. Then, the corresponding matching criteria are determined by matching the basic speech type of the speech data with the speech type to be translated. Matching coefficients are used to determine the intelligent translation mode of the speech content based on the matching coefficients, the type of speech to be translated, and the corresponding speech data. 4.The short video-based multilingual intelligent translation method of claim 3, wherein, The intelligent translation mode, which determines the speech content based on multiple speech data and the type of speech to be translated, triggers intelligent translation of each speech content and outputs the corresponding translated speech segment. The translated speech segment is marked with the corresponding country name, and also includes: In the intelligent translation mode of speech content, multiple translation paths are determined based on the recognition of the intelligent translation mode of speech content, and the optimal translation path is determined based on the multiple translation paths, the speech content of the speech data, and the speech scene. The system triggers autonomous translation of speech data based on the optimal translation path and performs intelligent translation of each speech content to output the corresponding translated speech segments. The translated speech segments are marked with the corresponding country name and the corresponding speech scene. 5.The short video based multilingual intelligent translation method according to claim 1, characterized in that, If multiple audio-visual videos need to be replaced, the multiple audio-visual videos are decomposed into multiple audio scene regions. The audio content to be translated is adjusted according to the change of audio scene regions, and new audio segments are determined based on the changed audio scene and the audio content to be translated to update the multiple audio-visual videos, including: Collect multiple audio and video recordings, determine audio replacement events based on the multiple audio and video recordings and the user's audio requirements, and determine multiple audio replacement items based on the parsing of the audio replacement events. Each audio replacement item contains corresponding sub-audio requirements. 6.The short video based multi-language intelligent translation method according to claim 5, characterized in that, If multiple audio-visual videos need to be replaced, the multiple audio-visual videos are decomposed into multiple audio scene regions. The audio content to be translated is adjusted according to the change of audio scene regions. New audio segments are determined based on the changed audio scene and the audio content to be translated to update the multiple audio-visual videos. The method also includes: In each voice replacement project, the specific requirements for voice replacement are determined based on the sub-voice needs and multiple voice / video inputs required for that project. The corresponding video decomposition mode is used to trigger the decomposition of multiple audio and video to output multiple audio scene regions. The system monitors the changes in multiple speech scene areas in real time and marks the speech content to be translated for further adjustment. Based on the changed speech scene and the speech content to be translated, new speech segments are determined to update the multi-speech video.

7. A short video based multilingual intelligent translation system, characterized in that, The short video-based multilingual intelligent translation system is applied to the short video-based multilingual intelligent translation method as described in any one of claims 1-6, wherein the short video-based multilingual intelligent translation system comprises: The voice module is used to determine multiple voice data based on the voice detection of short videos, and to determine the corresponding voice content and voice scene based on the recognition of multiple voice data. The speech category module is used to construct scene combinations based on various speech scenarios, determine the corresponding country name based on the recognition of the scene combination, and match the corresponding speech category based on the country name. The country name module is used for intelligent translation to determine the speech content based on multiple speech data and the type of speech to be translated. This intelligent translation mode triggers the intelligent translation of each piece of audio content and outputs the corresponding translated audio segments, which are marked with the corresponding country names. The multi-audio-video module is used to acquire multiple translated audio segments, determine the playback nodes of multiple translated audio segments based on the matching of multiple translated audio segments and short videos, and determine the multi-audio-video based on the playback nodes of multiple translated audio segments, the corresponding audio scenes, and the video content of the short videos. The update module is used to break down multiple audio and video segments into multiple audio scene regions if multiple audio and video segments need to be replaced. It adjusts the audio content to be translated according to the change of audio scene regions, and determines new audio segments based on the changed audio scene and the audio content to be translated, so as to update the multiple audio and video segments.