A dialogue processing method, device, storage medium and electronic equipment
By constructing dialogue director charts and performance charts, the problem of insufficient realism in multi-person dialogue scenarios using multilingual dubbing methods was solved, achieving natural interaction and accurate control of rhythm in dialogue, thus enhancing the realism of multilingual dubbing and the audience experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD
- Filing Date
- 2026-05-21
- Publication Date
- 2026-07-31
AI Technical Summary
Existing multilingual dubbing methods lack realism in multi-person dialogue scenarios, resulting in distorted dialogue. This is especially true in film and television-level multi-person dialogue scenarios, where real-life interactions such as interrupting, talking over, responding, and delayed reactions are forcibly eliminated as noise or errors. Furthermore, differences in sentence length and speaking speed habits between different languages amplify rhythmic differences.
By extracting multi-person interaction clues from the audio, visuals, and subtitles of the original film, a structured dialogue director's score is constructed, the overall dialogue scheduling is solved, a performance score is generated, and the mixing control is performed according to the type of interaction to ensure the realism of the dialogue and the accuracy of the rhythm.
It enhances the realism of multilingual dubbing, maintains the natural interaction and emotional progression of multi-person dialogues, avoids dialogue distortion, and enhances the audience's immersive experience.
Smart Images

Figure CN122224141B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of audio and video processing technology, and more specifically, to a dialogue processing method, apparatus, storage medium, and electronic device. Background Technology
[0002] Current multilingual dubbing methods typically use subtitle segments as the smallest unit, following a process of "sentence-by-sentence translation → sentence-by-sentence text-to-speech (TTS) → splicing or simple mixing according to time windows".
[0003] Existing multilingual dubbing methods are acceptable in single-person narration or sparse dialogue scenarios, but they have the following shortcomings in film-level multi-person dialogue scenarios: 1) Multi-person dialogues are "read from script in turn". Since real dialogues often include interactions such as interrupting, interrupting, responding, delayed reactions, and laughing together, sentence-by-sentence synthesis often treats these interactions as noise or errors and forcibly eliminates overlap, resulting in a lack of tension and realism in the dialogue; 2) Due to the large differences in sentence length and speaking speed habits between different languages, if only the subtitle time window is used as a hard boundary, it often happens that "the original film (i.e., the original material without post-processing) is arguing, but the target language becomes speaking in turn", which amplifies the differences in rhythm between languages and leads to the distortion of the dialogue.
[0004] Therefore, how to improve the realism of multilingual dubbing is a problem that this application urgently needs to solve. Summary of the Invention
[0005] In view of this, this application discloses a dialogue processing method, apparatus, storage medium, and electronic device, aiming to improve the realism of multilingual dubbing.
[0006] To achieve the above objectives, the disclosed technical solution is as follows:
[0007] The first aspect of this application discloses a dialogue processing method, the method comprising:
[0008] The obtained original footage is aggregated to obtain multiple dialogue fragments;
[0009] Extract a set of interactive cues from each dialogue segment;
[0010] The dialogue director spectrum corresponding to the original film is constructed based on the set of interactive clues; wherein, the dialogue director spectrum corresponding to the original film is a structured intermediate representation;
[0011] Under the constraints of the dialogue director's score, the overall dialogue scheduling is solved based on the dialogue director's score corresponding to the original film, and the performance score is generated.
[0012] Based on the interaction types in the performance spectrum and interaction clue set, the audio of each character in the original film is mixed and controlled, and the mixing control result is output.
[0013] A second aspect of this application discloses a dialogue processing apparatus, the apparatus comprising:
[0014] The aggregation unit is used to aggregate the acquired original clips to obtain multiple dialogue segments;
[0015] Extraction unit, used to extract a set of interaction cues for multi-person interaction from each dialogue segment;
[0016] A construction unit is used to construct the dialogue director spectrum corresponding to the original film based on the set of interactive clues; wherein the dialogue director spectrum corresponding to the original film is a structured intermediate representation;
[0017] The scheduling and solving unit is used to perform overall scheduling and solving of dialogue based on the dialogue director's spectrum corresponding to the original film under the constraints of the dialogue director's spectrum, and to generate the performance spectrum.
[0018] The control output unit is used to control the mixing of audio for each character in the original film according to the interaction type in the performance spectrum and the set of interaction clues, and output the mixing control result.
[0019] A third aspect of this application discloses a storage medium comprising stored instructions, wherein, when the instructions are executed, the device in which the storage medium resides executes the dialogue processing method as described in any one of the first aspects.
[0020] The fourth method of this application discloses an electronic device, including a memory and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described in any one of the first aspects of the dialogue processing method.
[0021] As can be seen from the above technical solution, this application discloses a dialogue processing method, apparatus, storage medium, and electronic device. The method aggregates the acquired original footage to obtain multiple dialogue segments. From each dialogue segment, an interaction clue set for multi-person interaction is extracted. Based on the interaction clue set, a dialogue director's score corresponding to the original footage is constructed, where the dialogue director's score is a structured intermediate representation. Under the constraints of the dialogue director's score, an overall dialogue scheduling solution is performed based on the dialogue director's score corresponding to the original footage to generate a performance score. Based on the performance score and the interaction types in the interaction clue set, the audio of each character in the original footage is mixed and controlled, and the mixing control result is output. Through the above solution, dialogue interaction clues are extracted from the audio / visual / subtitle information of the original footage to generate a reproducible and reusable dialogue director's score, transforming the realism of multi-person dialogue from subjective aesthetics into a computable and reusable structured object. Under the constraints of the dialogue director's score, an overall dialogue scheduling solution is performed based on the dialogue director's score corresponding to the original footage. While ensuring intelligibility and legal boundaries, a certain range of overlap and delay is allowed and precisely controlled, making the dubbing more like a real group dialogue rather than a turn-reading of a script. Furthermore, by solidifying the dialogue relationships and rhythm constraints into the dialogue director's score, and then making feasible scheduling and mixing in the target language, the dialogue rounds and emotional progression rhythm of the original film are maintained, thereby avoiding dialogue distortion in multilingual dubbing and improving the realism of multilingual dubbing. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0023] Figure 1 This is a schematic flowchart of a dialogue processing method disclosed in an embodiment of this application;
[0024] Figure 2 This is an architecture diagram of the dialogue processing system disclosed in the embodiments of this application;
[0025] Figure 3 This is a schematic diagram of the structure of the dialogue director spectrum disclosed in the embodiments of this application;
[0026] Figure 4 This is a schematic diagram of the performance spectrum solving and degradation link disclosed in the embodiments of this application;
[0027] Figure 5 This is a schematic diagram of the mixing control disclosed in the embodiments of this application;
[0028] Figure 6This is a schematic diagram of the structure of a dialogue processing device disclosed in an embodiment of this application;
[0029] Figure 7 This is a schematic diagram of the structure of the electronic device disclosed in the embodiments of this application. Detailed Implementation
[0030] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0032] As the background technology indicates, multilingual dubbing typically uses subtitle segments as the smallest unit, following a process of "sentence-by-sentence translation → sentence-by-sentence TTS → splicing or simple mixing according to time windows." In multi-person dialogue scenarios, existing multilingual dubbing methods often suffer from a lack of tension and realism. Real conversations frequently involve interruptions, interruptions, responses, delayed reactions, and shared laughter, but sentence-by-sentence synthesis often treats these interactions as noise or errors, forcibly eliminating overlaps. Furthermore, due to significant differences in sentence length and speaking speed across languages, using only subtitle time windows as hard boundaries often results in situations where "the original footage (the raw material without post-processing) is arguing, but the target language is presented as a series of speeches," amplifying cross-language rhythm differences and distorting the dialogue.
[0033] To address the aforementioned issues, this application discloses a dialogue processing method, apparatus, storage medium, and electronic device. It extracts dialogue interaction cues from the audio / video / subtitle information of the original film, generating a reproducible and reusable dialogue director's score. This transforms the realism of multi-person dialogue from subjective aesthetics into a calculable and reusable structured object. Under the constraints of the dialogue director's score, the overall dialogue scheduling is solved based on the dialogue director's score corresponding to the original film. While ensuring intelligibility and legal boundaries, it allows and precisely controls a certain range of overlap and delay, making the dubbing more like a real group dialogue rather than a series of script readings. Furthermore, by solidifying dialogue relationships and rhythm constraints into the dialogue director's score, and then performing feasible scheduling and mixing in the target language, the dialogue turns and emotional progression rhythm of the original film are maintained, thereby avoiding dialogue distortion in multilingual dubbing and improving the realism of multilingual dubbing. Specific implementation methods are described in detail in the following embodiments.
[0034] It should be noted that the dialogue processing method, apparatus, storage medium, and electronic device provided in this application relate to technical fields such as audio and video processing, artificial intelligence, audio and video processing, text-to-speech, multilingual content production and post-production mixing control, and are applicable to scenarios such as dramas, variety shows, and micro-dramas on Internet video platforms. The above are merely examples and do not limit the application field of the dialogue processing method, apparatus, storage medium, and electronic device provided in this application.
[0035] refer to Figure 1 The diagram shown is a flowchart illustrating a dialogue processing method disclosed in an embodiment of this application. The dialogue processing method mainly includes the following steps:
[0036] S101: Aggregate the obtained original footage to obtain multiple dialogue fragments.
[0037] In S101, the acquired original footage is forcibly aligned to obtain each minimum text segment (such as the minimum dialogue segment, i.e., the subtitle segment), and each minimum text segment is used as an initial unit; according to the temporal adjacency relationship, scene continuity relationship and semantic continuity relationship, each initial unit is aggregated into multiple dialogue segments.
[0038] It should be noted that the original footage refers to the raw material that has not undergone post-processing. The original footage includes the original video or audio, source language subtitles (or Automatic Speech Recognition (ASR) subtitles, etc.), target language information, optional speaker recognition results or role annotation information, etc.
[0039] The system aggregates subtitle segments according to temporal continuity (e.g., by adjacent intervals, shot boundaries, scene transitions, etc.) into several dialogue segments. Each dialogue segment contains multiple dialogue segments, which are also known as dialogue segment nodes. The overall system architecture is as follows: Figure 2 As shown. Figure 2 The document demonstrates the relationship between the modules of dialogue segmentation, interactive clue extraction, director's score construction, performance score solving, mixing control, and output tracking.
[0040] A dialogue segment is a range of dialogues aggregated according to temporal continuity. A dialogue segment contains multiple lines of dialogue and multiple speakers, and usually corresponds to continuous communication in the same context.
[0041] Dialogue segment nodes are the smallest units of dialogue within a dialogue segment. They can correspond to subtitle segments / ASR segments or their combined semantic segments, and record the speaker, time window, and text content.
[0042] To better understand the process of aggregating the acquired original footage to obtain multiple dialogue segments, an example is provided below:
[0043] For example, first use the subtitle / ASR segment as the initial unit, then determine whether adjacent segments belong to the same dialogue segment. Typically, at least two of the following conditions must be met:
[0044] 1. The interval between sentences is less than the continuous dialogue threshold;
[0045] 2. Although there are cuts, they are either reverse shots or continuous shots of the same scene;
[0046] 3. The set of speakers exhibits a continuity relationship;
[0047] 4. There is a clear sequential relationship in the text or semantics.
[0048] If the following situations occur, it is preferable to cut into a new dialogue segment:
[0049] 1. The spacing between sentences is too large;
[0050] 2. Scene / location switching;
[0051] 3. The original characters exit and enter another event flow;
[0052] 4. Insert longer pieces of music, empty shots, or explanatory narration in the middle.
[0053] For example: if character B immediately responds 150ms after character A finishes speaking, even if the scene is switched to a reverse shot, the dialogue will still be grouped into the same segment; if character A finishes speaking and there is a 4-second empty shot and background music, and then the dialogue of character C in another scene is introduced, it is preferable to switch to a new dialogue segment; in variety show group chats, multiple short phrases such as "yes, yes, yes," "that's right," and "haha" should usually be grouped into the same segment with the main dialogue to facilitate the overall solution later.
[0054] In one scenario implementation, the system first uses a subtitle segment, an ASR segment, or the smallest dialogue segment obtained by forced alignment as the initial unit (seg_i). Each initial unit has at least the following fields: speaker_id / role_id, t_in, t_out, shot_id, and text_id. Then, based on temporal adjacency, scene continuity, and semantic continuity, these initial units are aggregated into dialogue fragments (dialogue_k).
[0055] Specifically, it can be done in the following way:
[0056] 1. First, generate adjacency candidates. For two temporally adjacent or near-adjacent dialogue segments seg_i and seg_j, calculate the interval between sentences according to the formula gap(i,j) = t_in(j) - t_out(i);
[0057] Where gap(i,j) is the time interval between adjacent dialogue segments seg_i and seg_j; t_out(i) is the end time of the i-th segment; and t_in(j) is the start time of the j-th segment.
[0058] At the same time, auxiliary features such as whether there is a camera cut, whether the character continues, and whether there is a question-and-answer or dialogue semantics are extracted.
[0059] 2. Next, determine whether to retain them within the same dialogue segment. A segment can continue to be grouped into the same dialogue_k if at least two of the following conditions are met:
[0060] (1) gap(i,j) is less than the preset continuous dialogue threshold (τ_join);
[0061] (2) Although there is a camera switch, it is a reverse shot, a close-up switch, or a continuous shot of the same scene;
[0062] (3) The speaker set exhibits a continuity relationship;
[0063] (4) There are obvious connecting clues in the text or audio.
[0064] 3. A new dialogue segment can be directly started when any of the following conditions are met:
[0065] (1) gap(i,j) is greater than the break threshold (τ_break);
[0066] (2) Scene changes, location changes, obvious narrative jumps, or transitions to narration or subtitles occur;
[0067] (3) The original group of dialogue characters exits as a whole and is transferred to another group of characters or another event flow;
[0068] (4) Insert longer music or ambient sound into the audio to bridge the previous round of interaction, and the subsequent dialogue does not continue from the previous round of interaction.
[0069] 4. Make local corrections to the preliminary aggregation results. If a middle short phrase satisfies the continuity condition with both the preceding and following ends, it can be kept in the same dialogue_k to avoid accidentally cutting short phrases such as "yes, yes, yes," "wait a minute," and "that's it" into independent segments.
[0070] Example 1: Character A speaks between 10.00 and 11.20 seconds, and Character B responds immediately between 11.35 and 12.10 seconds. Although the shots are switched to reverse shots, the gap is 150ms, and the characters are consecutive, so they still merge into the same dialogue segment.
[0071] Example 2: Character A speaks between 32.00 and 33.00, followed by a 4-second empty shot and background music, and then at 37.20, the dialogue of character C in another scene is cut. Even if the text is still on the same theme, it is preferable to cut it into a new dialogue segment.
[0072] Example 3: In a variety show group chat, the host, guest A, and guest B talk and laugh together multiple times in a short period of time. The system can group multiple short sentences into the same dialogue_k for subsequent overall solution.
[0073] S102: Extract a set of interactive cues for multi-person interactions from each dialogue segment.
[0074] In S102, for each dialogue segment, a set of interaction cues for determining the interaction relationship can be extracted from audio, text, and / or video information using a multimodal interaction cue extraction method. This multimodal interaction cue extraction method is used to record the interaction relationships in multi-person dialogues and support a structured intermediate representation for subsequent scheduling and mixing control.
[0075] Specifically, audio cues, text cues, and / or visual cues are extracted from each dialogue segment, and an interactive cue set is formed based on the audio cues, text cues, and / or visual cues in each dialogue segment.
[0076] The set of interactive clues includes audio clues, visual clues, and text clues.
[0077] Audio cues include Voice Activity Detection (VAD), the start and end of adjacent speech, overlapping intervals; energy changes and spikes (often appearing as "suddenly starting to speak / stealing the end" when interrupted); and non-word content events such as laughter, inhalation, and sighing (used to identify interactions such as laughing together / breathing together).
[0078] Visual cues (optional) include whether the speaker is in the frame, whether they are obscured, whether it is a shot with a reverse shot, and other camera language such as multiple people in the same frame / single close-up, used to distinguish between the "leader of the conversation" and the "background echo".
[0079] Textual cues include punctuation and interjections (such as "ah," "yes," etc.) to prompt responses / agreements; repetition and emphasis and transition words to indicate emotional progression and motivations for interrupting.
[0080] S103: Construct the dialogue director spectrum corresponding to the original film based on the set of interactive clues; where the dialogue director spectrum corresponding to the original film is a structured intermediate representation.
[0081] Specifically, the dialogue director spectrum is a structured intermediate representation consisting of multiple dialogue segment nodes, interaction edges, and control parameters corresponding to the interaction edges of multiple dialogue fragments.
[0082] In S103, a Dialogue Directing Score (DDS) is constructed based on dialogue segment nodes. The Dialogue Directing Score is a structured "dialogue rhythm instruction" composed of dialogue segment nodes and interaction edges. It is used to express the turn relationship, interaction intensity and controllable range of the original dialogue. It can be stored in a library, versioned, and reused across languages.
[0083] The dialogue director's score can be reused across languages. For example, the "interaction relationship" of the same dialogue remains unchanged, and the target language only needs to choose appropriate lines and dubbing rhythm under the constraints of the director's score, thereby avoiding the disjointedness of "the plot is still arguing but the dubbing sounds like they are taking turns speaking" in cross-language versions.
[0084] An interaction edge represents the interaction relationship between two dialogue segment nodes, including the interaction type (such as interrupting, responding, reacting, echoing and laughing, following and echoing, etc.) and the controllable parameters corresponding to each edge (such as the allowed overlap range, the reaction delay range, the dominance weight, etc.). Among them, the dominance weight is used to describe the weight parameter of "who should be heard more clearly / who dominates the narrative more" in the interaction, and participates in scheduling and mixing control (e.g., the dominant one takes priority when overlapping).
[0085] The "turn relationships and interaction intensity" of a dialogue are abstracted into a director spectrum that can be stored in a database and is versionable. The dialogue director spectrum includes at least a set of nodes, a set of edges, interaction types, parameterization rules, credibility, and interaction signatures.
[0086] Node set: Each node corresponds to a dialogue segment, recording its source time window, speaker / role (if available), text content ID, etc.
[0087] Edge set: An edge is established between any two nodes with an interactive relationship, and the interaction type and parameters are labeled. The interaction types include at least the following: Interrupt, Follow, React, CoEvent, and Backchannel.
[0088] Break the edge: The following sentence enters before the previous sentence ends, forming a controllable overlap.
[0089] The next sentence is delivered quickly after the previous sentence, demonstrating a tight and concise response.
[0090] React edge: There is a perceptible delay in the response to the following sentence (such as a pause or response after inhalation).
[0091] Co-Event: Multiple characters laughing, sighing, or making a sound together.
[0092] Backchannel: Short backchannels such as "um," "yes," and "right" are inserted between main clauses.
[0093] Parameterization rules (example):
[0094] For broken edges, record the allowed overlap interval [O_min, O_max] (e.g., 80ms~250ms) and the dominant weight W_dom (which determines which is more prominent during mixing);
[0095] For the receiving side, record the response delay interval [D_min, D_max] (e.g., 0ms~180ms).
[0096] For the shared laughter edge, record the event alignment window [A_min, A_max] and whether unison continuation is allowed, etc.;
[0097] Regarding credibility: Conf_edge is calculated for each edge (weighted by the consistency and stability of audio / video / text cues, such as higher when audio side overlap and energy surge, video switching and speaker visibility, and text interjections are supported simultaneously, and lower when there is conflict or absence), which is used for the selection of strong / weak constraints in subsequent scheduling;
[0098] Interactive signature: For each dialogue segment, an interactive signature is generated based on the director's spectrum and the time relationship of the original film (such as concurrency curve, dominant weight order, overlapping envelope, reaction delay distribution, etc.) as a reference baseline for cross-language preservation and auditing, and is used for cross-language preservation and self-checking.
[0099] The structure of the dialogue director spectrum is as follows Figure 3 As shown.
[0100] Figure 3In this system, each node corresponds to a dialogue segment, recording at least the character, source language time window, and text content; edges between nodes represent interaction relationships. Different edges can correspond to different interaction types, such as interruption edges, response edges, agreement edges, and co-event edges. Edges can also carry parameters, such as the allowed overlap range and dominant weight of interruption edges, the delay range of response edges, and the alignment window of co-event edges. Subsequent performance spectrum solving and mixing control can directly read these nodes, edges, and their parameters, instead of simply treating the entire dialogue as a series of unrelated sentences.
[0101] Figure 3 In this context, each node corresponds to a dialogue segment; the edges between nodes represent the interaction relationship; the edges carry both the interaction type and parameters; for example, the interruption edge records the overlap window and the dominant weight, the continuation edge records the delay window, and the co-event edge records the alignment window; this means that the subsequent solution reads not only "who continues and who speaks", but also "how the conversation is allowed to continue, how much overlap is allowed, and who should be heard more clearly". Figure 3 This embodies the core intermediate assets of this plan.
[0102] The specific process of constructing a dialogue director spectrum based on the set of interactive clues is shown in A1-A4.
[0103] A1: Combine audio cues, text cues, and / or visual cues from the set of interactive cues into an interactive type.
[0104] The interaction types include strong / weak interruptions, immediate follow-up, hesitant responses, shared laughter / breathing, etc.
[0105] A2: Get the interaction intensity parameter of the interaction type.
[0106] The interaction intensity parameter represents the importance of a particular interaction edge to the audience's perception. It is preferably derived from a fusion of the following evidence: audio evidence such as overlap duration or response delay, sudden increases in volume at the beginning of the following sentence, and the end of the preceding sentence being suppressed; visual evidence such as whether the camera turns to the new speaker and who is in the dominant position on screen; textual / event evidence such as interjections, repeated words, laughter, and agreeing words; and whether there is a significant shift in dominance. Therefore, this scheme explicitly constructs the interaction relationships between these nodes into a graph structure.
[0107] A3: Determine the round relationship between multiple dialogue segment nodes on the timeline.
[0108] Among them, the round relationship refers to the order, entry, continuation, waiting and yielding relationship between multiple dialogue segment nodes on the timeline. The round relationship includes who speaks first, who is interrupted by whom, the response delay between two sentences, whether there is concurrent speaking, and whether the dominance is transferred between nodes.
[0109] A4: Construct the dialogue director spectrum corresponding to the original film based on the round relationship, interaction type, and interaction intensity parameters.
[0110] The specific process of multimodal interaction clue extraction and interaction type determination (automatically accumulating "dialogue dynamics") is as follows: Extract speaker start and end, overlapping intervals, energy surges (often accompanied by interruptions), laughter / inhalation, and other events from the audio side, as well as signs of "the previous sentence being suppressed / the ending being stolen"; extract (optional) information such as shot reversal, speaker visibility and occlusion, and multiple people in the same frame from the visual side to improve the credibility of interaction type determination; extract clues such as punctuation, interjections, transition words, and repetition from the subtitle / text side to correct the semantic rationality of "short sentences interrupting / interrupting"; integrate the above clues into interaction types (e.g., strong interruption / weak interruption, immediate follow-up, hesitant response, shared laughter / shared breathing, etc.) and their intensity parameters, and write them into the dialogue director's score corresponding to the original film.
[0111] To facilitate understanding of the process of abstracting turn relationships and interaction intensity into a database-compatible and versionable dialogue director spectrum, the following process is performed on each dialogue segment, using a scenario example:
[0112] 1. Node standardization. Organize dialogue segments belonging to the same dialogue_k into a node set V={v1, v2, ..., vn}, where each node includes at least node_id, speaker_id / role_id, src_text, t_in, t_out, duration, shot_id, and node_conf.
[0113] 2. Round skeleton generation. Nodes are sorted according to the time of origination and the temporal relationship between adjacent nodes is calculated. The specific calculation formula is as follows:
[0114] delay_ij=t_in(j)-t_out(i);
[0115] ovl_ij=max(0,t_out(i)-t_in(j));
[0116] Where delay_ij is the response delay from node i to node j. When delay_ij is greater than or equal to 0, it means that node j enters after node i finishes; when delay_ij is less than 0, it means that node j enters before node i finishes, indicating preemption or overlap. ovl_ij is the overlap duration.
[0117] 3. Interaction Edge Determination. Edge types are determined based on delay_ij, ovl_ij, and multimodal evidence, and may include:
[0118] (1) Interrupt the edge;
[0119] (2) Follow while responding;
[0120] (3) React edge;
[0121] (4) Attachment edge Backchannel;
[0122] (5) CoEvent.
[0123] 4. Interaction Strength Calculation. The interaction strength I_e is calculated for each edge, which can be understood as a comprehensive result of evidence such as overlap duration, sudden increase in volume, shift in dominance, textual cues, and visual dominance. A higher I_e indicates a more important interaction and a greater likelihood of it being retained as a strong interaction edge.
[0124] 5. Parameter Writing. After determining the edge type, write controllable parameters for each edge. For example:
[0125] (1) Interrupt writes allowable overlapping intervals [O_min, O_max], dominant weight W_dom;
[0126] (2) Follow the write delay interval [D_min, D_max];
[0127] (3) CoEvent writes to the alignment window [A_min, A_max].
[0128] 6. Credibility Calculation. The credibility of an edge, Conf_edge, is preferably determined by the consistency of audio, video, and text evidence. Conf_edge is higher when multiple sources of evidence are supported simultaneously; Conf_edge decreases when the evidence is singular or conflicting.
[0129] 7. Versioned Storage. Director's scores are preferably stored as versioned assets. Example fields may include:
[0130] (1) Node table: node_id, dialogue_id, speaker_id, t_in, t_out, text_ref, node_conf, dds_version;
[0131] (2) Edge table: edge_id, from_id, to_id, edge_type, intensity, params, conf_edge, evidence_refs, dds_version;
[0132] (3) Version table: dds_version, parent_version, build_rule_set, source_asset_hash, created_at, operator / model_id.
[0133] Thus, the "round relationship" corresponds to the temporal sequence and dominant handover structure between nodes, and the "interaction intensity" corresponds to the importance and retention priority of each side. Together, they form a reusable, traceable, and comparable director spectrum asset.
[0134] S104: Under the constraints of the dialogue director's score, solve the overall dialogue scheduling based on the dialogue director's score corresponding to the original film to generate the performance score.
[0135] The Dialogue Performance Plan (DPP) is the executable result obtained under the constraints of the Dialogue Director's Plane. The DPP includes at least the start time / relative offset of each line of dialogue, overlap strategies (such as allowed overlap range), suggested speaking speed / pauses, and degradation markers. The DPP serves as an explanatory label for the start time, allowed overlap range, suggested speaking speed / pauses, and necessary degradation markers (such as "why interrupt / delay is needed here").
[0136] Under the constraints of the dialogue director's spectrum, a holistic solution is performed based on the dialogue director's spectrum corresponding to the original film, outputting an executable performance spectrum. The holistic solution must at least satisfy interaction constraints, intelligibility constraints, boundary constraints, minimum modification, degradation path when infeasible, interaction signature preservation, and self-checking.
[0137] Interaction constraints: Try to satisfy the overlap / delay range of the director's spectrum edges; edges with high credibility are used as strong constraints, and edges with low credibility are used as soft constraints.
[0138] Intelligibility constraints: Limit the number of speakers who dominate the foreground at any given time and control the total concurrent load after background echoing is calculated to prevent the audience from not being able to hear clearly; echoing edges with low dominance weight can automatically reduce their presence or degenerate into sequential connections.
[0139] Boundary constraints: Each dialogue segment must not cross the boundary of its respective dialogue segment; minor adjustments may be made within the safety tolerance if necessary, but the amount of adjustment must be recorded.
[0140] Minimal changes: While satisfying the above constraints, try to maintain the relative rhythm of the source dialogue (e.g., the structure of "who speaks first, who speaks last, and who is interrupted") and reduce unnecessary overall drift.
[0141] Degradation path when infeasibility occurs: When some edge constraints cannot be satisfied, degradation occurs in the following order: "weak interaction first degradation → conformity first degradation → shared alignment relaxation → break overlap contraction → shorter candidate switching → complete sequence", and the degradation reason and the list of degraded edges are output.
[0142] Interactive signature retention and self-check: Among multiple feasible performance spectra, the scheme with the smallest difference from the source language interactive signature is selected first; when the difference exceeds the threshold, a callback is triggered to solve or enter the conservative degradation of "understandability priority", and the reason for the action is recorded; the difference can be represented by the weighted distance of indicators such as concurrency curve, overlapping envelope, and dominant order, and the threshold can be configured according to program type / dialogue style and tracked with version.
[0143] To better understand the solution that minimizes the differences in source language interactive signatures, an example is provided here:
[0144] In this scheme, interactive signatures may include concurrency curves; overlapping envelopes; response delay distribution; dominance order; and alignment of co-events such as laughter / shouts.
[0145] For any candidate performance spectrum, the system can calculate these metrics separately and compare them with the source language baseline. Preferably, the total difference can be expressed as a weighted distance, which includes D_conc: concurrency curve difference; D_ovl: overlap envelope difference; D_lat: response delay distribution difference; D_dom: dominance order reversal penalty; D_coe: co-event alignment window bias;
[0146] Finally, among all solutions that satisfy the basic constraints, the one with the smallest weighted distance is selected first. If the difference exceeds a threshold (which is set according to the actual situation, and this application does not specify a particular threshold), a callback solution is triggered or conservative degeneracy is entered.
[0147] The specific overall solution, concurrency limits, and interactive signature difference calculation process are as follows:
[0148] In one embodiment, the overall scheduling solution under the constraints of the dialogue director spectrum is no longer "sentence-by-sentence alignment," but rather the entire dialogue is treated as a joint optimization problem, solving for candidate selection, start and end arrangements, overlap, pauses, and speech rate together. This process can be summarized as a closed loop of "candidate generation—constraint layering—joint solution—acceptance—degenerate callback."
[0149] 1. First, prepare multiple target language candidates and their synthesizable attributes for each node;
[0150] 2. Then, the constraints are further layered according to strong interaction, weak interaction, and intelligibility budget (the intelligibility budget is a set of constraints on the number of concurrent speakers allowed at the same time, the allowable overlap intensity, etc., used to ensure that the audience can hear the main content clearly);
[0151] 3. Jointly search for the optimal combination of candidate options, start time, speech rate, and pauses within the feasible region;
[0152] 4. Perform understandability budget checks and interactive signature discrepancy checks on the solution results;
[0153] 5. If the result still does not meet the requirements, then solve the problem by relaxing the constraints of the degenerate link.
[0154] In this closed loop, the following quantities can be used as decision variables:
[0155] (1) y_i, c: Whether node v_i selects candidate c;
[0156] (2) x_i: the actual start time or relative offset of node v_i;
[0157] (3) r_i: the speech rate ratio of node v_i;
[0158] (4) p_i: the pause scheme for node v_i;
[0159] (5) z_e: Whether the interactive edge e is retained, reduced or degraded according to its original strength.
[0160] The constraints of the dialogue director spectrum include candidate uniqueness constraints, interaction edge constraints, node boundary constraints, speech rate and pause constraints, and intelligibility budget constraints.
[0161] The "concurrency limit" doesn't simply count how many people are speaking at the same time; it distinguishes between foreground concurrent speakers and background echoing. The concurrency limit differentiates between foreground concurrent speakers—the main speaker whom the audience must prioritize hearing—and background echoing speakers—those whose presence is sufficient but don't need to be equally clear. Therefore, the preferred approach is to calculate for each short-term analysis window: whether the number of foreground dominant speakers exceeds the scene threshold; whether the combined load of background echoing and foreground concurrency exceeds the limit; whether the relative clarity of the main dialogue in the key frequency band is sufficient; and whether the cumulative duration of strong overlap is excessive. For example, in a drama with strong conflict, two foreground dominant speakers can overlap for a short period, but three main dialogue lines cannot be allowed to compete for a long time simultaneously; in a variety show group chat, more background echoing can exist, but the main dialogue still needs to remain comprehensible.
[0162] The calculation can be performed for each short-time analysis window t, and the specific calculation formula is as follows:
[0163] N_fg(t)=Σ1(dom_u,t≥θ_fg);
[0164] N_bg(t) = Σ1(0 <dom_u,t<θ_fg);
[0165] And it requires that N_fg(t) ≤ N_fg_max(scene); N_fg(t) + λ × N_bg(t) ≤ N_total_max(scene);
[0166] Where dom_u,t represents the dominance of a character within the time window; θ_fg is the foreground judgment threshold; λ is the background congruence conversion factor; scene is the program type or dialogue style bucket; N_fg(t) is the number of foreground dominant speakers within the time window t; N_bg(t) is the number of background congruence speakers within the time window t; N_fg_max(scene) is the maximum allowed concurrent foreground dominant speakers in scene; and N_total_max(scene) is the maximum allowed total concurrent load in scene after background congruence conversion.
[0167] Among several feasible solutions, the system preferentially selects the one that differs least from the source language's interactive signature. This can be defined as:
[0168] Dist_sig=w1×D_conc+w2×D_ovl+w3×D_lat+w4×D_dom+w5×D_coe;
[0169] Where Dist_sig is the total difference between the target language performance spectrum and the source language interaction signature; D_conc is the concurrency curve difference; D_ovl is the overlap envelope difference; D_lat is the response delay distribution difference; D_dom is the dominant order difference; D_coe is the co-event alignment bias; w1, w2, w3, w4, and w5 are the weight coefficients corresponding to each difference item, used to reflect the relative importance of each difference in the overall evaluation.
[0170] Therefore, the objective function can be written as:
[0171] J=λ1×Drift+λ2×Dist_sig+λ3×DegradeCost+λ4×SpeedPenalty+λ5×BoundaryRisk;
[0172] Wherein, Drift is the offset of the overall rhythm of the source language; DegradeCost is the cost of degraded actions; SpeedPenalty is the cost of excessive speech rate adjustment; BoundaryRisk is the risk of touching the safety boundary or excessive shifting; J is the overall target value of the overall solution, used to measure the overall quality of the current solution; λ1, λ2, λ3, λ4, and λ5 are the weight coefficients corresponding to each evaluation item, which are used to adjust the proportion of rhythm drift, interaction signature difference, degraded action cost, speech rate adjustment cost, and boundary risk in the overall target, respectively.
[0173] The specific process of generating the performance score is shown in B1-B6.
[0174] B1: Generate multiple candidate words for each dialogue segment node in the target language.
[0175] For each dialogue segment node, at least one synthesizable candidate, i.e., candidate text, is generated in the target language, which can be obtained through machine translation, template rewriting, manual input, etc.
[0176] The "target language" in this case does not refer only to a language code, but to the complete dubbed version configuration to be delivered this time. In one embodiment, it at least includes fields such as lang_code (language code), locale (regional configuration), register (language style or register), style_tag (performance style label), tts_voice_id (selected dubbed voice timbre identifier), delivery_profile (overall delivery configuration plan), etc. For example, even for the same English language, it can be distinguished between American variety show spoken language, British drama tone, or a more exaggerated dubbed style for children.
[0177] For each node v_i, the system preferably generates multiple target language candidates. The candidates can be sourced from machine translation, length-controlled rewriting, template rewriting, or manual revision. For example, if the source language node is "You listen to me first", and the target language is English, multiple candidates such as "Hear me out.", "Wait, listen to me.", "Just let me finish." can be generated in the English target language. For strong interruption scenarios that require quick entry, the system can preferentially select candidates that are shorter, have fewer pauses, and start more neatly; for hesitant response scenarios, candidates with pause slots and hesitant tones can be preferentially retained.
[0178] It should be noted that the introduction of the target language candidate attribute in this case is only for providing a decision-making basis for the "overall arrangement under the constraints of the multi-person interaction director's score", and does not take the translation model, timbre model, or single-sentence performance model itself as the independent protection focus.
[0179] B2: Obtain the synthesis attributes of the candidates carried when multiple candidates enter the constraint solving of the corresponding dialogue director's score in the original film.
[0180] Among them, for each candidate, its synthesizable attributes are extracted / estimated, including but not limited to: expected duration, adjustable speed range, recommended pause positions, whether it contains non-word events (such as laughter), etc.
[0181] Example fields of the synthesis attributes can include the expected duration and its fluctuation range; the allowed speech speed adjustment range; the recommended pause positions; whether it contains non-word events such as laughter, inhalation, sighing, etc.; pragmatic roles such as statement, question, echo, strong interruption, exclamation, etc.; the confidence level of the candidate quality, etc.
[0182] The composition attributes are primarily used to solve the overall dialogue scheduling problem within the performance spectrum. The system uses these attributes to determine whether each candidate utterance is suitable for inclusion in the temporal relationships constrained by the current dialogue director's spectrum.
[0183] B3: Constrain the dialogue director spectrum according to the constraints of the dialogue director spectrum to obtain the feasible region.
[0184] It should be noted that the constraints of the dialogue director spectrum include, but are not limited to, candidate uniqueness constraints, interaction edge constraints, node boundary constraints, speech rate and pause constraints, and intelligibility budget constraints.
[0185] Interactive edge constraints include strong constraints and soft constraints. Interactive constraints aim to satisfy the overlap / delay range of the director's spectrum edges as much as possible; edges with high confidence are used as strong constraints, and edges with low confidence are used as soft constraints.
[0186] B4: Under the constraints of the dialogue director spectrum, the optimal combination is jointly searched within the feasible region based on the synthesis attributes of the candidate words.
[0187] In B4, the optimal combination of candidate choices, start time, speech rate, and pauses is jointly searched within the feasible region. The optimal combination refers to the overall optimal arrangement obtained after jointly evaluating candidate choices, start time, speech rate, and pauses, while satisfying interaction constraints, intelligibility constraints, boundary constraints, and speech rate and pause restrictions.
[0188] B5: Based on the optimal combination, perform a comprehensibility budget check and an interactive signature difference check on the original video.
[0189] B6: If the intelligibility budget check results and the interactive signature difference results both meet their respective requirements, generate the performance score.
[0190] If the interactive signature difference result does not meet its corresponding requirements, a callback solution / adjustment strategy will be implemented.
[0191] If the comprehensibility budget check results do not meet the corresponding requirements, the constraints will be relaxed according to the degradation link.
[0192] If the overall solution fails to find a solution that satisfies all constraints, or if post-solution acceptance reveals that the understandability budget is not met or the interaction signature discrepancy is too large, the system does not directly declare it a failure. Instead, it gradually relaxes the constraints along a preset degradation path. For example, in one embodiment, "whether the understandability budget is met" can be determined according to the following criteria:
[0193] 1. Check whether the number of concurrent users in the foreground exceeds the scene limit within a discrete time window;
[0194] 2. Check if the short-term relative clarity of the main speaker is below the threshold;
[0195] 3. Check whether the cumulative overlap rate within a unit time window exceeds the upper limit;
[0196] 4. Check if strong interaction edges have disappeared in large areas, or if the dominant order has been unacceptably reversed.
[0197] If any of the above conditions are not met, the degradation process begins. Degradation does not involve changing all interactions to sequential playback all at once, but rather proceeds layer by layer according to the principle of "resetting weaker interactions first, while preserving as many main interactions as possible." The degradation sequence may include the following:
[0198] 1. EDGE_LOWCONF_RELAX: First relax the restrictions on low-confidence interaction edges;
[0199] 2. `BACKCHANNEL_SEQUENTIALIZED`: Reduces the presence of the attached edges or processes them sequentially;
[0200] 3. COEVENT_WINDOW_WIDENED: Relaxes the alignment window for common events;
[0201] 4. INTERRUPT_OVERLAP_SHRUNK: The upper limit of allowed overlap for shrinking strong break edges;
[0202] 5. SHORTER_CANDIDATE_SWITCHED: Switch to a shorter target phrase candidate;
[0203] 6. STRICT_SEQUENTIAL_FALLBACK: Degenerates to strict sequential join if still not feasible.
[0204] Each degradation can be recorded with a reason code, the edges or nodes involved, the parameters before and after the change, and the impact on the signature difference. For example, it can be recorded as {reason_code, edge_id / node_id, before_param, after_param, budget_gain, signature_delta, review_flag}.
[0205] Traceable delivery and anomaly degradation (facilitating operations / legal / post-debriefing understanding and review), outputting the director's version number, interaction edge type and parameters, solution results and degradation link (if it occurs), forming a reproducible delivery basis; when the credibility of the interaction clue is low (such as strong noise, multiple people speaking at the same time, severe screen obstruction), it automatically enters a conservative strategy (reducing overlap, connecting according to the order of the subtitle window), and marks the segment as "recommended for manual review".
[0206] Cross-language preservation and self-checking driven by dialogue interaction signatures (avoiding "dynamic drift"). "Interaction signatures" (e.g., curves of concurrent speakers over time, dominant weight order, overlap envelope, response delay distribution, and shared laughter alignment window) are extracted from the original dialogue as a quantifiable baseline for the dialogue layer. The corresponding signature is calculated for the performance spectrum obtained from the target language and compared with the baseline. When the difference exceeds a threshold, a callback is prioritized for resolution (adjusting start / end / overlap / degradation strategies). If necessary, a conservative degradation based on "intelligibility priority" is triggered. Signature differences and processing results are recorded in the log file, making it explainable and reproducible why "this part became more sequential / less overlapping".
[0207] The preferred degradation strategy in this scheme follows the following degradation order:
[0208] 1. Relax restrictions on low-credibility interaction edges;
[0209] 2. Order or reduce the presence of low-dominance-weighted auxiliary edges;
[0210] 3. Expand the alignment window for events such as laughter / cheering;
[0211] 4. Shrink the upper limit of overlap of strong break edges;
[0212] 5. Switch to shorter target language candidates;
[0213] 6. Finally, it degenerates into a strict sequential connection and is marked for manual review.
[0214] Each time degradation occurs, the cause code, the edges or nodes involved, parameter changes, budget gains and signature differences are preferably output to ensure that it is interpretable, rollbackable and verifiable.
[0215] Specific performance spectrum solution and degradation path, such as Figure 4 As shown. Figure 4 A schematic diagram of the performance spectrum solution and degradation link is shown.
[0216] Figure 4 The system first reads the interaction edges, parameters, and credibility from the dialogue director's spectrogram. Then, it combines the target language candidates and their expected duration, speech rate range, pause slots, and other attributes to solve the entire dialogue as a whole. After solving, the system checks whether the intelligibility budget is met, such as whether the number of concurrent users in the foreground exceeds the limit, whether the clarity of the dominant lines is insufficient, and whether strong overlap lasts too long. It also checks whether the difference between the solution result and the source language interaction signature exceeds the threshold. If the requirements are met, the performance spectrogram is output; if the requirements are not met, the constraints are gradually relaxed according to the degradation path. The preferred order includes degrading low credibility edges, conforming to the sequence, widening the common event alignment window, shrinking interrupted overlaps, switching to shorter candidates, and strict sequence backtracking. The reasons for degradation and changes in related parameters are recorded.
[0217] Figure 4To check if the comprehensibility budget is met, check the following items:
[0218] Does the number of concurrent users on the front end exceed the scenario threshold?
[0219] Whether the relative clarity of the lead speaker in the key frequency band is below a threshold;
[0220] Does the cumulative duration of strong overlap exceed the upper limit of the unit time window?
[0221] Are the strong interaction edges being damaged too much, or has the dominant order been unacceptably reversed?
[0222] If any of the above conditions are not met, the constraints are gradually relaxed along the degradation path, with the preferred order being:
[0223] 1. Degradation of low-trusted edges;
[0224] 2. Sequentialization of agreement;
[0225] 3. Widen the window for event alignment;
[0226] 4. Disrupt overlapping contractions;
[0227] 5. Shorter candidate switching time;
[0228] 6. Strictly follow the order of rollback;
[0229] Simultaneously output the following:
[0230] 1. Degradation reason code;
[0231] 2. Degenerate edges;
[0232] 3. Relax the parameters before and after;
[0233] 4. Impact on differences in interactive signatures;
[0234] 5. Is manual review recommended?
[0235] Under the constraints of the dialogue director spectrum, one or more synthesizable candidate words (including expected duration, adjustable speed / pause position, etc.) are generated for each line of the target language. The start, end and overlap arrangements of the entire dialogue are solved with the goal of "maintaining dialogue relationship + prioritizing intelligibility". Hard / strong constraints are set for the edges that "must maintain interactive effect" (such as the host interrupting to control the scene, arguing and interrupting each other) and soft constraints are set for the edges that "can be moderately yielded" (such as weak responses, non-key echoes) to ensure stable delivery when the rhythm is not feasible. The "executable performance spectrum of the dialogue as a whole" is output: the start time of each sentence, the allowed overlap range, the suggested speech rate / pause, and the explanation label of "why interruption / delay is needed here".
[0236] In this implementation, the start time, allowed speed, and pauses of each dialogue segment can be used as decision variables. Interaction edges are transformed into temporal constraints of "allowed overlap / delay intervals," and the "interaction signature difference + total drift + degradation cost" is minimized under the intelligibility budget. The solution can be achieved using engineering-feasible methods such as integer programming, dynamic programming, or heuristic search. Specifically, the first layer constructs a candidate set and feasible region; the second layer performs constraint layering based on strong constraints, weak constraints, and intelligibility budget; the third layer searches for the optimal arrangement within the feasible region; the fourth layer verifies the results based on signature differences and budget; and the fifth layer solves the problem by calling back the degradation link when a solution fails, forming a complete closed loop of "candidate generation -> constraint layering -> joint solution -> acceptance -> degradation callback." The performance spectrum and its interpretation labels are output.
[0237] S105: Based on the interaction types in the performance score and interaction clue set, perform mixing control on the audio of each character in the original film and output the mixing control results.
[0238] It should be noted that mixing control includes interrupting mixing, echoing mixing, laughing mixing, global loudness and peak control, and generating mixing control scripts (mixing scripts are sets of executable mixing control instructions that convert interaction types and dominant weights into a set of instructions (such as sidechain suppression, fade-in / fade-out, frequency band yielding, and concurrency limit trigger points) for easy reuse, rollback, and auditing).
[0239] Interruption mixing: The interrupted speaker fades out shortly at the end, and key clarity bands are "give way" according to the dominant weight (to avoid blurring when two people speak at the same time); the interruptor is given priority in the short window to ensure that the main narrative line is understandable;
[0240] Echoing Mixing: Echoing elements are made less noticeable (e.g., positioned further back, on a narrower frequency band, or at a lower volume) to avoid overshadowing the main sentence; it can automatically degenerate into sequential linking when concurrency exceeds the budget;
[0241] Shared Laughter Recording: Maintain a certain level of synchronization and dynamism by aligning the window to avoid fake interaction where "one person finishes laughing before the next person laughs";
[0242] Global loudness and peak control: Ensure the final product meets the platform's loudness and engineering quality requirements;
[0243] Mixing control script generation: The above mixing control process is output as an executable mixing control script and bound to the director's score / performance score version for easy reuse, fine-tuning and rollback.
[0244] Specifically, based on the interaction types in the performance score and the set of interaction clues, the process of mixing and controlling the audio of each character in the original film and outputting the mixing control results is shown in C1-C5.
[0245] C1: Reads mixing control parameters contained in or associated with the performance score.
[0246] The mixed recording control parameters include, but are not limited to, the start time of each node, the overlap strategy, the degradation flag, and the dominant weight.
[0247] C2: Controls the mixing of audio for each character in the original film based on mixing control parameters and interaction type.
[0248] C3: During the mixing and control of the audio of each character in the original film, for the interactive type of the interruption edge, the end of the audio of the interrupted character in the original film is processed first to obtain the interruption mixing control action and output it.
[0249] The first processing step involves applying a smooth fade-out and sidechain compression to the end of the audio of the interrupted person in the original video.
[0250] C4: For the interactive type of shared laughter / shared breathing, perform a second processing on the original video within the alignment window to obtain the shared laughter mixing control action and output it.
[0251] The second process involves synchronously triggering multitrack laughter or shouts from the original film within the alignment window.
[0252] C5: For the interaction type of the echoing edge, the key frequency band of the echoing volume in the original film is processed in the third step to obtain the echoing mixing control action and output it.
[0253] The third processing step involves narrowing the key frequency band of the echo volume in the original film.
[0254] The mixing control rules for interactive types (ensuring overlap is "clearly audible and doesn't interrupt the narrative") are as follows: For interrupted characters: a short fade-out is applied to the end of the interrupted character, while the interruptor is slightly boosted or sidechained to ensure the audience's auditory focus is on the "leading speaker"; For shared laughter / breathing: a certain sense of immediacy is preserved (e.g., a wider dynamic range or a more natural proportion of bedtime sounds), avoiding fake interactions where "one person finishes laughing before the next person laughs"; For high-density dialogue segments: "intelligibility priority" constraints are adopted, such as limiting the maximum number of concurrent speakers at any given time, and automatically degrading weak interactive edges to sequential connections when necessary, with the reason for the degradation output. Specifically, during the mixing control of the audio of each character in the original film based on the interaction type in the performance spectrum and interaction clue set, the performance spectrum is first read to obtain the entry point, overlap strategy, and degradation marker for each sentence; then, mixing event groups are established according to the interactive edges; then, different action templates are called for different interaction types; the different action templates include the following:
[0255] For interrupted edges: short fade-out / downward pressure on the tail of the interrupted element, short window dominant boost, sidechain suppression, and necessary frequency band relinquishment for the interrupted element;
[0256] For the receiving end: the key is to maintain a tight connection and avoid artificially widening the gap;
[0257] For the embracing edge: reduce its presence, narrow its critical frequency band occupancy, and if necessary, move it to the background layer;
[0258] For shared laughter / same event edges: trigger synchronously within the alignment window to preserve the sense of shared presence;
[0259] These actions will eventually be written into a mixed control script, which will record at least the time window, action type, parameters, reason code, and rollback point.
[0260] For a better understanding of the mixed recording control process, please refer to [reference needed]. Figure 5 As shown.
[0261] Figure 5 In the process, the start time, overlap strategy, and degradation marker of each node in the performance score are first read. Then, combined with the interaction type and dominance weight in the director's score, the audio of each character is grouped and controlled. For interruption edges, a short fade-out or downward pressure can be applied to the tail of the interrupted character, and a short dominance boost and necessary frequency band yielding can be applied to the interruptor. For echoing edges, their presence can be reduced and moved to the background layer when necessary. For shared laughter or the same event edges, they can be triggered synchronously within the alignment window to preserve the shared sense of presence. The above control actions can finally be written into the mixing control script and used for subsequent rendering, review, and rollback.
[0262] The output mixing control results should include at least the target language dubbing finished audio track or video, updated subtitle timing information (if required), dialogue director's score version, performance score, interactive signatures and their differences, mixing control script, and reasons for degradation and suggested review points.
[0263] Mixing control scripts are output and reused (solidifying listening experience into executable assets), transforming interactive mixing rules such as "interrupting / agreeing / laughing together" into executable mixing control scripts (e.g., sidechain suppression curves, fade-in / fade-out curves, priority / dominant weights, and concurrent upper limit trigger points), which are bound to the director's version; script templates can be reused for similar programs / similar dialogue styles, reducing repeated trial and error in post-production, and supporting one-click rollback to historical stable scripts.
[0264] The scripting process for mixed recording control is as follows:
[0265] When performing mixed recording control based on the performance score and interaction type, the system can execute the process in the order of "reading the performance score—compiling actions according to the interaction type—generating control curves—rendering and acceptance," instead of simply performing a single overall compression. The process can be further broken down as follows:
[0266] 1. Track Preparation. Organize the target speech audio, non-word event audio, and necessary ambient bed sounds for each role into a multi-track object, and read the start_time, edge_type, W_dom, and degrade_flags corresponding to each node.
[0267] 2. Interactive Grouping. Create mixing event groups (mix_group_g) based on the interactive edges in the director's chart. Each group should include at least the participating roles, time window, leader, interaction type, and set of allowed control actions.
[0268] 3. Action compilation. Different rule templates are invoked for different interaction types:
[0269] (1) Disruption: Apply short fade-out or energy down pressure to the tail of the interrupted one, and do a dominant boost, side chain suppression and necessary frequency band concession to the interrupted one within the key start-up window;
[0270] (2) Responding to the next sentence: While maintaining the integrity of the end of the previous sentence, control the next sentence to enter the scene as soon as possible;
[0271] (3) Adhering edge: Reduce the volume of the adhering edge, narrow its key frequency band occupation, and move it to the background layer if necessary;
[0272] (4) Shared laughter or simultaneous event edge: Simultaneously trigger multi-track laughter or shouts within the alignment window while preserving the sense of shared environment.
[0273] 4. Parameter instantiation. Action templates are preferably converted into scripted parameters, such as fade_in_ms, fade_out_ms, duck_depth_db, eq_yield_band, priority, rollback_ref, etc.
[0274] 5. Script Output. Writes the mixing actions to the mixing control script. Example fields may include:
[0275] script_id, group_id, track_id, t_in, t_out, action_type, action_params, reason_code, expected_effect, rollback_ref`.
[0276] 6. Rendering and Acceptance. After mixing, check whether the clarity, loudness, peak value, and key interactions of the main dialogue are preserved. If the acceptance fails, rollback to the previous stable script can be performed based on rollback_ref, or the solution can be recalculated. Therefore, "mixing control based on performance score and interaction edge type" corresponds to a closed-loop process from director score parameters and performance score results to automatic mixing control scripts, rather than a post-production description that relies on general human experience.
[0277] In one implementation scenario, the director's lineup and mixing control for a variety show's "forced interruption" scene:
[0278] 1) Scene: The host frequently interrupts the guests to control the scene, and there are obvious instances of interrupting and suppressing the ending in the original video.
[0279] 2) Construction of the dialogue director spectrum: The system judges the strong interruption edge based on clues such as "the next sentence enters before the previous sentence ends", "sudden increase in energy" and "the camera cuts to the host's close-up", and sets a high dominant weight W_dom.
[0280] 3) Performance spectrum solution: In the target language, prioritize ensuring that the interruption overlap falls within [O_min, O_max]. If the target language is too long to be feasible, first compress the weak attachment edges, then shrink the interruption overlap range, and output the reason for the degradation.
[0281] 4) Mixing control: The interrupted person's end is faded out briefly, and the host is slightly raised to ensure that the audience perceives the real interaction of "the host interrupting and controlling the scene".
[0282] 5) Interactive signature self-check: Compare signature indicators such as "dominant weight order, overlapping envelope, and concurrent peak"; if the difference exceeds the threshold, the system will prioritize callback to solve or further degrade weak interactive edges, and write the signature difference into the log.
[0283] In another implementation scenario, the intelligibility constraints of a group argument involving "high-density responses and agreement" are as follows:
[0284] 1) Scenario: Multiple people quickly respond, interspersed with short echoes ("Yes, yes, yes / that's right").
[0285] 2) Construction of dialogue director spectrum: The system identifies a large number of responding and echoing sides and counts the potential number of concurrent users at the same time.
[0286] 3) Performance spectrum solution: For time windows where concurrency exceeds the threshold, prioritize degenerate the connecting edges (reduce volume or insert sequentially) to ensure the comprehensibility of core plot lines.
[0287] 4) Output: Simultaneously output the list of degraded embellishments and the reasons, and generate the corresponding mixing control script, which makes it easier to decide later whether to retain some embellishments to enhance the atmosphere.
[0288] Examples of detailed algorithms and data structures:
[0289] The key data structures involved in this solution include, but are not limited to:
[0290] Director's Spectrum Node Table (Example fields: node_id, dialogue_id, speaker_id / role_id, src_time_range, text_ref, node_conf, version number);
[0291] Director's Spectrum Interactive Edge Table (Example fields: from_id, to_id, edge_type, params<overlap / delay / alignment window / dominant weight>, Conf_edge, evidence_pointers, version number);
[0292] Candidate attribute table (example fields: candidate_id, node_id, text_variant_ref, pred_duration, speed_range, pause_slots, event_flags <laugh / sigh, etc.>);
[0293] Performance spectrogram (example fields: node_id, start_time, overlap_policy, speed, pause_plan, degrade_flags, explain_tags, signature difference summary);
[0294] Record the script table (example fields: script_id, events<time window / track / action / curve parameters>, priority, rollback_ref, bound director's score / acting score version).
[0295] Compared to existing solutions that rely on sentence-by-sentence synthesis and simple splicing / mixing, this solution offers the following key differences and benefits: 1) It structures the "realism of dialogue" into a dialogue director's score and performance score, making it a solvable object and reusable asset, rather than relying on human experience; 2) It allows and controls "reasonable overlap and reaction delay," transforming interaction from "errors" to "controllable effects," making high-density dialogues such as variety show control and group arguments closer to reality; 3) The overall dialogue solution avoids global distortion caused by sentence-by-sentence local optimization, and cross-language versions are more likely to maintain the original dialogue rounds and emotional progression; 4) It provides clear degradation paths and readable reasons, and outputs interaction signature differences, making "conservative degradation" in infeasible / uncertain scenarios explainable and reproducible; 5) Through interaction signature-driven cross-language preservation and self-checking, it reduces the risk of dialogue dynamics drift and rework; 6) Through mixed recording control, it controls scripted delivery and reuse, making the mixed recording strategy replicable, rollbackable, and auditable, reducing reliance on human post-production listening experience.
[0296] This solution reduces post-production costs by shifting the crucial work of ensuring "realistic dialogue" to the system's automatically generated director's score and mixing control. This transforms post-production from "line-by-line editing + repeated listening" to "refining a small number of segments," significantly reducing manpower investment. It also shortens delivery cycles and improves replicability: the director's score, as a reusable asset, can be quickly migrated to similar content and new language versions once optimized for a particular program / dialogue style, reducing repeated trial and error. Furthermore, it enhances user experience and platform brand: multi-person dialogues are more natural, with interruptions, responses, and emotional progression more closely mirroring the original film's rhythm, reducing negative feedback such as "fake dubbing" or "reading from a script," thus improving completion rates and reputation. Finally, it reduces rework and communication costs: the system outputs readable information such as "where the degradation occurred and why," reducing cross-team conflicts and improving operational decision-making efficiency.
[0297] This solution enables more natural multi-person dialogue: interactions such as interrupting, responding, and laughing together are closer to the rhythm of the original film; greater controllability: while retaining interaction, mixing rules and concurrency limits ensure that the audience can hear clearly; significantly reduced post-production costs: the dialogue director's score and performance score transform post-production from "remaking" to "fine-tuning"; cross-language reusability: the director's score, as a dialogue layer asset, can be migrated to new language versions and similar programs, improving the efficiency of large-scale production; more stable and measurable interaction consistency: interaction signatures provide a quantifiable basis for "whether the interactive tension of the original film has been retained," facilitating acceptance and iteration; and more controllable delivery: the mixing control script and record keeping make post-production fine-tuning, gray-scale testing, and rollback more lightweight, reducing the risk of going live.
[0298] In this embodiment, dialogue interaction cues are extracted from the audio, visuals, and subtitles of the original film to generate a reproducible and reusable dialogue director's score. This transforms the realism of multi-person dialogue from a subjective aesthetic into a calculable and reusable structured object. Under the constraints of the dialogue director's score, the overall dialogue scheduling is solved based on the original film's corresponding dialogue director's score. While ensuring intelligibility and legal boundaries, a certain range of overlap and delay is allowed and precisely controlled, making the dubbing more like a real group dialogue rather than a series of script readings. Furthermore, by solidifying dialogue relationships and rhythm constraints into the dialogue director's score, and then performing feasible scheduling and mixing in the target language, the dialogue turns and emotional progression rhythm of the original film are maintained, thereby avoiding dialogue distortion in multilingual dubbing and improving the realism of multilingual dubbing.
[0299] Based on the above embodiments Figure 1 The present application discloses a dialogue processing method and a corresponding dialogue processing apparatus, such as... Figure 6 As shown, the dialogue processing device includes:
[0300] Aggregation unit 601 is used to aggregate the acquired original pieces to obtain multiple dialogue fragments;
[0301] Extraction unit 602 is used to extract a set of interactive cues for multi-person interaction from each dialogue segment;
[0302] Construction unit 603 is used to construct the dialogue director spectrum corresponding to the original film based on the set of interactive clues; wherein, the dialogue director spectrum corresponding to the original film is a structured intermediate representation;
[0303] The scheduling solution unit 604 is used to perform overall dialogue scheduling solution based on the dialogue director's spectrum corresponding to the original film under the constraint of the dialogue director's spectrum, and generate the performance spectrum.
[0304] The control output unit 605 is used to control the mixing of audio for each character in the original film according to the interaction type in the performance score and the set of interaction clues, and output the mixing control result.
[0305] Furthermore, aggregation unit 601 obtains multiple dialogue fragments, including:
[0306] The alignment module is used to force alignment of the acquired original image to obtain each smallest text segment, and uses each smallest text segment as an initial unit.
[0307] The aggregation module is used to aggregate the initial units into multiple dialogue fragments according to temporal adjacency, scene continuity, and semantic continuity.
[0308] Furthermore, the extraction unit 602 includes:
[0309] The extraction module is used to extract audio cues, text cues, and / or visual cues from each dialogue segment;
[0310] The component module is used to assemble a set of interactive cues based on audio cues, text cues, and / or visual cues in each dialogue segment.
[0311] Furthermore, building unit 603 includes:
[0312] The fusion module is used to merge audio cues, text cues, and / or visual cues from the set of interactive cues into an interactive type;
[0313] The first acquisition module is used to acquire the interaction intensity parameter of the interaction type;
[0314] The determination module is used to determine the round relationship between multiple dialogue segment nodes on the timeline;
[0315] The construction module is used to build the dialogue director spectrum corresponding to the original film based on the round relationship, interaction type and interaction intensity parameters.
[0316] Furthermore, the scheduling solution unit 604 includes:
[0317] The first generation module is used to generate multiple candidate words corresponding to each dialogue segment node in the target language;
[0318] The second acquisition module is used to acquire the synthesis attributes of the candidate words when multiple candidate words enter the constraint solution of the dialogue director spectrum corresponding to the original film;
[0319] The constraint layering module is used to perform constraint layering on the dialogue director spectrum according to the constraints of the dialogue director spectrum to obtain the feasible region.
[0320] The search module is used to jointly search for the optimal combination within the feasible domain based on the synthesis attributes of candidate words, under the constraints of the dialogue director spectrum.
[0321] The inspection module is used to perform intelligibility budget checks and interactive signature difference checks on the original video based on the optimal combination.
[0322] The second generation module is used to generate a performance score if the intelligibility budget check results and the interaction signature difference results both meet their respective requirements.
[0323] Furthermore, the control output unit 605 includes:
[0324] The reading module is used to read the mixing control parameters contained in or associated with the performance score;
[0325] The mixing control module is used to control the mixing of audio for each character in the original film according to the mixing control parameters and interaction type.
[0326] The first operation output module is used to perform the first processing on the end of the audio of the interrupted character in the original film during the process of mixing and controlling the audio of each character in the original film. For the interaction type of the interruption edge, the module obtains the interruption mixing control action and outputs it.
[0327] The second operation output module is used to perform a second processing on the original video simultaneously within the alignment window for the interaction type of laughing / breathing together, to obtain and output the laughing mixing control action;
[0328] The third operation output module is used to perform third processing on the key frequency band of the volume of the echoer in the original film for the interaction type of the echoer, to obtain the echo mixing control action and output it.
[0329] In this embodiment, dialogue interaction cues are extracted from the audio, visuals, and subtitles of the original film to generate a reproducible and reusable dialogue director's score. This transforms the realism of multi-person dialogue from a subjective aesthetic into a calculable and reusable structured object. Under the constraints of the dialogue director's score, the overall dialogue scheduling is solved based on the original film's corresponding dialogue director's score. While ensuring intelligibility and legal boundaries, a certain range of overlap and delay is allowed and precisely controlled, making the dubbing more like a real group dialogue rather than a series of script readings. Furthermore, by solidifying dialogue relationships and rhythm constraints into the dialogue director's score, and then performing feasible scheduling and mixing in the target language, the dialogue turns and emotional progression rhythm of the original film are maintained, thereby avoiding dialogue distortion in multilingual dubbing and improving the realism of multilingual dubbing.
[0330] This application embodiment also provides a storage medium, the storage medium including stored instructions, wherein, when the instructions are executed, the device where the storage medium is located is controlled to perform the dialogue processing method as described above.
[0331] This application also provides an electronic device, the structural schematic diagram of which is shown below. Figure 7 As shown, it specifically includes a memory 701 and one or more instructions 702, wherein one or more instructions 702 are stored in the memory 701 and configured to be executed by one or more processors 703 to perform the above-mentioned dialogue processing method.
[0332] The steps in the methods of the various embodiments of this application can be adjusted, combined, and deleted according to actual needs. It should be noted that, in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations.
[0333] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0334] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A dialog processing method characterized by, The method includes: The acquired original footage is aggregated to obtain multiple dialogue fragments; Extract a set of interactive cues from each dialogue segment; The dialogue director spectrum corresponding to the original film is constructed based on the set of interactive clues; wherein, the dialogue director spectrum corresponding to the original film is a structured intermediate representation; the dialogue director spectrum is a structured dialogue rhythm instruction composed of dialogue segment nodes and interactive edges; Under the constraints of the dialogue director spectrum, the overall dialogue scheduling is solved based on the dialogue director spectrum corresponding to the original film to generate the performance spectrum; the performance spectrum includes at least one of the following for each line of dialogue: start time / relative offset, overlap strategy, suggested speech rate / pause, and degradation marker; Based on the performance spectrum and the interaction type in the interaction clue set, the audio of each character in the original film is mixed and controlled, and the mixing control result is output; the interaction clue set consists of audio clues, text clues and / or visual clues in each dialogue segment; the interaction type includes at least one of interruption edge, response edge, reaction edge, laughter / event edge and agreement edge.
2. The method according to claim 1, characterized in that, The obtained original footage is aggregated to obtain multiple dialogue fragments, including: Force alignment is performed on the obtained original image to obtain each smallest text segment, and each smallest text segment is used as an initial unit; Based on temporal adjacency, scene continuity, and semantic continuity, the initial units are aggregated into multiple dialogue fragments.
3. The method according to claim 1, characterized in that, The set of interactive cues extracted from each dialogue segment for multi-person interaction includes: Extract audio cues, text cues, and / or visual cues from each dialogue segment; The interactive cues are composed of audio cues, text cues, and / or visual cues in each dialogue segment.
4. The method according to claim 1, characterized in that, The step of constructing the dialogue director spectrum corresponding to the original film based on the set of interactive clues includes: Merge audio cues, text cues, and / or visual cues from the set of interactive cues into an interactive type; Get the interaction intensity parameter of the interaction type; Determine the round-based relationship between multiple dialogue segment nodes on the timeline; Based on the round relationship, the interaction type, and the interaction intensity parameter, construct the dialogue director spectrum corresponding to the original film.
5. The method according to claim 1, characterized in that, The process of generating a performance score by solving the overall dialogue scheduling problem based on the dialogue director's score corresponding to the original film, under the constraints of the dialogue director's score, includes: Generate multiple candidate words for each dialogue segment node in the target language; When multiple candidate words are entered into the constraint solution of the dialogue director spectrum corresponding to the original film, the synthesis attributes of the candidate words are obtained; the dialogue director spectrum is constrained and layered according to the constraints of the dialogue director spectrum to obtain the feasible region. Under the constraints of the dialogue director spectrum, the optimal combination is jointly searched within the feasible domain based on the synthesis attributes of the candidate words; Based on the optimal combination, the original video is subjected to intelligibility budget check and interaction signature difference check; the intelligibility budget is a set of constraints on the number of concurrent speakers allowed at the same time and the allowable overlap intensity; the interaction signature is generated for each dialogue segment based on the dialogue director spectrum and the time relationship of the original video; If both the intelligibility budget check results and the interactive signature difference results meet their respective requirements, a performance score is generated.
6. The method according to claim 1, characterized in that, The step of mixing and controlling the audio of each character in the original film according to the interaction type in the performance spectrum and interaction clue set, and outputting the mixing and control result, includes: Read the mixing control parameters contained in or associated with the performance score; The audio of each character in the original film is mixed and controlled according to the mixing control parameters and the interaction type. During the process of mixing and controlling the audio of each character in the original film, for the interaction type of the interrupted side, the end of the audio of the interrupted character in the original film is processed first to obtain the interruption mixing control action and output it. For the interactive type of shared laughter / shared breathing, the original video is simultaneously processed in the alignment window to obtain and output the shared laughter mixing control action; For the interaction type of the echoing edge, the key frequency band of the echoing volume in the original film is processed in the third way to obtain the echoing mixing control action and output it.
7. A dialogue processing device, characterized in that, The device includes: The aggregation unit is used to aggregate the acquired original clips to obtain multiple dialogue segments; Extraction unit, used to extract a set of interaction cues for multi-person interaction from each dialogue segment; A construction unit is used to construct the dialogue director spectrum corresponding to the original film based on the set of interactive clues; wherein, the dialogue director spectrum corresponding to the original film is a structured intermediate representation; the dialogue director spectrum is a structured dialogue rhythm instruction composed of dialogue segment nodes and interactive edges; The scheduling and solving unit is used to perform overall dialogue scheduling and solving based on the dialogue director spectrum corresponding to the original film under the constraints of the dialogue director spectrum, and generate the performance spectrum; the performance spectrum includes at least one of the following for each line of dialogue: start time / relative offset, overlap strategy, suggested speech rate / pause, and degradation marker; The control output unit is used to control the mixing of audio for each character in the original film according to the performance spectrum and the interaction type in the interaction cue set, and output the mixing control result; the interaction cue set consists of audio cue, text cue and / or visual cue in each dialogue segment; the interaction type includes at least one of interruption cue, response cue, reaction cue, laughter / event cue and agreement cue.
8. The apparatus according to claim 7, characterized in that, The aggregation unit includes: The alignment module is used to force alignment of the acquired original image to obtain each minimum text segment, and to use each minimum text segment as an initial unit. The aggregation module is used to aggregate the initial units into multiple dialogue fragments according to temporal adjacency, scene continuity, and semantic continuity.
9. A storage medium, characterized in that, The storage medium includes stored instructions, wherein, when the instructions are executed, the device in which the storage medium resides controls the execution of the dialogue processing method as described in any one of claims 1 to 6.
10. An electronic device, characterized in that, It includes a memory, and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors as described in any one of claims 1 to 6.