Cross-modal content generation scheduling method based on shot data structure

CN122802754APending Publication Date: 2026-09-22BEIJING ZHIXUN HIVE INFORMATION TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202610955283.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

[0004]有鉴于此,本发明提供了一种基于分镜驱动的多模态生成调度方法,用以解决现有技术只能重新执行素材获取、分镜头视频生成和剪辑合成,重新生成全部分镜头视频片段,或人工手动替换整个视频片段,无法实现仅对单一模态流数据的定向替换的问题

Benefits of technology

[0026]通过模态参数映射关系表实现了输入数据在生成前的按模态预分流,使不同生成智能体仅接收各自所需的纯净描述子集,消除了跨模态数据串扰;具体而言,步骤S1中建立的模态参数映射关系表在输入阶段即规定了内容描述数据集中各数据与目标生成模态的显式映射关系。步骤S2中,根据该模态参数映射关系表将每个分镜数据的内容描述数据集拆分为画面描述子数据集、配音描述子数据集和字幕描述子数据集,并分别构建第一模态调度指令、第二模态调度指令和第三模态调度指令,并将其发送至对应的画面生成智能体、配音生成智能体和字幕生成智能体。使同一段自然语言描述文本中的不同数据在进入生成流程前即被显式归属至对应的目标生成模态,画面生成智能体、配音生成智能体和字幕生成智能体仅接收其生成任务所需的画面描述子数据集、配音描述子数据集和字幕描述子数据集,从根本上消除了跨模态数据在输入阶段的耦合与串扰。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802754A_ABST
    Figure CN122802754A_ABST
Patent Text Reader

Abstract

The application discloses a cross-modal content generation scheduling method based on a split shot data structure, comprising the following steps: associating and binding split shot video stream data, split shot audio stream data and split shot subtitle stream data to generate a single split shot multi-modal data package; receiving a modal rescheduling signal, wherein the modal rescheduling signal comprises a target split shot identifier, a to-be-replaced modal type identifier and an update description data set; based on the to-be-replaced modal type identifier, locating a to-be-updated modal stream in the single split shot multi-modal data package; sending the update description data set to a generation intelligent agent corresponding to the to-be-replaced modal type identifier to obtain a generated update modal stream; overwriting the update modal stream to a storage location pointed to by the to-be-replaced modal type identifier in the single split shot multi-modal data package, and keeping non-target modal stream data in the single split shot multi-modal data package unchanged. The method solves the problem that only directional replacement of single modal stream data cannot be realized in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital media technology, and more specifically, to a cross-modal content generation and scheduling method based on a storyboard data structure. The invention may also be titled a multimodal generation and scheduling method based on storyboard-driven data. Background Technology

[0002] The short video industry is currently experiencing explosive growth, and the market demand for efficient video generation tools is increasing daily.

[0003] Existing technology (application number: 202511044744.8) discloses an AI-powered intelligent short video generation method based on multi-agent collaboration. This method performs semantic retrieval in a media library based on the semantic content of each shot in the storyboard. If the similarity is lower than a preset threshold, it calls an image or video generation model to generate the video. Simultaneously, it combines the acquired or generated media with script parameters to generate storyboard video clips. However, the generated clips are already mixed audiovisual data (directly outputting video clips with visuals). In these clips, the visuals and audio are coupled in the same generation channel, making it impossible to independently split and replace the video clips according to modality type after output. When a user needs to modify a specific modality, this method can only re-execute the media acquisition, storyboard video generation, and editing synthesis to regenerate all storyboard video clips, or manually replace the entire video clip. It cannot achieve targeted replacement of data from a single modality stream. Therefore, this is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0004] In view of this, the present invention provides a multimodal generation scheduling method based on storyboard-driven approach to solve the problem that existing technologies can only re-execute material acquisition, storyboard video generation and editing synthesis to regenerate all storyboard video segments, or manually replace the entire video segment, and cannot achieve targeted replacement of only single modal stream data.

[0005] This application provides a multimodal generation scheduling method based on storyboard-driven approach, including the following steps:

[0006] Natural language description text and scene constraints are input into a large language model to generate structured scene script data. The structured scene script data contains scene sequences, each of which includes at least two scene data sets. Each scene data set corresponds to a unique scene identifier, scene duration data, content description dataset, and modality parameter mapping table. The modality parameter mapping table is used to route different data in the content description dataset to different target generation modalities.

[0007] The modal parameter mapping table is parsed, and based on the mapping relationship between each data in the modal parameter mapping table and the target generation modality, the content description dataset of each storyboard data is split into a scene description subset, a dubbing description subset, and a subtitle description subset. A first modal scheduling instruction is constructed based on the scene description subset and sent to the scene generation agent. A second modal scheduling instruction is constructed based on the dubbing description subset and sent to the dubbing generation agent. A third modal scheduling instruction is constructed based on the subtitle description subset and sent to the subtitle generation agent. The first, second, and third modal scheduling instructions all include the unique identifier of the storyboard and the storyboard duration data.

[0008] The scene description subset and the scene duration data are input into the scene generation agent to generate scene video stream data; the dubbing description subset and the scene duration data are input into the dubbing generation agent to generate scene audio stream data; the subtitle description subset and the scene duration data are input into the subtitle generation agent to generate scene subtitle stream data; the scene video stream data, scene audio stream data, and scene subtitle stream data are received, and the scene video stream data, scene audio stream data, and scene subtitle stream data are associated and bound using the scene unique identifier as the association key to generate a single scene multimodal data package;

[0009] The system receives a modality rescheduling signal, which includes a target segment identifier, a modality type identifier to be replaced, and an update description dataset. Based on the modality type identifier to be replaced, the system locates the modality stream to be updated in the single-segment multimodal data packet. The system sends the update description dataset to the generating agent corresponding to the modality type identifier to be replaced to obtain the generated update modality stream. Using the target segment identifier as the primary key, the system overwrites the update modality stream to the storage location pointed to by the modality type identifier to be replaced in the single-segment multimodal data packet, while keeping the non-target modality stream data in the single-segment multimodal data packet unchanged.

[0010] Optionally, the modal parameter mapping table includes the following mapping rules:

[0011] The content description dataset includes scene description data, character action description data, character dialogue text data, character identity identification data, and storyboard duration data;

[0012] The scene description data and the character action description data are mapped to the screen description sub-dataset; the character dialogue text data and the character identity data are mapped to the dubbing description sub-dataset; the character dialogue text data and the storyboard duration data are mapped to the subtitle description sub-dataset.

[0013] Furthermore, there is a cross-reference relationship between the screen description sub-dataset, the dubbing description sub-dataset, and the subtitle description sub-dataset: when the character dialogue text data in the dubbing description sub-dataset changes, the subtitle description sub-dataset automatically synchronizes and obtains the changed character dialogue text data as input, while the screen description sub-dataset is not affected by the change in the character dialogue text data.

[0014] Optionally, the third mode scheduling instruction includes:

[0015] The subtitle stream data includes at least two subtitle data sets, each of which includes subtitle presentation duration data and the number of subtitle characters;

[0016] The start and end timestamps of each subtitle data are constrained to fall within the range of the start and end timestamps of the target storyboard data, and the subtitle presentation duration is not less than the preset minimum subtitle duration; the subtitle data is segmented according to the semantic punctuation of the character dialogue text data, and the number of characters in each subtitle data does not exceed the preset maximum number of subtitle characters.

[0017] Optionally, the first modal scheduling instruction includes a set of screen constraint parameters, which includes:

[0018] The intelligent agent for generating the scene is constrained to use the unique identifier of the scene as a tracking token. When iteratively regenerating the video stream data of the scene corresponding to the same unique identifier of the scene, the consistency offset of the main skeleton features and color distribution features between the previous generation and the subsequent regeneration shall not exceed a preset perception threshold. The appearance of the same scene data before and after the iteration shall not be abruptly changed due to the regeneration operation.

[0019] The start and end frames of the storyboard video stream data output by the image generation agent are aligned with the start and end timestamps of the storyboard duration data carried in the second and third modal scheduling instructions, and the total number of frames of the storyboard video stream data satisfies the following conditions: Where T is the segment duration data and fps is the preset frame rate;

[0020] The intelligent agent generating the scene is constrained to maintain the continuity of the optical flow field and the consistency of color between adjacent frames when generating the storyboard video stream data, and the structural similarity index of adjacent frames is not lower than a preset threshold.

[0021] Optionally, the second modal scheduling instruction includes a set of dubbing constraint parameters, which includes:

[0022] The voice-generating agent is constrained to retrieve and load a timbre embedding vector globally bound to the character identity from a preset timbre index library based on the character identity identifier. The same character identity identifier always maintains the same timbre embedding vector throughout all calls to the voice-generating agent.

[0023] The voice-generating agent is constrained to use the storyboard duration data as the upper time limit, and the actual duration of the output voice-over audio data satisfies the following conditions: The segmented audio stream data, where L represents the actual duration of the dubbing audio data, and Δ is the preset silence protection interval; when the base duration of the dubbing description subset calculated according to the base speech rate exceeds At that time, the voice-generating agent activates the prosody scaling mechanism, and the actual output duration L is adjusted by the prosody scaling factor γ. , This represents the prosodic scaling factor, and the range of values ​​for γ is limited to . , Indicates the base duration for calculating the base speech rate;

[0024] The voice-over generation agent is constrained to extract the emotional tags corresponding to the storyboard data from the structured storyboard script data, and select the fundamental frequency envelope, energy envelope and speech rate curve from the preset prosodic parameter space according to the emotional tags, so that the prosodic features of the storyboard audio stream data match the emotional tags.

[0025] Compared with existing technologies, the multimodal generation and scheduling method based on storyboard-driven approach provided by this invention achieves at least the following beneficial effects:

[0026] The modal parameter mapping table enables modal pre-splitting of input data before generation, ensuring that different generating agents receive only their required clean description subsets, thus eliminating cross-modal data crosstalk. Specifically, the modal parameter mapping table established in step S1 specifies the explicit mapping relationship between each data in the content description dataset and the target generating modality at the input stage. In step S2, based on this modal parameter mapping table, the content description dataset of each storyboard data is split into a scene description subset, a dubbing description subset, and a subtitle description subset. First modal scheduling instructions, second modal scheduling instructions, and third modal scheduling instructions are then constructed and sent to the corresponding scene generating agent, dubbing generating agent, and subtitle generating agent, respectively. This ensures that different data in the same natural language description text are explicitly assigned to their corresponding target generating modality before entering the generation process. The scene generating agent, dubbing generating agent, and subtitle generating agent only receive the scene description subset, dubbing description subset, and subtitle description subset required for their generation tasks, fundamentally eliminating cross-modal data coupling and crosstalk at the input stage.

[0027] Of course, any product implementing the present invention does not need to achieve all of the above additional technical effects while solving the background technical problem. It is sufficient for the product to solve the background technical problem first. The additional technical effects are effects that are beyond the understanding of those skilled in the art when the specific structure of the present invention is combined with a specific environment.

[0028] Other features and advantages of the invention will become clear from the following detailed description of exemplary embodiments of the invention with reference to the accompanying drawings. Attached Figure Description

[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments of the invention and, together with their description, serve to explain the principles of the invention.

[0030] Figure 1 This is a flowchart illustrating the multimodal generation and scheduling method based on scene-driven approach provided by the present invention. Detailed Implementation

[0031] Various exemplary embodiments of the present invention will now be described in detail with reference to the accompanying drawings. It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values ​​of the components and steps set forth in these embodiments do not limit the scope of the invention.

[0032] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the invention or its application or use.

[0033] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, they should be considered part of the specification.

[0034] In all the examples shown and discussed herein, any specific values ​​should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.

[0035] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.

[0036] See Figure 1 As shown, Figure 1 This is a flowchart illustrating the multimodal generation and scheduling method based on storyboard-driven approach provided by the present invention. This embodiment provides a multimodal generation and scheduling method based on storyboard-driven approach, including the following steps:

[0037] Step S1: Input the natural language description text and storyboard constraints into the large language model to generate structured storyboard script data; the structured storyboard script data contains a storyboard sequence, and the storyboard sequence includes at least two storyboard data. Each storyboard data corresponds to a unique storyboard identifier, storyboard duration data, content description dataset, and modality parameter mapping table; the modality parameter mapping table is used to route different data in the content description dataset to different target generation modalities;

[0038] Specifically, before executing step S1, it is necessary to obtain the natural language description text input by the user, such as the user input "an orange cat is making fish soup in the kitchen". This natural language description text is then input into the large language model, and scene constraints are applied during input (such as rules for constraining the range of scene duration data, character consistency, and scene logic continuity), causing the large language model to output structured scene script data. The structured scene script data contains a scene sequence, which includes multiple scene data sets. Each scene data set corresponds to a unique scene identifier, scene duration data, content description dataset, and modal parameter mapping table. The modal parameter mapping table is used to route different data in the content description dataset to different target generation modalities. Target generation modalities include visual generation modalities, dubbing generation modalities, and subtitle generation modalities. For example, scene description data and character action description data are assigned to the visual generation modality; character dialogue text data and character identity data are assigned to the dubbing generation modality; and character dialogue text data and scene duration data are assigned to the subtitle generation modality.

[0039] The large language model described above can be either the Doubao model or the DeepSeek model; this embodiment does not specifically limit it.

[0040] Step S2: Parse the modal parameter mapping table. Based on the mapping relationship between each data in the modal parameter mapping table and the target generation modality, split the content description dataset of each storyboard data into a visual description subset, a dubbing description subset, and a subtitle description subset. Construct a first modal scheduling instruction based on the visual description subset and send it to the visual generation agent. Construct a second modal scheduling instruction based on the dubbing description subset and send it to the dubbing generation agent. Construct a third modal scheduling instruction based on the subtitle description subset and send it to the subtitle generation agent. The first, second, and third modal scheduling instructions all include a unique storyboard identifier and storyboard duration data.

[0041] Specifically, the modal parameter mapping table is parsed, and based on the mapping relationship between each data point in the table and the target generation modality, the content description dataset of each storyboard data is split into a visual description subset, a dubbing description subset, and a subtitle description subset. A first modal scheduling instruction is constructed based on the visual description subset and sent to the visual generation agent; a second modal scheduling instruction is constructed based on the dubbing description subset and sent to the dubbing generation agent; a third modal scheduling instruction is constructed based on the subtitle description subset and sent to the subtitle generation agent. All three modal scheduling instructions include a unique storyboard identifier and storyboard duration data, ensuring that the visual generation agent, dubbing generation agent, and subtitle generation agent are aware of the time boundaries of the storyboard data during generation.

[0042] The aforementioned intelligent agents for generating images, voiceovers, and subtitles can be implemented using a large language model. The large language model can be either the Doubao model or the DeepSeek model, and this embodiment does not make any specific limitations on it.

[0043] Optionally, the first modal scheduling instruction includes a set of screen constraint parameters, which includes:

[0044] The intelligent agent responsible for generating the scene uses a unique storyboard identifier as a tracking token. When iteratively regenerating the video stream data of the same storyboard identifier, it ensures that the consistency offset of the main skeleton features and color distribution features between the previous and subsequent regenerations does not exceed a preset perception threshold. It also prevents abrupt changes in the appearance of the same storyboard data before and after iteration due to the regeneration operation. This can be understood as a cross-scene identifier binding constraint. The preset perception threshold can be limited according to actual conditions; this embodiment does not impose a specific limitation on it.

[0045] By using cross-camera identifier binding constraints and employing the unique identifier of each scene as a tracking token, the video stream data of the scene corresponding to the same unique identifier maintains the consistency of the main skeleton features and color distribution features before and after iterative regeneration.

[0046] The start and end frames of the storyboard video stream data output by the constraint image generation agent are aligned with the start and end timestamps of the storyboard duration data carried in the second and third modal scheduling instructions, and the total number of frames of the storyboard video stream data satisfies the following conditions. Where T represents the storyboard duration and fps represents the preset frame rate; this can be understood as: the above constraint is a multimodal boundary alignment constraint.

[0047] By using multimodal boundary alignment constraints, the start and end frames of the storyboard video stream data output by the image generation agent are aligned with the start and end timestamps of the storyboard duration data carried in the second and third modal scheduling instructions.

[0048] Spatiotemporal coherence constraints, such as requiring the scene-generating agent to maintain the continuity of optical flow and color consistency between adjacent frames when generating storyboard video stream data, and ensuring that the structural similarity index of adjacent frames is not lower than a preset threshold. This preset threshold can be 0.85.

[0049] By constraining the spatiotemporal coherence, the continuity of the optical flow field and the color consistency between adjacent frames in the storyboard video stream data generated by the image generation agent are kept above a preset threshold.

[0050] Optionally, the modal parameter mapping table includes the following mapping rules:

[0051] The content description dataset includes scene description data, character action description data, character dialogue text data, character identification data, and storyboard duration data;

[0052] Map scene description data and character action description data to the screen description sub-dataset; map character dialogue text data and character identity data to the dubbing description sub-dataset; map character dialogue text data and storyboard duration data to the subtitle description sub-dataset;

[0053] Furthermore, there is a cross-reference relationship between the screen description sub-dataset, the dubbing description sub-dataset, and the subtitle description sub-dataset: when the character dialogue text data in the dubbing description sub-dataset changes, the subtitle description sub-dataset automatically synchronizes and obtains the changed character dialogue text data as input, while the screen description sub-dataset is not affected by the change in character dialogue text data.

[0054] By explicitly mapping each data point in the modality parameter mapping table to the target generation modality, the content description dataset is precisely split by modality during the input stage. Each generation agent only receives a pure description subset belonging to its modality. Specifically, the modality parameter mapping table specifies that scene description data and character action description data belong to the screen description subset, character dialogue text data and character identity data belong to the dubbing description subset, and character dialogue text data and storyboard duration data belong to the subtitle description subset. Based on this mapping relationship, in step S2, the content description dataset of each storyboard data is split into screen description subset, dubbing description subset, and subtitle description subset, and first modality scheduling instructions, second modality scheduling instructions, and third modality scheduling instructions are constructed and sent to the corresponding screen generation agent, dubbing generation agent, and subtitle generation agent, respectively. The screen generation agent, dubbing generation agent, and subtitle generation agent obtain screen description subsets, dubbing description subsets, and subtitle description subsets that precisely match their own generation tasks during the input stage, eliminating data mismatch and intermodal data crosstalk from the source.

[0055] By leveraging cross-referencing relationships among the scene description, voice-over description, and subtitle description datasets, automatic synchronization of character dialogue text data is achieved between the voice-over and subtitle description datasets, while maintaining the independence and isolation of the scene description dataset. Specifically, in the modal parameter mapping table, character dialogue text data is mapped to both the voice-over and subtitle description datasets, and cross-referencing relationships exist between them. When the character dialogue text data in the voice-over description dataset changes, the subtitle description dataset automatically synchronizes and obtains the changed character dialogue text data as input, while the scene description dataset remains unaffected by the changes. This mechanism ensures that the voice-over and subtitle description datasets are always generated based on the same character dialogue text data, eliminating inconsistencies between voice-over and subtitle text caused by separate maintenance. Simultaneously, the scene description dataset remains isolated from the voice-over / subtitle link, and changes to scene description data and character action description data do not trigger the regeneration of the radio frequency description subset or the subtitle description dataset.

[0056] Third, by defining the modality parameter mapping table once at the input stage, the subsequent modality rescheduling signal accurately guides the target modality. Specifically, the modality parameter mapping table established in step S1 completes the attribution of each data point to the target generated modality at the input stage of the generation process. In step S4, the modality rescheduling signal carries an identifier of the modality type to be replaced, which directly corresponds to the screen modality, dubbing modality, or subtitle modality defined in the modality parameter mapping table. Based on the identifier of the modality type to be replaced, the system accurately locates the target modality stream data that matches the definition in the modality parameter mapping table from the single-scene multimodal data packet, and reroutes the updated description dataset to the corresponding target generating agent. This mechanism ensures that the data and modality mapping relationship at the input stage are consistent throughout the entire generation and rescheduling process, guaranteeing that the modality attribution in the initial generation and subsequent targeted recalculation remains consistent.

[0057] Optionally, the second modal scheduling instruction includes a set of dubbing constraint parameters, which includes:

[0058] Phono and identifier binding constraints, such as constraints that the voice-over generation agent retrieves and loads the phono embedding vector that is globally bound to the character's identity from a preset phono index library based on the character's identity identifier, and the same phono embedding vector is always maintained for the same character's identity identifier in all voice-over generation agents.

[0059] By binding timbre with identifier, the system retrieves and loads the timbre embedding vector that is globally bound to the character identifier from the preset timbre index library, using the character identifier as the index. This ensures that the same character identifier maintains the same timbre embedding vector throughout the calling process of all voice-over generation agents.

[0060] Flexible duration adaptation constraints, such as constraining the voice-over generation agent to use the storyboard duration data as the time limit, and ensuring that the actual duration of the output voice-over audio data meets the requirements. The audio stream data of the scene segment, where L represents the actual duration of the dubbing audio data, and Δ is the preset silence protection interval; when the base duration of the dubbing description subset calculated according to the base speech rate exceeds At this time, the voice-generating agent activates the prosodic scaling mechanism, and the actual output duration L is adjusted by the prosodic scaling factor γ. , This represents the prosodic scaling factor, and the range of values ​​for γ is limited to . , Indicates the base duration for calculating the base speech rate;

[0061] Through flexible duration adaptation constraints, the actual duration L of the storyboard audio stream data output by the voice-over generation agent is jointly constrained by the storyboard duration data T and the preset silence protection interval Δ; when the voice-over description subset dataset is calculated according to the baseline speech rate, the baseline duration is determined by the baseline duration. Exceed At that time, the prosodic stretching mechanism is activated. and will The range of values ​​is limited to .

[0062] Emotional feature mapping constraints, such as constraining the voice-over generation agent to extract the emotional tags corresponding to the storyboard from the structured storyboard script data, and select the fundamental frequency envelope, energy envelope and speech rate curve from the preset prosodic parameter space according to the emotional tags, so that the prosodic features of the storyboard audio stream data match the emotional tags.

[0063] By using emotion feature mapping constraints, the voice-over generation agent extracts emotion tags from the structured storyboard data and selects the corresponding fundamental frequency envelope, energy envelope, and speech rate curve from the preset prosodic parameter space based on the emotion tags, so that the prosodic features of the storyboard audio stream data match the emotion tags.

[0064] Optionally, the third mode scheduling instructions include:

[0065] The storyboard subtitle stream data includes at least two subtitle data sets, each of which includes the subtitle presentation duration and the number of subtitle characters.

[0066] The start and end timestamps of each subtitle data are constrained to fall within the range of the start and end timestamps of the target storyboard data, and the subtitle presentation duration is not less than the preset minimum subtitle duration (e.g., 0.8 seconds). The subtitle data is segmented according to the semantic punctuation of the character dialogue text data, and the number of characters in each subtitle data does not exceed the preset maximum number of subtitle characters (e.g., 20 characters).

[0067] Step S3: Input the scene description subset and scene duration data into the scene generation agent to generate scene video stream data; input the dubbing description subset and scene duration data into the dubbing generation agent to generate scene audio stream data; input the subtitle description subset and scene duration data into the subtitle generation agent to generate scene subtitle stream data; receive the scene video stream data, scene audio stream data, and scene subtitle stream data, and associate and bind the scene video stream data, scene audio stream data, and scene subtitle stream data with the unique identifier of the scene as the association key to generate a single scene multimodal data package;

[0068] After receiving the storyboard video stream data, storyboard audio stream data, and storyboard subtitle stream data, the storyboard video stream data, storyboard audio stream data, and storyboard subtitle stream data are bound into a single storyboard multimodal data packet using the unique identifier of the storyboard as the association key. This enables the physical aggregation of the storyboard video stream data, storyboard audio stream data, and storyboard subtitle stream data of the same storyboard data at the data level.

[0069] Step S4: Receive a modality rescheduling signal, which includes: target segment identifier, modality type identifier to be replaced, and update description dataset; locate the modality stream to be updated in the single segment multimodal data packet based on the modality type identifier to be replaced; send the update description dataset to the generating agent corresponding to the modality type identifier to be replaced to obtain the generated update modality stream; overwrite the update modality stream to the storage location pointed to by the modality type identifier to be replaced in the single segment multimodal data packet, using the target segment identifier as the primary key, while keeping the non-target modality stream data in the single segment multimodal data packet unchanged.

[0070] Specifically, a modal rescheduling signal is received, which includes a target storyboard identifier, a modal type identifier to be replaced, and an updated description dataset (e.g., the user specifies "change the scene of the 3rd storyboard from 'washing fish' to 'stir-frying'"). Based on the modal type identifier to be replaced, the modal stream data to be replaced is located from the single-storyboard multimodal data packet (e.g., the scene modal stream data is located), and the updated description dataset is rerouted to the target generating agent corresponding to the modal type identifier to be replaced (e.g., the description of "stir-frying" is routed to the scene generating agent), triggering the target generating agent to generate the updated modal stream data. Using the target storyboard identifier as the key, the updated modal stream data is backfilled into the single-storyboard multimodal data packet to replace the modal stream data to be replaced, while keeping other modal stream data in the single-storyboard multimodal data packet (such as storyboard audio stream data and storyboard subtitle stream data) unchanged.

[0071] Compared with existing technologies, the multimodal generation and scheduling method based on storyboard-driven approach provided in this embodiment achieves at least the following beneficial effects:

[0072] First, a modal parameter mapping table is used to pre-divide input data by modality before generation, ensuring that different generating agents only receive their respective pure description subsets, thus eliminating cross-modal data crosstalk. Specifically, the modal parameter mapping table established in step S1 specifies the explicit mapping relationship between each data in the content description dataset and the target generating modality at the input stage. In step S2, based on this modal parameter mapping table, the content description dataset of each storyboard data is split into a scene description subset, a dubbing description subset, and a subtitle description subset, and first, second, and third modal scheduling instructions are constructed and sent to the corresponding scene generating agent, dubbing generating agent, and subtitle generating agent, respectively. This ensures that different data in the same natural language description text are explicitly assigned to the corresponding target generating modality before entering the generation process. The scene generating agent, dubbing generating agent, and subtitle generating agent only receive the scene description subset, dubbing description subset, and subtitle description subset required for their generation tasks, fundamentally eliminating cross-modal data coupling and crosstalk at the input stage.

[0073] Second, by using the unique identifier of the scene as the association key, the three modal stream data generated independently in parallel are bound into a single scene multimodal data package, realizing the physical aggregation of the full modal data of the same scene at the output end;

[0074] In step S2, the first, second, and third modal scheduling instructions are sent in parallel to the scene generation agent, the dubbing generation agent, and the subtitle generation agent, enabling each agent to independently and concurrently execute its generation task without waiting for the others. In step S3, after receiving the storyboard video stream data, storyboard audio stream data, and storyboard subtitle stream data, the three are associated and bound into a single-storyboard multimodal data package using the unique storyboard identifier as the association key. This allows the storyboard video stream data, storyboard audio stream data, and storyboard subtitle stream data of the same storyboard data to be physically aggregated at the output end using the unique storyboard identifier as the key. Subsequent editing operations can use the unique storyboard identifier as an index to locate and manipulate all modal data of the same storyboard data at once.

[0075] Third, by using the modality type identifier to be replaced in the modality rescheduling signal, the precise location, directional routing and directional replacement of single modality stream data in single-scene multimodal data packets are realized, while keeping other modality stream data intact.

[0076] In step S4, a modality rescheduling signal is received. This signal includes a target scene identifier, a modality type identifier to be replaced, and an updated description dataset. Based on the modality type identifier to be replaced, the modality stream data to be replaced is located from the single-scene multimodal data packet. The updated description dataset is then rerouted to the target generating agent corresponding to the modality type identifier to be replaced, triggering the agent to generate the updated modality stream data. Using the target scene identifier as the key, the updated modality stream data is backfilled into the single-scene multimodal data packet to replace the modality stream data to be replaced, while keeping other modality stream data in the single-scene multimodal data packet unchanged. This achieves precise backflow of modified content to the corresponding generating agent, triggering only a single target generating agent to recalculate a single modality for a single scene. The computational cost of a single modification is compressed to the generation cost of a single modality, and other modality stream data does not need to be regenerated.

[0077] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of the equivalent elements of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.

[0078] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multimodal generation and scheduling method based on storyboard-driven modeling, characterized in that, Includes the following steps: Natural language description text and scene constraints are input into a large language model to generate structured scene script data. The structured scene script data contains scene sequences, each of which includes at least two scene data sets. Each scene data set corresponds to a unique scene identifier, scene duration data, content description dataset, and modality parameter mapping table. The modality parameter mapping table is used to route different data in the content description dataset to different target generation modalities. The modal parameter mapping table is parsed, and based on the mapping relationship between each data in the modal parameter mapping table and the target generation modality, the content description dataset of each storyboard data is split into a scene description subset, a dubbing description subset, and a subtitle description subset. A first modal scheduling instruction is constructed based on the scene description subset and sent to the scene generation agent. A second modal scheduling instruction is constructed based on the dubbing description subset and sent to the dubbing generation agent. A third modal scheduling instruction is constructed based on the subtitle description subset and sent to the subtitle generation agent. The first, second, and third modal scheduling instructions all include the unique identifier of the storyboard and the storyboard duration data. The scene description subset and the scene duration data are input into the scene generation agent to generate scene video stream data; the dubbing description subset and the scene duration data are input into the dubbing generation agent to generate scene audio stream data; the subtitle description subset and the scene duration data are input into the subtitle generation agent to generate scene subtitle stream data; the scene video stream data, scene audio stream data, and scene subtitle stream data are received, and the scene video stream data, scene audio stream data, and scene subtitle stream data are associated and bound using the scene unique identifier as the association key to generate a single scene multimodal data package; Receive a modality rescheduling signal, the modality rescheduling signal including: target segment identifier, modality type identifier to be replaced, and updated description dataset; Based on the modality type identifier to be replaced, locate the modality stream to be updated in the single-scene multimodal data packet; send the updated description dataset to the generating agent corresponding to the modality type identifier to be replaced to obtain the generated updated modality stream; using the target scene identifier as the primary key, overwrite the updated modality stream to the storage location pointed to by the modality type identifier to be replaced in the single-scene multimodal data packet, and keep the non-target modality stream data in the single-scene multimodal data packet unchanged.

2. The multimodal generation and scheduling method based on storyboard-driven approach according to claim 1, characterized in that, The modal parameter mapping table contains the following mapping rules: The content description dataset includes scene description data, character action description data, character dialogue text data, character identity identification data, and storyboard duration data; The scene description data and the character action description data are mapped to the screen description sub-dataset; the character dialogue text data and the character identity data are mapped to the dubbing description sub-dataset; the character dialogue text data and the storyboard duration data are mapped to the subtitle description sub-dataset. Furthermore, there is a cross-reference relationship between the screen description sub-dataset, the dubbing description sub-dataset, and the subtitle description sub-dataset: when the character dialogue text data in the dubbing description sub-dataset changes, the subtitle description sub-dataset automatically synchronizes and obtains the changed character dialogue text data as input, while the screen description sub-dataset is not affected by the change in the character dialogue text data.

3. The multimodal generation and scheduling method based on storyboard-driven approach according to claim 2, characterized in that, The third mode scheduling instruction includes: The subtitle stream data includes at least two subtitle data sets, each of which includes subtitle presentation duration data and the number of subtitle characters; The start and end timestamps of each subtitle data are constrained to fall within the range of the start and end timestamps of the target storyboard data, and the subtitle presentation duration is not less than the preset minimum subtitle duration; the subtitle data is segmented according to the semantic punctuation of the character dialogue text data, and the number of characters in each subtitle data does not exceed the preset maximum number of subtitle characters.

4. The multimodal generation and scheduling method based on storyboard-driven approach according to claim 1, characterized in that, The first modal scheduling instruction includes a set of screen constraint parameters, which includes: The intelligent agent for generating the scene is constrained to use the unique identifier of the scene as a tracking token. When iteratively regenerating the video stream data of the scene corresponding to the same unique identifier of the scene, the consistency offset of the main skeleton features and color distribution features between the previous generation and the subsequent regeneration shall not exceed a preset perception threshold. The appearance of the same scene data before and after the iteration shall not be abruptly changed due to the regeneration operation. The start and end frames of the storyboard video stream data output by the image generation agent are aligned with the start and end timestamps of the storyboard duration data carried in the second and third modal scheduling instructions, and the total number of frames of the storyboard video stream data satisfies the following conditions: Where T is the segment duration data and fps is the preset frame rate; The intelligent agent generating the scene is constrained to maintain the continuity of the optical flow field and the consistency of color between adjacent frames when generating the storyboard video stream data, and the structural similarity index of adjacent frames is not lower than a preset threshold.

5. The multimodal generation and scheduling method based on storyboard-driven approach according to claim 1, characterized in that, The second modal scheduling instruction includes a set of dubbing constraint parameters, which includes: The voice-generating agent is constrained to retrieve and load a timbre embedding vector globally bound to the character identity from a preset timbre index library based on the character identity identifier. The same character identity identifier always maintains the same timbre embedding vector throughout all calls to the voice-generating agent. The voice-generating agent is constrained to use the storyboard duration data as the upper time limit, and the actual duration of the output voice-over audio data satisfies the following conditions: The segmentation audio stream data, where L represents the actual duration of the dubbing audio data, and Δ is the preset silence protection interval; when the base duration of the dubbing description subset calculated according to the base speech rate exceeds At that time, the voice-generating agent initiates a prosodic scaling mechanism, and the actual output duration L is adjusted by a prosodic scaling factor γ. , This represents the prosodic scaling factor, and the range of values ​​for γ is limited to . , Indicates the base duration for calculating the base speech rate; The voice-over generation agent is constrained to extract the emotional tags corresponding to the storyboard data from the structured storyboard script data, and select the fundamental frequency envelope, energy envelope and speech rate curve from the preset prosodic parameter space according to the emotional tags, so that the prosodic features of the storyboard audio stream data match the emotional tags.

Citation Information

Patent Citations

  • An AI intelligent short video generation method and system based on multi-agent collaboration

    CN120547420B