Image generation system adapted for large model streaming

CN122820874APending Publication Date: 2026-09-25ZHONGJIN SHUYOU TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611338169.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-31
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0002]大语言模型常与图像生成模型协同处理多模态内容创建任务,模型通常会输出包含实体描述、内容片段及全局约束的结构化文本来指导下游的图像生成;为处理这些跨模态生成任务,现有方案普遍采用同步阻塞架构,即通过调用服务等待大语言模型完成全量文本输出,在接收到结束标识后,再对完整的文档结构进行统一解析、提取节点属性,进而触发对应的图像生成模块;虽然此方案在短文本或单次图文生成任务下具备一定处理能力,但由于其高度依赖文本流的完整终结状态及全量解析机制,且缺乏对输出片段的流式识别与前置处理,造成整体处理链路冗长、端到端响应延迟高、前置大模型输出期间后置计算资源利用率低,难以支撑长篇内容的并发处理与图像任务的低延迟渲染

Benefits of technology

[0046]1.本发明针对现有技术依赖全量文本输出导致的同步阻塞和响应延迟高的问题,通过流式可扩展标记语言解析模块在接收增量文本片段时实时识别已闭合可扩展标记语言节点,并转换为结构化节点数据,同时将未闭合的可扩展标记语言片段保留在文本缓冲区以进行拼接;结合任务分发模块将图像任务输入参数构造为异步图像生成任务进行发送,该机制无需等待大语言模型输出结束标识即可前置启动图像任务;所述机制避免了后置计算资源未被充分利用的问题,大幅提升了流式情境下节点解析与异步分发的及时性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122820874A_ABST
    Figure CN122820874A_ABST
Patent Text Reader

Abstract

The application relates to the field of artificial intelligence large model application and image generation technology, in particular to an image generation system adaptive to large model streaming output, which comprises text processing, model prompt word construction, large language model, streaming extensible markup language analysis, node type determination, idempotent deduplication, prompt word rendering, task distribution, image generation and result processing modules; the system receives to-be-processed text, constructs preset extensible markup language format prompt words and submits the same to large model streaming output; through a streaming extensible markup language analysis module, a closed node is identified in real time and converted into structured data, and in combination with an idempotent deduplication and an asynchronous task distribution mechanism, an image generation task is started in advance before a receiving end identifier; the application does not need to wait for full-text output, effectively avoids the problem that computing resources are not fully utilized, and improves the response efficiency of node analysis and asynchronous distribution in a streaming context.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence large model application and image generation technology, specifically to an image generation system adapted to the streaming output of large models. Background Technology

[0002] Large language models often collaborate with image generation models to handle multimodal content creation tasks. The models typically output structured text containing entity descriptions, content fragments, and global constraints to guide downstream image generation. To handle these cross-modal generation tasks, existing solutions generally adopt a synchronous blocking architecture. This involves calling a service to wait for the large language model to complete the full text output. After receiving the end marker, the complete document structure is uniformly parsed, node attributes are extracted, and the corresponding image generation module is triggered. Although this solution has certain processing capabilities for short text or single-time image-text generation tasks, it relies heavily on the complete final state of the text stream and the full parsing mechanism. Furthermore, it lacks streaming recognition and pre-processing of output fragments, resulting in a lengthy overall processing chain, high end-to-end response latency, and low utilization of post-processing computing resources during the output of the large model. This makes it difficult to support concurrent processing of long content and low-latency rendering of image tasks.

[0003] Therefore, improving the real-time performance of structured node parsing and the timeliness of asynchronous image task distribution during the streaming output of large models has become an urgent technical problem to be solved. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides an image generation system adapted for large-scale model streaming output. Specifically, the technical solution of this invention includes:

[0005] The system includes a text processing module, a model prompt word construction module, a large language model module, a streaming scalable markup language parsing module, a node type determination module, an idempotent deduplication module, a prompt word rendering module, a task distribution module, an image generation module, and a result processing module.

[0006] The text processing module receives the task of generating the text to be processed and maintains the text buffer and the set of processed nodes;

[0007] The model prompt word construction module constructs prompt words based on the generated task, and the constraint large language model module outputs them in a preset extensible markup language format.

[0008] The text processing module inputs the prompt word into the large language model module and streams the incremental text fragments and end markers.

[0009] The streaming extensible markup language parsing module appends the incremental text fragment to the text buffer for parsing, converts the parsed closed nodes into structured node data containing node identification information, and retains unclosed fragments to splice subsequent incremental fragments.

[0010] The node type determination module divides the structured node data into entity description nodes, content fragment nodes, and global constraint nodes, and extracts the corresponding attribute information.

[0011] The idempotent deduplication module compares the structured node data with the set of processed nodes based on the node identifier information. If the node identifier information is not found in the set, it is added.

[0012] The prompt word rendering module renders the image task input parameters based on the extracted attribute information;

[0013] The task distribution module constructs the input parameters into an asynchronous task and sends it to the image generation module to generate the target image;

[0014] After obtaining the end marker, the result processing module summarizes the target image to obtain the processing result.

[0015] As a further aspect of the present invention: the text processing module, the model prompt word construction module, and the large language model module are specifically configured as follows:

[0016] The text processing module receives a generation task containing business data, which includes at least one of the following: text to be processed, a list of content fragments, style identifiers, image ratios, entity description information, and contextual constraint information. The task type of the generation task includes at least one of the following: entity description extraction task, content fragment image generation task, global constraint extraction task, and image redrawing task.

[0017] The model prompt word construction module reads the corresponding preset prompt word template according to the task type, and defines an Extensible Markup Language (XML) output format in the preset prompt word template. The XML output format requires that the entity description node include at least one sub-node of entity identifier, entity category, entity visual attribute description, style trigger information, style parameter, and description text; the content fragment node includes at least one sub-node of fragment identifier, entity reference, generated prompt information, style parameter, style trigger information, and reference description; and the global constraint node includes at least one sub-node of style constraint, context constraint, and content list.

[0018] The text processing module inputs the prompt words into the large language model module in a streaming mode, and initializes the text buffer and the set of processed nodes. Whenever the large language model module returns an incremental text fragment, the text processing module appends the incremental text fragment to the text buffer to form the current cumulative output text.

[0019] As a further aspect of the present invention: the streaming extensible markup language parsing module is specifically configured as follows:

[0020] The streaming Extensible Markup Language (EXPLAIN) parsing module identifies text segments with matching start and end tags in the currently accumulated output text according to preset parsing rules, wherein the start and end tags have the same node name;

[0021] When an Extensible Markup Language (XML) fragment is detected to contain only a start tag but no corresponding end tag is received, the XML fragment is determined to be an unclosed XML fragment and is kept in the text buffer to wait for it to be concatenated with the subsequent incremental text fragments.

[0022] When an Extensible Markup Language (XML) fragment is detected to contain matching start and end tags, the XML fragment is determined to be a closed XML node, and its node content is further checked to see if it contains child XML nodes. If it contains child XML nodes, the node content is recursively parsed. If it does not contain child XML nodes, the node content is used as the leaf field value.

[0023] The streaming extensible markup language parsing module uses a preset depth-first algorithm to traverse the node tree hierarchy, merging nodes with the same name: when there are multiple closed extensible markup language nodes with the same name in the same level structure, the multiple node results are organized into a list or array according to the order of appearance, and the corresponding structured node data is output.

[0024] As a further aspect of the present invention: when the structured node data is the entity description node, the specific processing steps include:

[0025] The node type determination module extracts entity identifiers, style trigger information, style parameters, and entity visual attribute descriptions from the structured node data as the corresponding attribute information; the idempotent deduplication module determines whether the entity identifier already exists in the processed node set. If it already exists, the entity description node is skipped; if it does not exist, the entity identifier is added to the processed node set, and a corresponding subtask identifier is generated based on the entity identifier and the generation task.

[0026] The prompt word rendering module merges the extracted style trigger information, style parameters, and entity visual attribute descriptions to render and generate the main image task input parameters, and records the association between the generation task and the main image task; the task distribution module sends the asynchronous image generation task containing the main image task input parameters to the image generation module.

[0027] As a further aspect of the present invention: when the structured node data is the content fragment node, the specific processing steps include:

[0028] The node type determination module extracts fragment identifiers, entity references, generated prompt information, style parameters, style trigger information, and reference descriptions from the structured node data as the corresponding attribute information;

[0029] The idempotent deduplication module determines whether the fragment identifier already exists in the processed node set. If it already exists, the content fragment node is skipped; if it does not exist, the fragment identifier is added to the processed node set.

[0030] The prompt word rendering module merges and extracts the fragment identifier, entity reference, generated prompt information, style parameters, style trigger information, and reference description, and renders the content fragment image task input parameters.

[0031] The task distribution module sends an asynchronous image generation task containing the input parameters of the content fragment image task to the image generation module, and the result processing module records the task status, start time, and the input parameters of the content fragment image task.

[0032] As a further aspect of the present invention: the system further includes a global constraint and entity context extraction module, which is specifically configured as follows:

[0033] Context missing judgment: When the received generation task does not carry context constraint information or entity description information, the global constraint and entity context extraction module generates supplementary prompt words through the model prompt word construction module, and calls the large language model module to generate an incremental text fragment containing the global constraint node and the entity description node;

[0034] Context parsing: The streaming extensible markup language parsing module parses the global constraint node and entity description node to obtain a list of global context constraints and entity descriptions;

[0035] Unified reference: When generating the image task input parameters for each content fragment, the prompt word rendering module uniformly references the global context constraints and entity description list.

[0036] As a further aspect of the present invention: the system further includes a long-content concurrent processing module, which is specifically configured as follows:

[0037] Content batching: For the generation task in which the text to be processed contains multiple content fragments, the text processing module divides the text to be processed into multiple paragraph batches according to a preset content fragment number threshold. Each paragraph batch carries its start and end positions in the complete content.

[0038] And this occurs as follows: each paragraph batch calls the large language model module to stream generate the corresponding content fragment nodes;

[0039] Concurrent processing: Multiple paragraph batches are processed through concurrent tasks, including streaming scalable markup language parsing, node deduplication, prompt word rendering, and the distribution of the asynchronous tasks.

[0040] As a further aspect of the present invention, the result processing module and the task distribution module are also configured to perform the following exception handling:

[0041] Results Summary: After the large language model module streams the end marker, the result processing module summarizes the target images corresponding to the processed structured node data and outputs the processing results.

[0042] Anomaly detection: When network anomalies, model anomalies, or task distribution anomalies occur during processing, the result processing module records the anomaly information and the current task status;

[0043] Retry on failure: The task distribution module will resubmit the asynchronous task that has encountered an error to the waiting queue according to the preset number of task retries, or write failure information to the preset result queue.

[0044] As a further aspect of the present invention: the task distribution module and the image generation module communicate through an asynchronous communication mechanism; the image generation module is configured with at least one of a text-to-image generation model, an image-to-image generation model, a subject generation model, an illustration generation model, or a video keyframe generation model; the processed node set uses an idempotent deduplication table to record the node identification information that has been processed.

[0045] The present invention has the following beneficial effects:

[0046] 1. This invention addresses the problems of synchronous blocking and high response latency caused by existing technologies that rely on full text output. It utilizes a streaming Scalable Markup Language (SML) parsing module to identify closed SML nodes in real time when receiving incremental text fragments and convert them into structured node data. Simultaneously, unclosed SML fragments are retained in a text buffer for concatenation. Combined with a task distribution module, image task input parameters are constructed into asynchronous image generation tasks for transmission. This mechanism allows image tasks to be started beforehand without waiting for the large language model to output an end marker. This mechanism avoids the problem of underutilization of post-processing computing resources and significantly improves the timeliness of node parsing and asynchronous distribution in streaming scenarios.

[0047] 2. This invention addresses the problems of excessively long processing times for long text generation tasks and the tendency for streaming appending to cause duplicate rendering. The long text concurrent processing module of this invention divides the text to be processed into multiple paragraph batches, with each batch concurrently performing streaming parsing and task distribution. Simultaneously, the system employs an idempotent deduplication module to compare node identifier information and add the identifier of new nodes to the set of processed nodes. This mechanism not only achieves low-latency rendering of long text content but also ensures that when re-parsing incremental text fragments, asynchronous image generation tasks are not repeatedly created for already processed structured node data, effectively guaranteeing concurrency efficiency and resource utilization. Attached Figure Description

[0048] Figure 1 This is a structural diagram of the system of the present invention. Detailed Implementation

[0049] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0050] like Figure 1 As shown, the image generation system adapted for large model streaming output includes:

[0051] The system includes a text processing module, a model prompt word construction module, a large language model module, a streaming scalable markup language parsing module, a node type determination module, an idempotent deduplication module, a prompt word rendering module, a task distribution module, an image generation module, and a result processing module.

[0052] The text processing module receives the task of generating the text to be processed and maintains the text buffer and the set of processed nodes;

[0053] The model prompt word construction module constructs prompt words based on the generated task, and the constraint large language model module outputs prompt words in a preset extensible markup language format.

[0054] The text processing module inputs the prompt word into the large language model module and streams the incremental text fragments and end markers.

[0055] The streaming extensible markup language parsing module appends the incremental text fragment to the text buffer for parsing, converts the parsed closed nodes into structured node data containing node identification information, and retains unclosed fragments to splice subsequent incremental fragments;

[0056] The node type determination module divides the structured node data into entity description nodes, content fragment nodes, and global constraint nodes, and extracts the corresponding attribute information.

[0057] The idempotent deduplication module compares the structured node data with the set of processed nodes based on the node identifier information. If the node identifier information is not found in the set, it is added.

[0058] The prompt word rendering module renders the image task input parameters based on the extracted attribute information;

[0059] The task distribution module constructs the input parameters into an asynchronous task and sends it to the image generation module to generate the target image;

[0060] After obtaining the end marker, the result processing module summarizes the target image to obtain the processing result.

[0061] The text processing module, the model prompt word construction module, and the large language model module are specifically configured as follows: The generation task received by the text processing module contains business data, which includes at least one of the following: text to be processed, a list of content fragments, style identifiers, image ratios, entity description information, and contextual constraint information. The task type of the generation task includes at least one of the following: entity description extraction task, content fragment image generation task, global constraint extraction task, and image redrawing task.

[0062] The model prompt word construction module reads the corresponding preset prompt word template according to the task type, and defines the Extensible Markup Language (XML) output format in the preset prompt word template. The XML output format requires that the entity description node include at least one of the following child nodes: entity identifier, entity category, entity visual attribute description, style trigger information, style parameter, and description text; the content fragment node includes at least one of the following child nodes: fragment identifier, entity reference, generated prompt information, style parameter, style trigger information, and reference description; and the global constraint node includes at least one of the following child nodes: style constraint, context constraint, and content list.

[0063] The text processing module inputs prompt words into the large language model module in a streaming mode and initializes the text buffer and the set of processed nodes. Whenever the large language model module returns an incremental text fragment, the text processing module appends the incremental text fragment to the text buffer to form the current cumulative output text.

[0064] The streaming Scalable Markup Language (SML) parsing module is specifically configured to: in the currently accumulated output text, the streaming SML parsing module identifies text segments with matching start tags and end tags according to preset parsing rules, wherein the start tags and end tags have the same node name;

[0065] When an Extensible Markup Language (XML) fragment is detected to contain only a start tag but no corresponding end tag is received, the XML fragment is determined to be an unclosed XML fragment and is kept in the text buffer to wait for it to be concatenated with subsequent incremental text fragments.

[0066] When an Extensible Markup Language (XML) fragment is detected to contain matching start and end tags, the XML fragment is determined to be a closed XML node, and its node content is further checked to see if it contains child XML nodes. If it contains child XML nodes, the node content is recursively parsed. If it does not contain child XML nodes, the node content is used as the leaf field value.

[0067] The streaming Extensible Markup Language (XML) parsing module uses a preset depth-first algorithm to traverse the node tree hierarchy, merging nodes with the same name: when there are multiple closed XML nodes with the same name in the same level structure, the results of multiple nodes are organized into a list or array according to their order of appearance, and the corresponding structured node data is output.

[0068] In the application environment: The task submission end submits a text-to-image task, which carries the text to be processed and can optionally carry at least one of the following: a list of content fragments, style identifiers, image ratios, entity description information, and contextual constraint information; the system is used to parse the fully closed extensible markup language nodes before the large language model has completed all outputs, and to convert the directly usable node content into an image generation task, so as to reduce the situation where the image task must wait for the complete response to be completed before it can be started.

[0069] After receiving the generation task, the text processing module establishes an independent text buffer and a set of processed nodes for the task. The text buffer is used to store the current cumulative output text that has not yet been parsed or has not yet been removed from the parsing range, and the set of processed nodes is used to record the node identification information that has entered subsequent business processing.

[0070] The model prompt word construction module reads the corresponding preset prompt word template according to the task type, and limits the output of the large language model to adopt the preset extensible markup language format in the template; the extensible markup language format constrains at least three types of nodes, namely entity description nodes, content fragment nodes and global constraint nodes;

[0071] The entity description node can contain entity identifier, entity category, visual attributes, and descriptive text; the content fragment node can contain fragment identifier, entity reference, generated prompt information, and reference description; the global constraint node can contain style constraints, context constraints, and content list; the text processing module submits the generated prompt words to the large language model module in streaming mode, and appends each incremental text fragment to the text buffer when receiving each incremental text fragment to form the current cumulative output text;

[0072] As the cumulative output text continues to expand, the streaming Extensible Markup Language (XML) parsing module does not require the current text to constitute a complete XML document. Instead, it only scans the text for start and end tags that match the tag names according to preset parsing rules.

[0073] The preset parsing rules include at least the following: Firstly, the node names of entity description nodes, content fragment nodes, and global constraint nodes are prioritized as primary identification objects; upon detecting a start tag, a search is conducted for an end tag with the same name from the start tag position onwards, and the fragment that first forms a complete closed relationship and has a balanced tag hierarchy is selected as a candidate parsing fragment; if the candidate parsing fragment contains a lower-level node with the same name, parsing continues inwards according to the tag pairing hierarchy until a complete closed node is obtained.

[0074] Tag level balancing refers to counting the net difference between the start and end tags with the same name from front to back within the same candidate segment. The difference returns to zero for the first time at the end of the segment, thus determining that the candidate segment has been closed. The parsing module maintains the parsing position in the order from front to back.

[0075] For closed Extensible Markup Language (XML) nodes that have completed structured transformation and node identification records, their corresponding read text is removed from the subsequent scanning range or marked as read and processed. Only unclosed XML fragments and text that have not yet formed complete nodes are retained. Subsequent incremental text fragments can be directly concatenated with the retained content. Repeated scanning will not change the boundaries of unclosed fragments.

[0076] The text buffer is maintained using a combination of sequential appending and front-end pruning. After each round of scanning, the parsing module divides the currently accumulated output text into three logical states: processed area, pending confirmation area, and reserved area. The processed area consists of text parts that have formed closed Extensible Markup Language (XML) nodes and whose corresponding node identification information has been written into the processed node set. The pending confirmation area consists of text parts that have started tags but have not yet formed a complete closed relationship. The reserved area consists of text parts that are newly appended after the pending confirmation area but have not yet participated in this round of scanning.

[0077] Only when a text portion simultaneously meets the three conditions of completing closure, completing structured transformation, and completing node identifier registration will it be clipped and removed from the front of the text buffer; if it has only completed closure but its internal child nodes have not yet been recursively parsed, it will be temporarily kept in the pending confirmation area and processed first in the next round of scanning.

[0078] Within the same task, the system sets the maximum scan length and the maximum number of closed nodes processed in a single round. When either upper limit is reached, the current round of scanning ends and the remaining text is retained to enter the next round; the next round continues processing from the starting position of the previous round.

[0079] The maximum scan length per round is selected from 2 to 5 times the average length of the three most recent consecutive incremental text segments of the current task, and is limited to no less than 1024 characters and no more than 16384 characters; the maximum number of closed nodes processed per round is 1 to 20.

[0080] The current task records the default value when it is created. During the operation, if there are still unprocessed closed nodes after two consecutive rounds of scanning reach the maximum scan length of a single round, the maximum scan length of a single round will be increased by 1.5 times the original value. If the calculated value after the increase exceeds 16384 characters, it will be forcibly truncated and the value will be taken as 16384 characters.

[0081] When the number of closed nodes is less than one-third of the maximum number of closed nodes processed in a single round after two consecutive rounds of scanning, or when the number of closed nodes is less than one-third of the maximum number of closed nodes processed in a single round. When one-third of the maximum number of closed nodes processed in a single round is processed and there are no unprocessed closed nodes remaining, the original value is maintained; this provides a clear basis for the upper limit of scanning, while keeping the processing cycle within the same task within the preset time range;

[0082] The preset parsing rules are implemented using a deterministic sequential scanning method, which includes: setting the current cumulative output text as the target string and the scanning starting point as the target pointer; first, starting from the target pointer position, searching for the earliest starting label in the first-level recognition object, and recording it as the starting position;

[0083] Then, scan character by character from the starting position. Each time a starting tag with the same name as the starting tag is encountered, the level count is incremented by 1. Each time an ending tag with the same name is encountered, the level count is decremented by 1. When the level count first decreases to zero, the substring from the starting position to the end of the ending tag is identified as a closed Extensible Markup Language (XML) node. If the level count is still greater than zero when the current cumulative output text is scanned to the end, the remaining substring starting from the starting position is retained as an unclosed XML segment.

[0084] For a closed Extensible Markup Language (XML) node, the parsing module first records its parent node identifier, node level, and node path, and then determines whether it still contains preset XML tags. If it does, the same rules are recursively executed using the node content as a new string to be parsed. If it does not, the node content is directly written into the corresponding field value.

[0085] For multiple child nodes with the same name under the same parent node, the parsing module assigns sequence numbers according to the order of appearance when generating structured node data for the first time, and organizes them into a list or array according to the sequence numbers, so that subsequent deduplication and rendering processes are performed based on the determined node positions.

[0086] The parent node identifier refers to the node identifier information of the node directly above the current node; the node level refers to the integer value obtained by incrementing the level of the root node (1), the level of its direct child nodes (2), and so on; the node path refers to the path string that connects the current node in order of tag name starting from the root node, with adjacent levels connected by a fixed separator / ; the node content refers to the text value or set of child nodes retained between the start and end tags of the current node after removing the child node tag structure.

[0087] During structured transformation, if the current node does not contain child nodes, the node content is directly recorded as the text value after removing leading and trailing whitespace characters; if the current node contains child nodes, the node content is recorded as a set of fields arranged in the order of child nodes, and is no longer mixed with the leaf text; the source of values ​​and writing method of each node field are fixed in the parsing stage, and subsequent modules do not need to infer them again.

[0088] For an Extensible Markup Language (XML) segment that only detects a start tag but has not yet received a corresponding end tag, the parsing module identifies it as an unclosed segment and keeps it in the text buffer, waiting for subsequent incremental text segments to complete it; for an XML segment that has formed a complete start and end tag pair, the parsing module identifies it as a closed XML node and identifies the corresponding node content from the accumulated text.

[0089] If the node content still contains child Extensible Markup Language nodes, the node content is recursively parsed; if the node content no longer contains child nodes, this part of the content is written as a leaf field value into the structured node data; if there are multiple closed nodes with the same name in the same level, they are organized into a list or array according to the order of appearance so that the subsequent node type determination module can handle them uniformly.

[0090] For plain text that does not belong to a valid Extensible Markup Language node, the parsing module skips the text and does not write it into the closed node result; to avoid ambiguity when merging nodes with the same name at different levels, the merging of nodes with the same name is only performed under the condition that the parent node identifier is the same and the node level is the same; when the node names are the same but the parent node identifiers are different, or although they are under the same parent node but the node paths are different, they are retained as different fields and no list merging is performed;

[0091] Structured node data includes at least the node name, parent node identifier, node level, node path, node content, and node identifier information; among them, the node path is generated according to the label order from the root node to the current node, which is used to distinguish nodes with the same name under different levels or different parent nodes when merging nodes with the same name and deduplicating them later.

[0092] Node identification information is generated in a fixed order during the parsing phase: the entity identifier, fragment identifier, or other standard identifier fields explicitly given inside the node are read first; if no explicit identifier field exists, a supplementary identifier is generated by combining the task identifier, parent node identifier, node path, and same-level sequence number in sequence; the same-level sequence number is an integer obtained by incrementing from 1 according to the order of closure of nodes with the same name under the same parent node; even if adjacent incremental text fragments trigger repeated scanning, the same node identification information output by the parsing module remains consistent;

[0093] During the scanning process, the parsing module maintains anomaly markers and recursion depth counts. If an end tag is detected to appear before the corresponding start tag in the same segment to be parsed, or if the nesting level of tags is detected to exceed the preset upper limit consecutively, the segment is marked as an anomaly candidate segment and recursion is stopped. For the anomaly candidate segment, it is not directly discarded, but only the subsequent unaffected text is retained for scanning, and the anomaly candidate segment is kept at the end of the text buffer to wait for subsequent incremental text segments to be repaired.

[0094] The maximum recursion depth is set to 8 to 32 levels, and is determined at the start of the task by adding 2 to the maximum node level allowed in the preset prompt word template. If the maximum node level is not declared in the template, the default is 16 levels. If the abnormal candidate fragment has not been restored to a level balance state after several rounds of scanning, it will be switched to the skip state and will no longer be used as a source of resolvable nodes in the subsequent processing of this task, but it will not affect the continued extraction of normally closed nodes.

[0095] A series of consecutive scans can be 3 to 10 consecutive scans; the specific value is determined by the maximum scan length of a single scan in the current task. When the maximum scan length of a single scan does not exceed 4096 characters, 3 scans are used; when the maximum scan length of a single scan is greater than 4096 characters, 5 scans are used.

[0096] The routing directions for plain text, abnormal candidate segments, unclosed segments, and consumed text in the control flow are different: plain text is skipped directly, abnormal candidate segments are temporarily deferred from recursion, unclosed segments are retained for splicing, and consumed text is allowed to be trimmed and removed, thereby avoiding a single abnormal segment blocking the subsequent processing of the entire stream output;

[0097] After receiving the above structured node data, the node type determination module distinguishes entity description nodes, content fragment nodes, and global constraint nodes according to node names or preset mapping relationships, and extracts the corresponding node attribute information. The preset mapping relationship prioritizes the one-to-one correspondence table between node names and node types. When there are aliases for node names, the standard node names pre-registered in the prompt word template are used as the basis for determining the normalized node type.

[0098] The idempotent deduplication module compares node identifier information with the set of processed nodes. For node identifier information that has not yet been recorded in the set, the identifier is written into the set of processed nodes, and the node is allowed to enter the subsequent prompt word rendering and task distribution stages. For node identifier information that already exists in the set, the image generation task will not be created again.

[0099] For global constraint nodes in Extensible Markup Language that do not have a direct node identifier, the system generates stable node identifier information based on the task identifier, node type, and their order in the same task. The order is determined according to the order in which the node first forms a closure in the cumulative output text of the current task. If subsequent repeated parsing yields global constraint nodes of the same type and in the same order, the same node identifier information is used to perform deduplication.

[0100] For entity description nodes and content fragment nodes, if the entity identifier or fragment identifier is directly provided in the Extensible Markup Language, then the entity identifier or fragment identifier is used as the core field of the node identifier information, and combined with the task identifier and written into the processed node set; if multiple fields that can be used as identifiers appear in a node at the same time, the priority is as follows: entity identifier or fragment identifier, standard identifier field predefined in the prompt word template, and supplementary identifier field generated by the system according to the node path and the first closure order;

[0101] Node identification information is generated using reproducible concatenation rules: when a node is an entity description node, the node identification information is formed by concatenating the task identifier, constant string entity, and entity identifier in sequence; when a node is a content fragment node, the node identification information is formed by concatenating the task identifier, string fragment, and fragment identifier in sequence; when a node is a global constraint node and lacks an explicit identification field, the node identification information is formed by concatenating the task identifier, global string, node path, and first closure sequence number in sequence.

[0102] The initial closure sequence number starts from 1 and increments, remaining unchanged within the same task scope; the set of processed nodes stores the above node identification information using a key-value table or set table structure, and the query rule is exact matching: if the query matches, a duplicate processing flag is output and the subsequent process of the current node is terminated; if the query does not match, the node identification information is written and a processing allowed flag is output; based on this rule, even if the same incremental text fragment is repeatedly sent to the parsing module, as long as the node identification information does not change, the judgment result of the idempotent deduplication module will remain consistent;

[0103] The prompt word rendering module uses the node attribute information extracted by the node type determination module to form the image task input parameters. The task distribution module encapsulates these input parameters into an asynchronous image generation task and sends it to the image generation module via a task queue, event bus, task table, or other asynchronous communication mechanisms. The image generation module employs a text-to-image generation model based on a latent diffusion model. The model's network topology includes a contrastive language-image pre-trained text encoder and a U-shaped denoising network. The feature vector dimension of the input text is fixed at [dimensional value missing]. The loss function for model training is designed as follows:

[0104]

[0105] in, For loss function, Input parameters for the encoded image task. Noise predicted for the U-shaped network For the network parameters of the U-shaped network, For time steps Noisy latent variables at time; in the formula Represents the mathematical expectation. For the initial noiseless latent variables, To obtain from the standard normal distribution The actual noise in the mid-sample, express The square of the norm; It follows a Markov forward noise-adding process and satisfies:

[0106]

[0107] in, The preset noise scheduling cumulative parameters; The random time step is uniform sampling; the model generates target images according to the received task input parameters through this network structure; the result processing module continuously records the processing status of the started image task, and after obtaining the end marker output by the large language model module, it summarizes the target images generated in this task to form the final processing result.

[0108] The summary is based on the task status record, and only target images that have achieved successful results are included in the completed results set. For asynchronous image generation tasks that are still pending or being processed when the end marker is reached, the result processing module retains its task record and continues to receive status updates. After the task is successful or is determined to have failed according to the exception handling rules, the corresponding status is included in the final processing result. This ensures that the time difference between the end of the large language model output and the completion of the asynchronous image task will not cause result omission or misidentify ungenerated target images as completed.

[0109] If the cumulative output text in a certain round of parsing contains only unclosed tag segments at the end, these segments will not enter the node type determination and task distribution stage, but will remain in the text buffer. If an incremental text segment contains irrelevant text that does not constitute a valid node, the streaming extensible markup language parsing module will only identify segments that meet the paired tag conditions, and the rest will be skipped in subsequent scans without affecting the extraction of closed nodes. If the same node reappears due to repeated parsing of cumulative text, the idempotent deduplication module will directly intercept the duplicate processing request based on the already processed node set, thereby preventing the image generation module from repeatedly processing the same entity or the same content segment.

[0110] In this text-to-image task, when the large language model continuously outputs Scalable Markup Language content, the system does not wait for all responses to finish, but begins parsing and processing as soon as the first entity description node, content fragment node, or global constraint node is closed; the subsequent output of the large language model, streaming Scalable Markup Language parsing, asynchronous task delivery, and the execution of the image generation module can be carried out in parallel.

[0111] When the number of characters in the text to be processed exceeds the preset text length threshold, even if there are missing local labels or truncation at the end in the subsequent output, the closed and deduplicated nodes can still be used to create image generation tasks, and the previously completed execution process will not be interrupted due to the incompleteness of the reply tail.

[0112] When the structured node data is an entity description node, the specific processing steps include: the node type determination module extracts the entity identifier, style trigger information, style parameters and entity visual attribute description as corresponding attribute information from the structured node data; the idempotent deduplication module determines whether the entity identifier already exists in the processed node set. If it already exists, the entity description node is skipped. If it does not exist, the entity identifier is added to the processed node set, and a corresponding subtask identifier is generated based on the entity identifier and the generation task.

[0113] The prompt word rendering module merges the extracted style trigger information, style parameters, and entity visual attribute descriptions, renders the input parameters for the main image task, and records the association between the generation task and the main image task; the task distribution module sends the asynchronous image generation task containing the main image task input parameters to the image generation module.

[0114] When the structured node data is a content fragment node, it is specifically configured as follows: the node type determination module extracts fragment identifier, entity reference, generated prompt information, style parameters, style trigger information and reference description from the structured node data as corresponding attribute information;

[0115] The idempotent deduplication module determines whether the fragment identifier already exists in the set of processed nodes. If it already exists, the content fragment node is skipped; if it does not exist, the fragment identifier is added to the set of processed nodes.

[0116] The prompt word rendering module merges and extracts fragment identifiers, entity references, generated prompt information, style parameters, style trigger information, and reference descriptions, and renders the generated content fragment image task input parameters.

[0117] The task distribution module sends the asynchronous image generation task containing the content fragment image task input parameters to the image generation module, and the result processing module records the task status, start time, and content fragment image task input parameters.

[0118] The system also includes a global constraint and entity context extraction module. Specifically, the global constraint and entity context extraction module is configured as follows: context missing judgment: when the received generation task does not carry context constraint information or entity description information, the global constraint and entity context extraction module generates supplementary prompt words through the model prompt word construction module and calls the large language model module to generate incremental text fragments containing global constraint nodes and entity description nodes.

[0119] Context parsing: The streaming extensible markup language parsing module parses the global constraint nodes and entity description nodes to obtain a list of global context constraints and entity descriptions;

[0120] Unified referencing: When generating image task input parameters for each content fragment, the prompt word rendering module uniformly references the global context constraints and entity description list.

[0121] In the application environment: The text-image task submitted by the task submission end needs to generate multiple related images. Different images involve the same subject, style and continuous content. At this time, simply converting the nodes in the streaming output into image tasks one by one can start the processing in advance. However, if there is a lack of independent management at the entity level and unified context reference, the appearance and visual style of the subjects of the images corresponding to the content fragments are likely to be inconsistent.

[0122] When the structured node data output by the streaming extensible markup language parsing module is identified as an entity description node, the node type determination module extracts the entity identifier, style trigger information, style parameters, and entity visual attribute description from the node. The above information is not directly used for the content fragment image task, but is first entered into the idempotent deduplication module.

[0123] The idempotent deduplication module determines whether the entity identifier already exists in the set of processed nodes. If the entity identifier already exists, it means that the corresponding entity description node has already been processed in the previous cumulative parsing, and the main image task will not be constructed in this round. If the entity identifier does not yet exist, it is written into the set of processed nodes and combined with the current task identifier to generate the corresponding subtask identifier.

[0124] Subtask identifiers are generated primarily using a combination of task identifiers and entity identifiers. When an entity identifier appears repeatedly in different tasks, different subtask identifiers are still generated based on the different task identifiers to avoid cross-task interference. The prompt word rendering module merges style trigger information, style parameters, and entity visual attribute descriptions to form the main image task input parameters, while recording the parent-child task relationship between the main image task and the current generation task. The task distribution module then sends the asynchronous image generation task containing the main image task input parameters to the image generation module so that the main image or main reference result can be generated in advance.

[0125] When writing the input parameters of the main image task into the fields of the entity description node, a fixed field mapping rule is adopted: the entity identifier is written into the main reference field, the style trigger information is written into the style trigger field, the style parameter is written into the style parameter field, and the entity visual attribute description is written into the main description field.

[0126] When the visual attribute description of an entity includes descriptions of appearance, clothing, posture, and color, the prompt word rendering module concatenates them in the order of appearance, clothing, posture, and color; if any item is missing, it is skipped while the remaining items remain in the same order.

[0127] The parent-child task relationship record is stored using a binary mapping table. The parent task key is the generation task identifier, and the child task key is the child task identifier. Each time a new main image task is added, a corresponding relationship is written so that subsequent content fragments can look up the corresponding main description source according to the entity identifier when referencing the entity. Style trigger information refers to keywords, style tags, or style phrases used to directly trigger the image generation module to call specific style capabilities. Style parameters refer to field values ​​that limit style intensity, brushstroke tendency, color tendency, lens performance, or image ratio. Entity visual attribute description refers to a set of text descriptions used to depict the appearance of the subject.

[0128] If multiple style trigger information appears in the same entity description node, the style trigger field is written in the order of appearance in the node content; if multiple style parameters appear, the parameter name is deduplicated and the later one is kept, overwriting the earlier one; if a style parameter is missing a parameter name, only the first one that appears is kept.

[0129] The source, deduplication method, and writing order of entity node input items are all fixed; through the deduplication judgment of entity identifiers and the recording of parent-child task relationships, the same subject only creates a corresponding sub-task when it is first identified, and subsequent fragments only refer to the determined entity description information and do not reconstruct the subject task;

[0130] When structured node data is identified as a content fragment node, the node type determination module extracts the fragment identifier, entity reference, generated prompt information, global context constraints, style parameters, style trigger information, and reference descriptions; the idempotent deduplication module compares the fragment identifier; if the fragment identifier already exists in the processed node set, the current fragment is considered a duplicate identification result and is skipped; if the fragment identifier has not yet been recorded, it is written into the processed node set.

[0131] The prompt word rendering module merges global context constraints, style parameters, style trigger information, entity references, and generated prompt information to form the input parameters for the content fragment image task. The merging is performed in a fixed order: first, global context constraints are written; then, entity descriptions corresponding to the referenced entities are written; style parameters and style trigger information are written; and finally, the generated prompt information and reference descriptions for the content fragment itself are written. When the same field exists in different sources, the priority is as follows: fields within the content fragment node, fields in the entity description node, and fields in the global constraint node.

[0132] The correspondence between entity references and entity descriptions is established based on entity identifiers. When there are multiple entity references in a content fragment node, the corresponding entity descriptions are concatenated in the order in which the references appear, and the same entity identifier is written only once in the input parameters of the same content fragment image task.

[0133] If an entity reference does not find a corresponding entity identifier in the parsed entity description list, the entity reference field is retained but the construction of the current fragment task is not blocked. Subsequent fragments continue to be rendered according to the updated entity description list. The task distribution module encapsulates the image task input parameters into an asynchronous image generation task and sends it to the image generation module. The result processing module synchronously records the task status, start time, and content fragment image task input parameters so that the task status of the corresponding fragment can be traced back when summarizing the processing results later.

[0134] The input parameters for the content fragment image task are generated according to a defined field assembly table, which includes at least a context field, an entity field, a style field, a fragment hint field, and a reference description field. Among them, the context field receives global context constraints, the entity field receives entity descriptions concatenated in the order of entity references, the style field receives style parameters and style trigger information, the fragment hint field receives generated hint information, and the reference description field receives reference descriptions.

[0135] When assembling fields, a sequential append rule is used: first, it is checked whether the previous field is empty. If it is not empty, a preset separator is inserted between the fields before writing the next field. If it is empty, it is skipped directly without generating a blank placeholder. The preset separator is at least one of the following: comma, semicolon, newline character or combination thereof, and it is kept consistent within the same task.

[0136] In cases where the entity reference is missing a corresponding entity description, the prompt word rendering module only writes the existing entity description and does not generate alternative or presumed descriptions for the missing entity, so that the task input parameters are constructed entirely based on the parsed data.

[0137] Global context constraints refer to shared constraint texts applicable to multiple content segments, including at least one of environmental background, era setting, narrative context, visual tone, or unified prohibitions; generated prompt information refers to the visual actions, compositional intentions, subject relationships, or plot points corresponding to the current content segment itself; reference descriptions refer to weak prompt texts that do not directly participate in strong constraint control but can serve as supplementary references.

[0138] The null value judgment of each field in the field assembly table is performed according to whether the character length is greater than zero; before writing, the original text is first removed from the leading and trailing whitespace characters and continuous blank lines are compressed; if the same repeated substring appears in the same field, only the first occurrence of the substring is retained; the sources of the task input parameters and how the sources are organized before assembly all have clear execution boundaries.

[0139] When generating input parameters for both the main image task and the content fragment image task, the prompt word rendering module executes the same stateful assembly process. This process is carried out sequentially in four stages: candidate collection, sequential merging, conflict resolution, and result freezing. It collects usable descriptive content from the currently closed and deduplicated structured node data.

[0140] Write the available content to the corresponding positions in the predetermined order; check again for semantic conflicts or duplicate constraints. If they exist, retain the source closest to the current task and suppress the other sources; freeze the input parameters of this image task before task distribution so that they will not be written back and modified after the subsequent addition of new nodes.

[0141] The source closest to the current task refers to the source that is more directly related to the image to be generated in the processing chain. The fields in the content fragment node are higher than the fields in the entity description node, and the fields in the entity description node are higher than the fields in the global constraint node. If there are multiple descriptions of the same type within the same source, the later closing one will overwrite the earlier closing one.

[0142] In cases where there is a discrepancy between style trigger information and style parameters, the more restrictive description is preferred. Here, "more restrictive" means that the description that leads to a narrower range of candidate generation is adopted first, while the description with a wider range is only added as a supplement when there is no conflict.

[0143] The rules for determining semantic conflicts include: when two descriptions are written to the same target field and the field names are the same but the field values ​​are different, they are determined to be conflicts of the same kind; when one description belongs to an explicitly prohibited item and the other description belongs to the corresponding allowed item, and both point to the same object, they are determined to be constraints conflicts.

[0144] For similar conflicts, they are handled according to the source priority and the post-closure overriding rule; for constraint conflicts, prohibited items are retained first; result freezing means that before the asynchronous communication mechanism is written into the task distribution module, an unwriteable version record is generated for the input parameters of this task. Subsequent addition of nodes only affects subsequent tasks that have not yet been frozen, and does not rewrite frozen tasks; conflict judgment conditions, overriding relationships and freezing time points are all based on preset rules.

[0145] Based on this, if the generated task itself does not carry contextual constraint information or entity description information, and only relies on the local content inside the content fragment node for prompt word rendering, it will result in a lack of unified constraints between different fragments. To handle this situation, the system sets up a global constraint and entity context extraction module. When the missing contextual constraint information or entity description information is detected, the global constraint and entity context extraction module generates supplementary prompt words through the model prompt word construction module, and calls the large language model module to generate incremental text fragments containing global constraint nodes and entity description nodes.

[0146] The streaming extensible markup language parsing module parses these nodes to obtain a global context constraint and entity description list. Context missing judgment is performed according to the following rules: when the generation task does not carry context constraint information, it is considered that the global constraint is missing; when the generation task does not carry entity description information, or although it carries entity description information, it is not possible to establish a correspondence between entity identifiers and entity references in the content fragment, it is considered that the entity description information is missing; supplementary prompt words are triggered only when at least one of the above missing situations exists.

[0147] When the prompt word rendering module generates image task input parameters for each content fragment, it uniformly references this global context constraint and entity description list, so that although each content fragment triggers image generation separately, it is still processed under the same subject setting and the same style constraint.

[0148] The inability to establish a correspondence between entity identifiers and entity references in content fragments means that, in the list of entity descriptions parsed by the current task, after searching item by item by entity identifier, at least one entity reference in a content fragment node does not find an entity identifier with the same name, or multiple entity identifiers with the same name are found but cannot be uniquely located based on existing fields within the task; if the latter occurs, the first closed item in the order of appearance in the entity description nodes will be taken as the temporary reference source, while the missing entity description information will be maintained to allow for subsequent supplementation and extraction to continue to improve the information;

[0149] If only some global constraint nodes are obtained during the supplementary extraction process and the entity description nodes are not yet closed, the system will continue to retain the corresponding unclosed segments until the subsequent incremental text is completed before entering the unified reference stage; if a content segment node closes before the corresponding entity description node, the segment node can still complete the recognition and deduplication first, but the prompt word rendering module will call the obtained global context constraints and entity description list when constructing its image task input parameters.

[0150] If a certain part has not yet formed a usable result, only the parsed constraint content will be used for rendering, and the updated context information will continue to be referenced in subsequent fragment processing; the usable result is structured data that has formed closed extensible markup language nodes, completed node type determination, and undergone idempotent deduplication; data that has not yet met this condition will not participate in the current round of prompt word rendering;

[0151] When the missing entity description information of a content segment image task is filled in after the task has been distributed, the newly closed content segment nodes will uniformly reference the supplemented entity description list, and the distributed task will keep the original task input parameters without rollback and modification, thereby ensuring that the task status record is consistent with the input parameters received by the image generation module.

[0152] The context missing detection and supplementary extraction adopt a control method that combines event triggering and suppression: when a missing situation is detected for the first time, a supplementary extraction request is immediately generated and the corresponding task is set to the supplementary processing state; before this state is lifted, even if a new content fragment node is received later, the same type of supplementary extraction request will not be repeatedly initiated.

[0153] A new supplementary extraction is only allowed when the supplementary extraction has ended and the missing content still exists. If the node returned by the supplementary extraction only covers part of the missing content, the system will put the completed part into the referenceable state and keep the uncompleted part in the missing state. Subsequent prompt word rendering will continue to be executed in the manner of prioritizing the referenceable state and skipping the missing state. This allows the supplementary extraction process to be kept in parallel with the content fragment task construction process, while avoiding frequent changes in the context source within the same task due to repeated supplementary calls.

[0154] In this text-to-image task, entity description nodes correspond to the subject setting, content fragment nodes correspond to each image unit, and global constraint nodes correspond to the style and context information that the entire batch of images follows. Through the above processing method, the system can first form subject information and global constraints during the streaming output process, and then continuously pass them to subsequent content fragment tasks, so that the generated multiple target images are consistent in terms of subject image, picture style and context expression.

[0155] The system also includes a long content concurrent processing module, which is specifically configured as follows: content batching: for a task to generate text containing multiple content fragments, the text processing module divides the text to be processed into multiple paragraph batches according to a preset content fragment number threshold. Each paragraph batch carries its start and end positions in the complete content.

[0156] Concurrent generation: Each paragraph batch calls the large language model module to generate corresponding content fragment nodes in a streaming manner;

[0157] Concurrent processing: Multiple paragraph batches are distributed and processed through concurrent tasks, including streaming scalable markup language parsing, node deduplication, prompt word rendering, and asynchronous image generation.

[0158] The result processing module and the task distribution module are also configured to perform the following exception handling: Result summarization: After the large language model module finishes streaming, the result processing module summarizes the target images corresponding to the processed structured node data and outputs the processing results.

[0159] Anomaly detection: When network anomalies, model anomalies, or task distribution anomalies occur during processing, the result processing module records the anomaly information and the current task status;

[0160] Retry on failure: The task distribution module will resubmit the asynchronous image generation task that has an error to the waiting queue according to the preset number of task retries, or write the failure information to the preset result queue.

[0161] The task distribution module and the image generation module communicate through an asynchronous communication mechanism. The image generation module is configured with at least one of the following: text-to-image generation model, image-to-image generation model, subject generation model, illustration generation model, or video keyframe generation model. The processed node set uses an idempotent deduplication table to record the node identification information that has been processed.

[0162] In the application environment: When the text to be processed submitted by the task submission end contains multiple original content fragments, if all the content is still processed in the order of a single streaming session, the starting order of subsequent image tasks will be completely subject to the output progress of that single session, which may cause the later fragments to remain in the waiting state for a long time.

[0163] For a generation task containing multiple original content fragments, the text processing module divides the text to be processed into multiple paragraph batches according to a preset threshold for the number of content fragments; each paragraph batch carries its start and end positions in the complete content, which are used to indicate the corresponding range of the batch in the original text.

[0164] The content fragment number threshold is determined based on the processing capability of completing one streaming generation and parsing of the text to be processed within a single batch within the target latency. When the number of original content fragments is not greater than the content fragment number threshold, single batch processing is maintained. When the number of original content fragments is greater than the content fragment number threshold, the original content fragments are divided into multiple paragraph batches according to their order in the text to be processed, and each paragraph batch contains at least one original content fragment.

[0165] After the batching is completed, each paragraph batch calls the large language model module to generate the corresponding content fragment nodes in a streaming manner. Since each batch has an independent streaming output process, the text processing module maintains the corresponding text buffer and processed node record for each batch, and receives the corresponding incremental text fragments in parallel.

[0166] The segmentation of paragraph batches adopts a deterministic sequential segmentation rule: first, read the total number of original content segments N and the content segment number threshold T. When N is less than or equal to T, only one paragraph batch is generated; when N is greater than T, according to the natural order of the original content segments in the text to be processed, every T original content segments from the beginning to the end are divided into a paragraph batch, and the part with less than T original content segments at the end constitutes the last paragraph batch separately.

[0167] Each paragraph batch is assigned a unique paragraph batch identifier after it is generated. This identifier includes at least two parts: a task identifier and a batch number. The batch number starts from 1 and increments. The text buffer, processed node record, and task status record of each paragraph batch are maintained separately with the paragraph batch identifier as the index, thereby isolating the data range between concurrent batches.

[0168] The content fragment quantity threshold T is determined by the number of original content fragments submitted with the task and the number of available concurrent threads in the system: when the number of available concurrent threads is 1, T is the total number of original content fragments; when the number of available concurrent threads is greater than 1, T is the quotient of N (rounded up) divided by the number of available concurrent threads, and is limited to between 1 and 10; if T obtained according to the above rules is greater than 10, then 10 is used; if it is less than 1, then 1 is used; the batching threshold is not arbitrarily specified, but can be directly determined by the task input size and the current concurrent processing capacity;

[0169] During concurrent execution of each paragraph batch, streaming Extensible Markup Language (XML) parsing, node deduplication, prompt word rendering, and asynchronous image generation task distribution are all executed independently according to the batch. After each batch receives a new incremental text segment, it first identifies closed XML nodes in the cumulative output text of this batch, and then compares them with the corresponding set of processed nodes according to the node identification information to generate image task input parameters and send them to the image generation module.

[0170] In this way, content fragment nodes in different paragraph batches can trigger image generation tasks separately without waiting for all original content fragments to be generated sequentially in the same streaming output link; if the same original content fragment needs to form a semantic continuous unit with adjacent content due to upstream segmentation rules, the original content fragment is only assigned to the paragraph batch where its starting position is located, avoiding duplicate assignment across batches.

[0171] The task distribution module and the image generation module are decoupled using an asynchronous communication mechanism. The image task input parameters are encapsulated into asynchronous image generation tasks in the task distribution module and then written into the task queue, event bus, task table, callback interface or other asynchronous scheduling mechanism. The image generation module consumes tasks according to its own processing capacity.

[0172] The image generation module can be configured with at least one of the following models: text-to-image generation model, image-to-image generation model, subject generation model, illustration generation model, or video keyframe generation model, to adapt to different types of image generation needs. Meanwhile, the processed node set uses an idempotent deduplication table to record the node identification information that has been processed. For node identification information in concurrent batches, a combination of task identifier, paragraph batch identifier, and node self identifier is preferentially used to write it into the idempotent deduplication table.

[0173] When a node is a global constraint node and does not directly carry its own identifier, the task identifier, paragraph batch identifier, node type and its closing order in this batch are used to generate a record item; when the cumulative output text is parsed again in any batch, as long as the corresponding node identifier information already exists in the idempotent deduplication table, the structured node data will not trigger the creation of the asynchronous image generation task again.

[0174] The parsing, deduplication, and task distribution of each paragraph batch are independent of each other, while the result processing module updates the status of each batch in an event-driven manner; once a node closure event is generated in any batch, the subsequent processing corresponding to that batch can be triggered without waiting for other batches to reach the same processing progress. Therefore, clock-level synchronization is not required between different batches. It is only required that the text buffer and processed node records be maintained serially within the same batch according to the arrival order of incremental text fragments.

[0175] The idempotent deduplication table operates using a query-by-item and write-by-item rules: when a batch of structured node data is parsed, the node identifier information is used as the query key to retrieve the idempotent deduplication table; if there is a completely identical query key, the duplicate result is returned and the task creation for that node is terminated; if there is no completely identical query key, a record with a status of "registered" is written first, and then the task distribution module is allowed to proceed.

[0176] For asynchronous image generation tasks written to the processing queue by the task distribution module, the task identifier is generated by combining the task identifier, paragraph batch identifier, node identifier information and task type in sequence, so that retry judgment, result summary and deduplication interception are all performed based on the same identifier system.

[0177] When the result processing module receives status updates from the task distribution module and the image generation module, it first locates the generated task according to the task identifier, and then locates the corresponding record according to the paragraph batch identifier and node identifier information, so as to update the task status, start time, end time and result address of the record.

[0178] The results processing module continuously receives task status information during the concurrent execution of each batch. When the corresponding large language model module outputs the end marker, the results processing module summarizes the target images corresponding to the processed structured node data and outputs the processing results. This summarization process can be aggregated by task identifier or organized according to the correspondence between paragraph batches and original content fragments so that the upstream system can obtain complete processing results.

[0179] Because the image generation module and the text processing module use an asynchronous communication mechanism, the result processing module does not rely on a single synchronous return path when summarizing, but integrates based on the task status record and the target image result; if there are asynchronous image generation tasks that have not yet been completed, the result processing module keeps the current task in the state of pending summary completion, and continues to perform the summary after the corresponding task status is updated.

[0180] In terms of exception handling, if a network exception, model exception, or task distribution exception occurs during concurrent processing, the result processing module records the exception information and the current task status. This record can be associated with and saved with the corresponding task identifier, segment batch information, and node identifier information to maintain traceability when retrying or returning failure results. The task distribution module resubmits the asynchronous image generation task that has an exception to the waiting queue according to the preset number of task retries. If the task still fails to recover after the preset retry conditions are met, failure information is written to the result queue.

[0181] The preset number of task retries is determined using a fixed-number strategy and remains consistent for the same asynchronous image generation task throughout its lifecycle. Before each resubmission, the idempotent deduplication table and task status record are queried, and resubmission is only performed when the asynchronous image generation task is in a failed state and the preset number of task retries has not been reached.

[0182] Task status is at least divided into pending, processing, success, and failure; resubmission judgment is only allowed when the task status changes from processing or pending to failure, and asynchronous image generation tasks in the successful state will not enter the retry process; exception handling and normal result summarization are carried out separately and do not affect the return of processing results for other batches that have been successfully completed.

[0183] The preset number of retry attempts for a task is 1 to 3; the number of retry attempts for network anomalies is 3, the number of retry attempts for task distribution anomalies is 2, and the number of retry attempts for model anomalies is 1; before a task status changes from failure to resubmission, the waiting time since the last failure record must reach the preset retry interval.

[0184] The preset retry interval increases in sequence with the retry number. The first retry waits for 10 to 30 seconds, the second retry waits for 30 to 90 seconds, and the third retry waits for 90 to 180 seconds. If the corresponding task status is updated to success by an external callback during the waiting period, the current resubmission is canceled. The exception type, retry limit, and retry triggering time are all clearly defined.

[0185] Anomaly handling employs a fixed state transition rule: the asynchronous image generation task is initially written to a pending state, transitions to a processing state when it is taken by the image generation module, transitions to a successful state after successfully generating the target image, and transitions to a failed state when a network anomaly, model anomaly, or task distribution anomaly occurs and the current processing is not yet complete; the task distribution module only re-delivers the asynchronous image generation task when the task status is failed and the corresponding retry count is less than the preset number of task retries, and increments the retry count by 1 before re-delivering;

[0186] If the retry count reaches the preset number of task retries, the result processing module marks the task as a final failure and writes failure information to the result queue and stops submitting it. Through the above state transition rules, the triggering conditions, stopping conditions and result output conditions for failure retries all have definite judgment boundaries.

[0187] The processing rhythm between concurrent batches is coordinated by batch-level rate limiting rules; the number of batches allowed to enter the task distribution module at any given time does not exceed the number of currently available concurrent threads; when the number of concurrent batches exceeds the number of available concurrent threads, batches with a larger number of cumulative closed nodes and a smaller number of unfinished tasks are given priority to be distributed first, while the remaining batches remain in a waiting state.

[0188] A batch in the waiting state can still receive incremental text fragments and complete the internal parsing of this batch, but the input parameters of its newly generated image task will not be written into the asynchronous communication mechanism until processing resources are regained; if the number of undistributed tasks of a batch continues to increase and exceeds the preset limit during the waiting period, the continued fetching of subsequent streaming calls for that batch will be suspended; the streaming calls for that batch will be resumed after the number of distributed tasks decreases.

[0189] The preset upper limit is determined by multiplying the current number of available concurrent threads by 2 to 5 and rounding up to the nearest integer. When the number of available concurrent threads changes, only the updated preset upper limit is applied to the batches that enter the waiting state later. The upper limit value corresponding to the batches that are already in the waiting state remains unchanged. This creates a dynamic rate limit between the text generation side and the image processing side, preventing the backlog of tasks from getting out of control when multiple batches of long content are processed concurrently due to insufficient processing capacity on one side.

[0190] If the streaming output of a certain batch has ended, while other batches are still being generated, the result processing module only records the structured node data and target image corresponding to the completed batch in stages, and outputs the whole task result after all relevant end markers are in place; if a batch is repeatedly entered into the parsing process due to abnormal retry, the idempotent deduplication table still intercepts the creation of duplicate tasks based on the node identifier information, thereby avoiding the problem of duplicate delivery under the combined effect of concurrent retries and duplicate parsing.

[0191] In this text-to-image task, the long text to be processed is split into multiple paragraph batches. Each batch can independently complete streaming generation, parsing of closed nodes, deduplication judgment, and image task distribution. The image generation module processes the tasks step by step through an asynchronous communication mechanism. Image generation for the preceding content segments can start before the subsequent segments. In the event of network anomalies, model anomalies, or task distribution anomalies, the system can still maintain the overall processing chain through task status recording, failure retry, and result queue writing.

[0192] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. An image generation system adapted for large-scale model streaming output, characterized in that, include: The system includes a text processing module, a model prompt word construction module, a large language model module, a streaming scalable markup language parsing module, a node type determination module, an idempotent deduplication module, a prompt word rendering module, a task distribution module, an image generation module, and a result processing module. The text processing module receives the task of generating the text to be processed and maintains the text buffer and the set of processed nodes; The model prompt word construction module constructs prompt words based on the generated task, and the constraint large language model module outputs them in a preset extensible markup language format. The text processing module inputs the prompt word into the large language model module and streams the incremental text fragments and end markers. The streaming extensible markup language parsing module appends the incremental text fragment to the text buffer for parsing, converts the parsed closed nodes into structured node data containing node identification information, and retains unclosed fragments to splice subsequent incremental fragments; The node type determination module divides the structured node data into entity description nodes, content fragment nodes, and global constraint nodes, and extracts the corresponding attribute information. The idempotent deduplication module compares the structured node data with the set of processed nodes based on the node identifier information. If the node identifier information is not found in the set, it is added. The prompt word rendering module renders the image task input parameters based on the extracted attribute information; The task distribution module constructs the input parameters into an asynchronous task and sends it to the image generation module to generate the target image; After obtaining the end marker, the result processing module summarizes the target image to obtain the processing result.

2. The image generation system adapted for large-model streaming output according to claim 1, characterized in that: The text processing module, model prompt word construction module, and large language model module are specifically configured as follows: The text processing module receives a generation task containing business data, which includes at least one of the following: text to be processed, a list of content fragments, style identifiers, image ratios, entity description information, and contextual constraint information. The task type of the generation task includes at least one of the following: entity description extraction task, content fragment image generation task, global constraint extraction task, and image redrawing task. The model prompt word construction module reads the corresponding preset prompt word template according to the task type, and defines an Extensible Markup Language (XML) output format in the preset prompt word template. The XML output format requires that the entity description node include at least one sub-node of entity identifier, entity category, entity visual attribute description, style trigger information, style parameter, and description text; the content fragment node includes at least one sub-node of fragment identifier, entity reference, generated prompt information, style parameter, style trigger information, and reference description; and the global constraint node includes at least one sub-node of style constraint, context constraint, and content list. The text processing module inputs the prompt words into the large language model module in a streaming mode, and initializes the text buffer and the set of processed nodes. Whenever the large language model module returns an incremental text fragment, the text processing module appends the incremental text fragment to the text buffer to form the current cumulative output text.

3. The image generation system adapted for large-scale model streaming output according to claim 2, characterized in that: The streaming extensible markup language parsing module is specifically configured as follows: The streaming Extensible Markup Language (EXPLAIN) parsing module identifies text segments with matching start and end tags in the currently accumulated output text according to preset parsing rules, wherein the start and end tags have the same node name; When an Extensible Markup Language (XML) fragment is detected to contain only a start tag but no corresponding end tag is received, the XML fragment is determined to be an unclosed XML fragment and is kept in the text buffer to wait for it to be concatenated with the subsequent incremental text fragments. When an Extensible Markup Language (XML) fragment is detected to contain matching start and end tags, the XML fragment is determined to be a closed XML node, and its node content is further checked to see if it contains child XML nodes. If it contains child XML nodes, the node content is recursively parsed. If it does not contain child XML nodes, the node content is used as the leaf field value. The streaming extensible markup language parsing module uses a preset depth-first algorithm to traverse the node tree hierarchy, merging nodes with the same name: when there are multiple closed extensible markup language nodes with the same name in the same level structure, the multiple node results are organized into a list or array according to the order of appearance, and the corresponding structured node data is output.

4. The image generation system adapted for large model streaming output according to claim 1, characterized in that: When the structured node data is the entity description node, the specific processing steps include: The node type determination module extracts entity identifiers, style trigger information, style parameters, and entity visual attribute descriptions from the structured node data as the corresponding attribute information; the idempotent deduplication module determines whether the entity identifier already exists in the processed node set. If it already exists, the entity description node is skipped; if it does not exist, the entity identifier is added to the processed node set, and a corresponding subtask identifier is generated based on the entity identifier and the generation task. The prompt word rendering module merges the extracted style trigger information, style parameters, and entity visual attribute descriptions to render and generate the main image task input parameters, and records the association between the generation task and the main image task; the task distribution module sends the asynchronous image generation task containing the main image task input parameters to the image generation module.

5. The image generation system adapted for large-scale model streaming output according to claim 1, characterized in that: When the structured node data is the content fragment node, the specific processing steps include: The node type determination module extracts fragment identifiers, entity references, generated prompt information, style parameters, style trigger information, and reference descriptions from the structured node data as the corresponding attribute information; The idempotent deduplication module determines whether the fragment identifier already exists in the processed node set. If it already exists, the content fragment node is skipped; if it does not exist, the fragment identifier is added to the processed node set. The prompt word rendering module merges and extracts the fragment identifier, entity reference, generated prompt information, style parameters, style trigger information, and reference description, and renders the content fragment image task input parameters. The task distribution module sends an asynchronous image generation task containing the input parameters of the content fragment image task to the image generation module, and the result processing module records the task status, start time, and the input parameters of the content fragment image task.

6. The image generation system adapted for large model streaming output according to claim 2, characterized in that: The system also includes a global constraint and entity context extraction module, which is specifically configured as follows: Context missing judgment: When the received generation task does not carry context constraint information or entity description information, the global constraint and entity context extraction module generates supplementary prompt words through the model prompt word construction module, and calls the large language model module to generate an incremental text fragment containing the global constraint node and the entity description node; Context parsing: The streaming extensible markup language parsing module parses the global constraint node and entity description node to obtain a list of global context constraints and entity descriptions; Unified reference: When generating the image task input parameters for each content fragment, the prompt word rendering module uniformly references the global context constraints and entity description list.

7. The image generation system adapted for large-scale model streaming output according to claim 1, characterized in that: The system also includes a long-content concurrent processing module, which is specifically configured as follows: Content batching: For the generation task in which the text to be processed contains multiple content fragments, the text processing module divides the text to be processed into multiple paragraph batches according to a preset content fragment number threshold. Each paragraph batch carries its start and end positions in the complete content. And this occurs as follows: each paragraph batch calls the large language model module to stream generate the corresponding content fragment nodes; Concurrent processing: Multiple paragraph batches are processed through concurrent tasks, including streaming scalable markup language parsing, node deduplication, prompt word rendering, and the distribution of the asynchronous tasks.

8. The image generation system adapted for large-scale model streaming output according to claim 1, characterized in that: The result processing module and the task distribution module are also configured to perform the following exception handling: Results Summary: After the large language model module streams the end marker, the result processing module summarizes the target images corresponding to the processed structured node data and outputs the processing results. Anomaly detection: When network anomalies, model anomalies, or task distribution anomalies occur during processing, the result processing module records the anomaly information and the current task status; Retry on failure: The task distribution module will resubmit the asynchronous task that has encountered an error to the waiting queue according to the preset number of task retries, or write failure information to the preset result queue.

9. The image generation system adapted for large-scale model streaming output according to claim 1, characterized in that: The task distribution module and the image generation module communicate through an asynchronous communication mechanism. The image generation module is configured with at least one of the following: text-to-image generation model, image-to-image generation model, subject generation model, illustration generation model, or video keyframe generation model. The processed node set uses an idempotent deduplication table to record the node identification information that has been processed.