Method, device and storage medium for controlling a text-to-image generation model

By integrating generation task constraint parameters and target concept dimensions into the text-to-image generation model, the problems of insufficient model adaptability and low generation efficiency are solved, achieving efficient and accurate image generation.

CN121328775BActive Publication Date: 2026-02-27PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511872540.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-02-27
Estimated Expiration
2045-12-12

AI Technical Summary

Technical Problem

Existing text-to-image generation models rely on expensive human prior knowledge and a large amount of manually labeled data, making them difficult to adapt to the image generation requirements of real-world scenarios and resulting in low efficiency.

Method used

By proposing text prompts based on the model's inherent knowledge, matching the image generation quality threshold and inference cost expectation, integrating generation task constraint parameters, filtering target concept dimensions, combining multiple dimensions and assigning contextual attribute values, completing simple sub-tasks corresponding to test prompts, evaluating alignment and learning, filtering target text into image tools, and finally generating the target image.

Benefits of technology

It improves the quality and semantic alignment of image generation, optimizes inference cost control, and enhances the accuracy of tool selection and the efficiency of the generation process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121328775B_ABST
    Figure CN121328775B_ABST
Patent Text Reader

Abstract

The application discloses a text-to-image generation model control method, device and storage medium, relates to the technical field of image generation, and the method comprises the following steps: independently proposing a text prompt word according to model internal knowledge, matching corresponding image generation quality threshold values and inference cost expectations, and integrating generation task constraint parameters; filtering a scene-based target concept dimension according to the prompt word and the constraint parameters, cross-combining and giving scene attribute value instances, obtaining a test prompt word and a task exploration decision; completing an image result of a simple subtask of the test prompt word, evaluating alignment degree and learning, and refining adaptive knowledge; combining the decision, adaptive knowledge and tool performance and task history to filter a target tool, and finally completing image generation of a user input prompt word through the tool and outputting a target image. Through the image generation process of prompt word driving constraint adaptation and tool screening, the technical problem of low image generation efficiency is solved, and the execution efficiency of a generation task is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image generation technology, and in particular to a control method, device and storage medium for a text-to-image generation model. Background Technology

[0002] The text-to-image model industry has experienced explosive growth, with a proliferation of models ranging from large-scale, high-precision to lightweight, real-time models, forming a diverse yet highly fragmented text-to-image generation ecosystem. In these technologies, model training relies on expensive human prior knowledge and requires manually written tool descriptions or large amounts of manually labeled data, making it difficult for models to adapt to the image generation requirements of real-world scenarios, thus resulting in low image generation efficiency.

[0003] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention

[0004] The main objective of this application is to provide a control method, device, and storage medium for a text-to-image generation model, aiming to solve the technical problem of low image generation efficiency.

[0005] To achieve the above objectives, this application proposes a control method for a text-to-image generation model, the method comprising:

[0006] Based on the model's inherent knowledge, text prompts are proposed, and the image generation quality threshold and inference cost expectation corresponding to the text prompts are matched according to preset indicators, and the generation task constraint parameters are integrated.

[0007] Based on the text prompts and their corresponding task constraint parameters, a set of target concept dimensions that conform to the characteristics of general scenarios are selected and defined.

[0008] By combining the target concept dimensions in multiple dimensions, assigning specific scenario-based attribute values ​​to the resulting combinations and instantiating them, test prompt words and task exploration decisions are obtained.

[0009] By completing simple sub-tasks corresponding to the test prompt words to obtain image results, the alignment degree between the test prompt words and the image results is evaluated and learning is carried out to obtain adaptation knowledge;

[0010] Based on the task exploration decision and the adaptation knowledge, combined with the performance characteristics of the text-to-image tool and the current task execution history, the target text-to-image tool is selected from the feasible set of text-to-image tools.

[0011] The target text to image tool is used to generate the image corresponding to the test prompt word and output the target image.

[0012] In an embodiment, the text prompt is received and verified, invalid and rule-violating text content is filtered out, and a text prompt that passes verification is obtained;

[0013] According to the text prompt scene classification rules and historical configuration data in the preset index library, an image generation quality threshold and inference cost expectation that fit the text prompt are matched out;

[0014] According to the parameter packaging specification, the image generation quality threshold, the inference cost expectation, and the corresponding text prompt are associated and bound, and the task constraint parameter is integrated and generated.

[0015] In an embodiment, the core semantic features of the text prompt are extracted, the quality requirement and cost requirement in the task constraint parameter are combined, the scene type corresponding to the text prompt is determined, and the analysis result is integrated;

[0016] Based on the association rules of different scenes and concept dimensions in the concept dimension library, the initial concept dimensions that fit the analysis result are filtered out;

[0017] According to the task constraint parameter, the initial concept dimensions are de-redundant and prioritized, dimensions that do not match the task constraint are removed, and basic and orthogonal core dimensions are retained, to obtain the target concept dimensions.

[0018] In an embodiment, the target concept dimensions are multi-dimensionally cross-combined according to a preset dimension combination rule, and a dimension combination set arranged according to an exploration priority is generated by combining a priority determination strategy;

[0019] According to the attribute value domains corresponding to different concept dimension combinations in the dimension combination set, a quantified scene attribute value is assigned to each dimension combination, to obtain a scene dimension combination;

[0020] The scene dimension combination and its corresponding attribute value are converted into a natural language description that conforms to the text-to-image generation task specification, forming the test prompt, and the task exploration decision corresponding to the test prompt is generated according to a preset exploration decision rule.

[0021] In an embodiment, according to a preset subtask execution rule, simple subtasks corresponding to the test prompt are completed in sequence, and tool call information and execution time of each simple subtask are recorded, to obtain the image result and its corresponding execution process data;

[0022] The multi-dimensional alignment scores of the test prompt and the corresponding image result in content semantics, feature attributes, and scene matching degree are calculated by a preset evaluation algorithm, and the single-generation compliance probability of each simple subtask corresponding to the test prompt is calculated, to generate a comprehensive evaluation result;

[0023] According to the exploration space pruning strategy, it is judged whether the current text-to-image tool has mastered all simple sub-tasks corresponding to the test prompt word, and based on the alignment score and the passing probability of each dimension in the comprehensive evaluation result, targeted learning is carried out, the adaptation knowledge between tool capability and sub-task demand is extracted, and the task exploration decision is generated according to the judgment result.

[0024] In an embodiment, the text-to-image tool performance characteristics and the current task execution history are called, and the task exploration decision and its corresponding adaptation knowledge, tool performance characteristics, and task execution history are associated and integrated to obtain tool selection comprehensive reference data;

[0025] According to the tool selection comprehensive reference data and the task constraint parameters, all text-to-image tools are screened, and tools with quality not meeting the standard and cost exceeding the expectation are proposed to obtain the text-to-image tool feasible set;

[0026] Based on the cost-quality trade-off decision algorithm, the target text-to-image tool that can meet the current task quality requirement and achieve the optimal reasoning cost is selected from the text-to-image tool feasible set according to the judgment of task complexity in the task exploration decision.

[0027] In an embodiment, the special adaptation interface of the target text-to-image tool is called, the tool running parameters are configured according to the quality precision requirement in the task constraint parameters, the complex image generation operation corresponding to the test prompt word is executed, and the initial generated image and its corresponding process data are obtained;

[0028] The quality threshold in the task constraint parameters is compared with the semantic alignment degree and the feature matching degree of the initial generated image and the test prompt word to generate a judgment result.

[0029] If the judgment result meets the quality threshold, the initial generated image is standardized and output as the target image.

[0030] In an embodiment, if the judgment result does not meet the quality threshold, the initial generated image, the tool used for generation, the generation time consumption, the review score, the semantic alignment degree, and the feature matching degree are associated to obtain single generation task data.

[0031] Through a preset difference analysis algorithm, the specific causes of insufficient semantic alignment degree, feature matching deviation, and excessive time consumption in the single generation task data are located, and a task data attribution result containing the problem root and the tool adaptation short board is output.

[0032] According to the task data attribution result, a tool selection strategy, tool generation precision and resource allocation parameter are formulated, and an optimization adjustment strategy for a new round of image generation is formed by integration, so as to call and configure a corresponding text-to-image tool according to the optimization adjustment strategy to complete the image generation operation of the test prompt word.

[0033] In addition, to achieve the above-mentioned purpose, the present application also provides a text-to-image generation device, which comprises a memory, a processor and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the control method of the text-to-image generation model as described above.

[0034] In addition, to achieve the above-mentioned purpose, the present application also provides a storage medium, which is a computer readable storage medium, and a computer program is stored on the storage medium, and the computer program is executed by a processor to implement the steps of the control method of the text-to-image generation model as described above.

[0035] The present application provides a control method of a text-to-image generation model, which comprises proposing a text prompt word by relying on the internal knowledge of the model, matching its corresponding image generation quality threshold and reasoning cost expectation to integrate generation task constraint parameters, screening and defining target concept dimensions that meet general scene characteristics according to the prompt word and constraint parameters, multi-dimensionally crossing and combining and giving scene attribute value instances to obtain test prompt words and task exploration decisions, completing simple sub-tasks corresponding to the test prompt words to obtain image results, evaluating alignment degrees and learning to obtain adaptive knowledge, and then selecting a target text-to-image tool from a tool feasible set based on the task exploration decisions, adaptive knowledge, the performance characteristics of the text-to-image tool and the current task execution history, and finally completing image generation corresponding to the user input prompt word by the tool and outputting the target image, which solves the technical problems of insufficient tool adaptability, unbalanced generation quality and reasoning cost, and lack of self-knowledge driven optimization mechanism in text-to-image generation, improves image generation quality and semantic alignment degree, optimizes reasoning cost control, and improves tool selection accuracy and generation process efficiency.

[0036] In summary, the present application receives a user text prompt word first, matches an image generation quality threshold and a reasoning cost expectation to integrate generation task constraint parameters, then screens and defines target concept dimensions and combines instances to obtain a test prompt word, and then completes a simple sub-task generation exploration decision, selects a target text-to-image tool and generates a target image, which solves the technical problem of low image generation efficiency, improves the execution efficiency of the generation task, and achieves high matching degree of the generation result and quality cost constraint and accurate detection of tool capability boundary. BRIEF DESCRIPTION OF DRAWINGS

[0037] The accompanying drawings, which are incorporated herein and form a part of the specification, illustrate embodiments consistent with the present application and, together with the description, further serve to explain the principles of the application.

[0038] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings required by the embodiments or prior art description will be briefly introduced as follows. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative effort.

[0039] Figure 1 The flowchart of the first embodiment of the control method of the text-to-image generation model of the present application;

[0040] Figure 2 The model selection diagram of the present application;

[0041] Figure 3 The flowchart of the eighth embodiment of the control method of the text-to-image generation model of the present application;

[0042] Figure 4 The algorithm flowchart of the present application;

[0043] Figure 5 The effect comparison diagram of the present application;

[0044] Figure 6 The structure diagram of the text-to-image generation device of the present application.

[0045] The purpose implementation, functional features and advantages of the present application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION

[0046] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application, and are not used to limit the present application.

[0047] In the related art, model training relies on expensive manual prior knowledge, and needs to be completed through tool description written by hand or a large amount of manually annotated data, which causes the model to be difficult to adapt to the actual scene image generation requirements, and further leads to low image generation efficiency.

[0048] The application provides a solution: first, based on the intrinsic knowledge of the model, the text prompt word is proposed, the image generation quality threshold corresponding to the text prompt word is matched according to the preset index and the expected reasoning cost, the generation task constraint parameter is integrated, then a group of target concept dimensions meeting the general scene characteristics are selected and defined according to the text prompt word and the corresponding task constraint parameter, then the target concept dimensions are combined in multiple dimensions, the combination obtained is given specific scene attribute values and instantiated to obtain test prompt words and task exploration decisions, then the image results are obtained by completing the simple subtasks corresponding to the test prompt words, the alignment degree of the test prompt words and the image results is evaluated and learning is carried out, the adaptive knowledge is obtained, then the target text-to-image tool is selected from the text-to-image tool feasible set based on the task exploration decision, the adaptive knowledge, the performance characteristics of the text-to-image tool and the current task execution history, and finally the image corresponding to the test prompt word is completed by the target text-to-image tool, and the target image is output.

[0049] It should be noted that the execution subject of the embodiment can be a computing service device with data processing, network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a text-to-image generation device, etc. capable of realizing the above functions. In the following, the text-to-image generation device is taken as an example to explain the embodiment and the following embodiments.

[0050] In order to better understand the technical solutions of the application, the following will be explained in detail in combination with the drawings in the specification and specific embodiments.

[0051] The application embodiment provides a control method of a text-to-image generation model, referring to Figure 1 , Figure 1 The flowchart of the first embodiment of the control method of the text-to-image generation model of the application.

[0052] In the embodiment, the control method of the text-to-image generation model includes steps S10-S60:

[0053] Step S10, based on the intrinsic knowledge of the model, the text prompt word is proposed, the image generation quality threshold corresponding to the text prompt word is matched according to the preset index and the expected reasoning cost, and the generation task constraint parameter is integrated.

[0054] In the embodiment, the model intrinsic knowledge is a set of implicit cognition internalized into parameters in the large language model training stage. The text prompt word is a natural language instruction proposed by the model itself to drive image generation. The preset index is a pre-set image generation quality judgment standard and cost control benchmark. The image generation quality threshold is the minimum quality qualified line that the image generation needs to reach. The inference cost expectation is the upper limit of acceptable resource and time consumption in the image generation process. The task constraint parameter is the core control basis of the image generation task formed after integrating the quality threshold and the inference cost expectation.

[0055] As an optional implementation, the fixed quality threshold and inference cost expectation corresponding to each type of prompt word in the preset index library are called. The core features of the text prompt word proposed by the model intrinsic knowledge are extracted. The features are accurately matched with the prompt word types in the index library to directly obtain the corresponding quality threshold and inference cost expectation. Then, the two types of data are integrated into task constraint parameters according to a fixed template, and the parameter is added with a prompt word type label. This method has fixed matching logic, simple execution process, fast response speed, and ensures the efficiency of the basic process.

[0056] As another optional implementation, the full-dimensional features of the text prompt word proposed by the model intrinsic knowledge are extracted, and the basic quality threshold and inference cost expectation of the preset index library are called. Then, the basic index is calibrated in combination with the generation data of the same type of prompt word in history, so that the quality threshold matches the complexity level of the prompt word, and the inference cost expectation matches the generation difficulty. Then, the rationality of the calibrated index is verified through a multi-dimensional verification algorithm. Then, it is integrated into the task constraint parameter with a dynamic calibration label. This method has strong index adaptability and can cover the individual needs of various prompt words.

[0057] Step S20, according to the text prompt word and its corresponding task constraint parameter, a set of target concept dimensions conforming to the general scene characteristics are screened and defined.

[0058] In the embodiment, the general scene characteristics are the core attribute characteristics common to various generation scenes. The target concept dimensions are a set of independent attribute dimensions supporting the generation of test prompt words.

[0059] As an optional implementation, according to the text prompt word and the task constraint parameter, a preset scene and dimension association rule library is called, and initial candidate dimensions are screened out according to the fixed scene feature and concept dimension mapping relationship in the scene and dimension association rule library. Then, based on the quality precision and cost requirement in the task constraint parameter, a fixed redundancy removal algorithm is executed to remove repeated or highly associated dimensions and retain basic and orthogonal core dimensions, and finally the core dimensions are defined to form the target concept dimensions. The method has fixed process and high execution efficiency, does not require complex computing resources, and can quickly output the target concept dimensions meeting the requirements in a conventional scene, thereby guaranteeing the efficient promotion of the subsequent test prompt word generation.

[0060] As another optional implementation, the text prompt word is subjected to deep semantic analysis, and the core generation requirement and the implied scene feature are extracted in combination with the task constraint parameter. The dimension usage data of the historical generation task and the cloud general dimension resource pool are called, and the adaptive candidate dimensions are dynamically screened out through feature association analysis. Then, the dimension orthogonality verification algorithm is used to judge the independence and task adaptability of the candidate dimensions in real time, remove the invalid dimensions and supplement the missing core dimensions, and finally the core dimensions are defined to form the target concept dimensions. The method has strong adaptability, can cover conventional and new scenes, and has high dimension precision.

[0061] In step S30, the target concept dimensions are combined in multiple dimensions, specific scene attribute values are assigned to the obtained combinations, and instantiation is performed to obtain test prompt words and task exploration decisions.

[0062] In this embodiment, the multi-dimensional cross combination is a hierarchical or free dimension fusion operation on each target concept dimension. The scene attribute value is a specific feature limit value assigned to each dimension combination. The instantiation is a semantic conversion process of converting the dimension combination and the attribute value into a standard generation instruction. The test prompt word is an instruction that can be directly input into a text-to-image generation model and has a clear generation direction. The task exploration decision is a core decision basis for guiding the subsequent subtask execution and tool selection.

[0063] As an optional implementation, first, the single dimension or double dimension is cross combined in priority according to the preset dimension combination priority rule, and the preset high-order complex dimension combination is removed. Then, the built-in scene attribute value fixed library is called, and the attribute values fixed in the library are matched for each dimension combination. Subsequently, the standardized semantic instantiation module is started, the dimension combination and the attribute value are converted into a natural language instruction with uniform format to obtain the test prompt word, and the corresponding task exploration decision is bound for each test prompt word according to the preset decision matching rule to form a one-to-one corresponding set of prompt words and decisions. The method has fixed combination logic, simple execution process, high generation efficiency and does not require additional computing power support, and can efficiently support the tool performance verification of simple subtasks.

[0064] As an alternative implementation, historical exploration data of dimension combinations is first retrieved, and the level and scope of the dimension combinations are dynamically adjusted based on the quality and cost requirements in the current generation task constraints. Simultaneously, a dynamically updated pool of scenario-based attribute values ​​is accessed to match more adaptable attribute values ​​for the dimension combinations. Then, the intelligent semantic instantiation module is activated, and based on the instruction adaptation features of the text-to-image generation model, the dimension combinations and attribute values ​​are transformed into natural language with model instruction adaptability, generating test prompt words. This method offers high flexibility, adaptability to various dimension exploration needs, and high model adaptability of the prompt words.

[0065] As an alternative implementation, the target concept dimensions are first divided into basic derivative levels, and the adaptation success rate of each dimension is retrieved from the tool's historical execution data. Basic dimension pairs are sorted from high to low according to this success rate, and basic dimension pairs with a success rate greater than or equal to a preset probability are prioritized to form basic test prompts. Then, derivative attribute dimensions are combined with the basic combinations using the same logic in a quadratic orthogonal combination, while simultaneously marking the exploration tags of the dimension combinations to be verified, generating a set of test prompts sorted by basic, derivative, and unverified. The exploration priority of the current dimension combination is associated with the tool's capability boundary information to generate a differentiated task exploration decision. This method's test prompts cover core attributes and have no redundant combinations, accurately matching the tool's existing capability boundaries, making subsequent subtask execution more focused on the direction of tool capability adaptation.

[0066] Step S40: Obtain image results by completing simple sub-tasks corresponding to the test prompt words, evaluate the alignment between the test prompt words and the image results, and perform learning to obtain adaptation knowledge.

[0067] In this embodiment, the simple subtask is a basic generation task with a single or low-order dimension combination. The image result is the visualization output after performing the simple subtask. Alignment degree is the semantic and feature matching degree between the test prompt and the image result. Learning is a cognitive process of extracting the task-tool adaptation rules based on the alignment degree evaluation conclusion; adaptation knowledge is a core cognitive set about the correlation between image generation task requirements and tool capabilities.

[0068] As an optional implementation, preset simple subtask execution rules are called, and simple subtasks corresponding to each test prompt word are started in sequence according to the rules. The execution flow information and the final output image result of the subtasks are recorded throughout. Then, a fixed multi-dimensional evaluation logic is called to calculate the alignment degree of the test prompt word and the image result from two aspects of semantic fit degree and feature matching degree, and a standardized evaluation conclusion is generated. Then, according to the preset knowledge refining template, the association information between the tool capability and the subtask demand embodied in the evaluation conclusion is converted into structured adaptive knowledge, and the solidification storage of knowledge is completed. This method has fixed execution logic, compact process connection and fast overall response speed, which guarantees the timeliness of knowledge output.

[0069] As another optional implementation, simple subtask information corresponding to the test prompt word is called and historical execution and evaluation data of similar subtasks are accessed. Simple subtasks are completed in sequence according to the subtask priority, and the execution data of this time and the historical data are associated and bound. Then, a dynamic evaluation logic is started, and the alignment degree of the test prompt word and the image result is determined from multiple dimensions of semantics, features and resource consumption, combined with historical alignment rules and this time execution features, to generate an evaluation conclusion with an association label. Then, based on the association label in the evaluation conclusion, deep rule refining is carried out combined with historical adaptive cases to form hierarchical adaptive knowledge with task adaptive priority, and the dynamic association link of the knowledge system is updated synchronously. This method has comprehensive evaluation dimensions, and the knowledge refining has depth and association, which can adapt to various task scenarios.

[0070] As another optional implementation, the basic configurations of all simple subtasks are called, and a real-time simple subtask pre-validation is performed on the tool. The pre-validation result is collected and the pre-validation result is cross-verified combined with the historical ability data of the tool. If all simple subtasks meet the pre-validation standard and the stability of the historical data meets the preset standard, it is determined that the tool has mastered all simple subtasks, and the agent starts complex task exploration while recording the pre-validation data to update the tool ability profile. If the pre-validation or the historical stability does not meet the requirement, the simple subtask is re-executed. This method can adapt to the dynamic change of tool capability, and the verification result is more accurate. The disadvantage is that additional pre-validation is required, which increases short-term resource consumption. The real-time capability boundary of the tool can be accurately controlled to avoid invalid complex task exploration caused by tool capability fluctuation, which improves the exploration accuracy, but occupies pre-validation resources and prolongs the exploration preparation period.

[0071] Step S50, based on the task exploration decision and the adaptive knowledge, and combined with the performance characteristics of the text-to-image tool and the current task execution history, a target text-to-image tool is selected from the text-to-image tool feasible set.

[0072] In this embodiment, the text-to-image tool performance features are inherent attributes of various types of generation tools, such as generation quality, scene adaptation range, resource consumption level, etc. The current task execution history is the relevant data of the tool call records, result scores, and execution time of the completed links in this generation process. The text-to-image tool feasible set is the tool set that meets the basic quality threshold of the generation task and the expected inference cost. The target text-to-image tool is the tool finally selected to execute the image generation task.

[0073] As an optional implementation, the preset tool screening weight rule is first called to convert the task complexity level in the task exploration decision and the tool adaptation dimension in the adaptation knowledge into fixed weights. Then the core indicators in the text-to-image tool performance features and the tool compliance rate data in the current task execution history are extracted, the comprehensive adaptation scores of each tool in the feasible set are calculated according to the weight rule, the tools with the highest scores are selected as the target text-to-image tools, and the corresponding execution parameters in the adaptation knowledge are bound to the tools. This method has fixed screening logic, simple execution process, and fast overall response speed.

[0074] As another optional implementation, the full-dimensional instructions of the task exploration decision and the hierarchical adaptation association information in the adaptation knowledge are first called. Then the tools in the text-to-image tool feasible set are associated and matched with the same type of task execution data in the current task execution history, the initially adapted tool subset is first screened out, and then multiple rounds of simulation verification are carried out based on the complexity requirements of the task exploration decision to verify the execution effectiveness of the tools in different task scenarios. At the same time, the adaptation judgment standard is adjusted in combination with the dynamic change data of the tool performance features, and finally the target text-to-image tool is selected from the tools that pass the verification and a customized execution scheme is generated. This method has high screening accuracy, and the adaptable tool capability dynamically changes and meets the complex task requirements.

[0075] Step S60, completing the image generation corresponding to the test prompt word by the target text-to-image tool, and outputting a target image.

[0076] In this embodiment, image generation is the process of converting text requirements into visual content by tools according to prompt words. The target image is the final visual result output after meeting the generation task constraint parameters and quality verification.

[0077] As an optional implementation, the target text to image tool is invoked, the initial running parameters are dynamically configured in combination with the current task execution history and tool performance characteristics, the test prompt word is input into the tool to generate the initial image in the first round. Subsequently, a multi-dimensional quality review process is started to synchronously check the semantic alignment degree, feature matching degree and generation stability. If the quality threshold is not reached, the running parameters are adjusted according to the review results, the image generation and review operations are repeatedly executed, and the final image is output after standardization processing, while the parameters, results and review data generated in each round are completely stored. This method has strong parameter adaptability, comprehensive review dimensions, accurate target image quality, and can cover complex generation requirements.

[0078] Exemplarily, in the text-to-image generation scenario, based on a new type of agent framework named OctoT2I, which redefines the agent T2I (text-to-image) task as the joint optimization of generation quality and reasoning efficiency, including two parts. One is a stateful, multi-round intelligent routing agent responsible for responding to user prompts. It queries the knowledge module (long-term tool capability) and the memory module (current task history), executes the "decision, action, evaluation" cycle, dynamically selects the most suitable T2I tool, until the quality threshold is met or the maximum round is reached. The second is a set of innovative self-evolution mechanisms. It is responsible for autonomously building and filling the "knowledge module" from scratch in the offline stage. The intelligent routing agent implementation idea: given a user input text prompt word; the agent's decision strategy is activated. It combines user prompts, long-term knowledge, and current memory to select a T2I model as a constraint to meet the quality threshold and as a goal to minimize cost; execute the selected tool to generate an image; the evaluation module evaluates the alignment of the input text prompt word and the output image and gives a continuous quality score; the tool used in the current round, the generated image, and the obtained evaluation score are stored in the memory module; if the evaluation score reaches the preset threshold or the maximum round is reached, terminate the cycle and output the image with the highest score; otherwise, return to step 1 and start the next round of decision-making. The offline knowledge acquisition implementation idea of self-evolution: the agent first defines a set of basic, orthogonal concept dimensions autonomously; the agent generates specific test prompts according to the combination of concept dimensions; uses the current T2I tool to execute these prompts to generate N images; the evaluation module calculates the score of each image and calculates the Pass@1 success rate of the prompt; the tool used in the current round, the generated image, and the obtained evaluation score are stored in the memory module; the agent analyzes the evaluation results and stores the prompts, Pass@1 success rates, and average time consumption in the knowledge base to form prompt exploration records and tool capability portraits; to avoid brute force search in the vast combination space, a pruning strategy for the exploration space is designed. The core principle is: only when a tool has been proven to master all simple subtasks, will the agent explore complex tasks composed of them. This makes the exploration process efficient and always focused on the capability boundary of each tool.

[0079] By explicitly constraining parameters, accurately screening dimensions and tools, the problems of fuzzy constraints, poor tool adaptation, and inefficient exploration in text-to-image generation are solved, improving the matching degree of generation results and demand, generation efficiency, and controlling the reasoning cost.

[0080] Based on any of the above embodiments, in Embodiment Two of the present application, the step S10 includes steps A11-A13:

[0081] Step A11, receive and verify the text prompt word, filter out invalid and illegal text content, and obtain a verified text prompt word.

[0082] In this embodiment, invalid text content is a prompt word with improper format or ambiguous semantics that cannot be recognized. Violation text content is a prompt word that violates the generation content control rules. Text prompt words that pass the verification are instructions that are compliant with the format, legal in semantics, and can support subsequent generation processes.

[0083] As an optional implementation, the user-submitted text prompt word is received, the built-in fixed verification rule library is invoked, the format integrity of the prompt word is verified, and it is checked whether it contains the basic elements of the generation subject and core features. Then, according to the semantic recognition rules fixed in the verification rule library, it is determined whether the semantics are clear and analyzable. Finally, the content matching is performed against the preset violation content keyword library, and the prompt words containing violation keywords are filtered out. The prompt words that meet the format integrity, clear semantics, and content compliance are marked as passing the verification. This method has fixed verification logic, simple execution process, fast response speed, and does not require complex computing power support.

[0084] Step A12, according to the text prompt word scene classification rules and historical configuration data in the preset index library, match the image generation quality threshold and inference cost expectation that are adapted to the text prompt word.

[0085] In this embodiment, the preset index library is a standardized data set that stores text prompt word scene classification rules and historical configuration data. The text prompt word scene classification rule is a preset judgment criterion for classifying different generation requirements by scene. The historical configuration data is the practical record of the image generation quality threshold and inference cost expectation corresponding to each scene in the past generation task. The image generation quality threshold is the semantic alignment and feature matching standard that the image generation result needs to meet. The inference cost expectation is the upper limit of the computing power and time consumption allowed in the generation process.

[0086] As an optional implementation, the text prompt word scene classification rules in the preset index library are invoked to complete the scene classification of the text prompt word. Then, all historical configuration data of the scene and similar scenes in the index library are extracted, the parameter variation law in the data is mined through a feature correlation analysis algorithm, and the parameters are modified in combination with the potential features of the current generation task to dynamically fit the image generation quality threshold and inference cost expectation that are highly adapted to the text prompt word. Finally, the fitting result is verified for reasonableness to complete the matching. This method has high matching accuracy, can adapt to regular and new scenes, and can avoid the limitations of historical data.

[0087] Step A13, according to the parameter packaging specification, the image generation quality threshold, the inference cost expectation, and the corresponding text prompt word are associated and bound, and the task constraint parameters are integrated.

[0088] In this embodiment, the parameter packaging specification is a standardized criterion for associating and binding various types of generated related parameters and integrating the format.

[0089] As an optional implementation, the dynamic tag system in the preset parameter packaging specification is invoked to generate exclusive associated tags for the text prompt word, image generation quality threshold, and inference cost expectation, respectively. Then, a label mapping algorithm is used to establish a deep association among the three, and according to the dynamic adaptation rules of the packaging specification, the display level and association weight of the parameters are adjusted for different types of generation requirements. After completing the label association and level adjustment, a full-dimensional check is started, and after confirming that the association logic and format meet the requirements, the task constraint parameters are integrated into a flexible and adaptable parameter.

[0090] Illustratively, in the text-to-image generation scenario, the offline knowledge acquisition (self-evolution mechanism) phase, the goal of this phase is to let the agent completely autonomously, from scratch, build a knowledge module about the T2I tools: the agent first autonomously defines a set of basic, orthogonal "concept dimensions" (such as color, shape, count, position, etc.), which constitute the basis of the exploration space; the agent iteratively executes the "propose, solve, evaluate, learn" (PSEL) cycle under the guidance of the "exploration space pruning" strategy for each T2I tool in the tool library: the agent "proposes" and instantiates a series of specific test prompt words according to the combination of concept dimensions. Use the T2I tool currently under investigation to execute N times for each proposed prompt word to generate N candidate images. Call the evaluation function to score the N images, and calculate the Pass@1 success rate (i.e. the probability of reaching the quality threshold in a single attempt) and the average time cost of the tool for this prompt word. The agent analyzes and summarizes the evaluation results, updates the raw data containing the success rate and time cost to the knowledge base, and forms a detailed "prompt exploration record" and an advanced "tool capability profile". This cycle continues until all tools and valuable dimension combinations are traversed. The exploration space pruning strategy ensures that the agent only explores more complex combined tasks after mastering simple sub-tasks, thereby ensuring efficiency; the online reasoning routing (responding to users) phase, the goal of this phase is to respond to users' real-time prompts to generate images in a way that optimizes quality and efficiency in collaboration: receive user prompts: the system receives the user's input text prompt and the expected quality (threshold) and cost; the "decision strategy" module of the agent is activated. It simultaneously queries the long-term knowledge module (to obtain the general capabilities of the tools), the cost, and the short-term memory module (to obtain the execution history of the current prompt word); based on the above information, the decision strategy selects a tool with the lowest estimated cost from all feasible tools that can meet the quality threshold; the system executes the selected tool to generate an image; the "evaluation module" scores the alignment of the generated image with the user's prompt, resulting in a score; if the quality threshold is met, the loop terminates, and the system returns the image as the final output to the user; if not, the system adds the results of this attempt to the memory module, and then returns to step 2 to start a new round of decision-making. Key modules and technologies: self-evolution mechanism: this mechanism enables the agent to learn autonomously without human annotation and supervision. The key lies in the PSEL cycle and the exploration space pruning strategy, which makes the knowledge acquisition process data-driven, efficient, and scalable. Dual-module memory system: a long-term database built through offline self-evolution. It stores the performance characteristics, applicable scenarios, and reasoning costs of all tools. A short-term workspace that is reset for each new prompt. It records the execution history (tools, scores, image tuples) of the current task and the best result so far.Cost and quality trade-off based decision strategy: at each decision round, based on the current state, this strategy first filters out all tools whose estimated quality meets the threshold to form a feasible set. Then, it selects the tool with the lowest estimated cost from the feasible set, which guarantees the maximum efficiency under the premise of meeting the user's quality requirements.

[0091] Due to the precise verification and scenario parameter matching, the problems of insufficient text prompt word compliance and poor constraint parameter adaptability are solved, and the constraint precision and content compliance of text-to-image generation are improved.

[0092] Based on any of the above embodiments, in Embodiment Three of the present application, the step S20 comprises steps B11-B13:

[0093] Step B11, extract the core semantic features of the text prompt word, combine the quality requirements and cost requirements in the task constraint parameters, determine the scene type corresponding to the text prompt word, and integrate to form an analysis result.

[0094] In this embodiment, the core semantic features are key semantic information that can reflect the core generation demand in the text prompt word. The scene type is the scene classification corresponding to the generation demand of the text prompt word. The analysis result is a comprehensive data set formed by integrating the core semantic features, scene type and constraint requirements.

[0095] As an optional implementation, a deep semantic analysis module is started to perform full-dimensional semantic mining on the text prompt word to extract its core semantic features, while the quality requirements and cost requirements in the task constraint parameters are retrieved. Then, the historical analysis data and the scene feature library are accessed, and the core semantic features, quality and cost requirements, and historical scene data are compared through a feature association algorithm to dynamically determine the scene type corresponding to the text prompt word. Finally, according to the dynamic adaptation rules, various types of information are deeply associated and formatted to complete the integration of the analysis result. This method has comprehensive analysis dimensions, can adapt to new semantics and special constraints, has high scene determination accuracy, and greatly improves the adaptation degree of the analysis result and the subsequent process.

[0096] Step B12, based on the association rules of different scenes and concept dimensions in the concept dimension library, filter out the initial concept dimensions that adapt to the analysis result.

[0097] In this embodiment, the concept dimension library is a standardized data set that stores the mapping relationship between various scenes and corresponding concept dimensions. The association rules of scenes and concept dimensions are preset criteria that define which basic attribute dimensions are suitable for different generation scenes. The initial concept dimensions are a set of basic attribute dimensions selected from the concept dimension library that adapt to the current analysis result.

[0098] As an optional implementation, the associated rules in the concept dimension library and the historical dimension filtering data are invoked to perform full-dimension semantic mining on the analysis result to extract explicit scene types and implicit semantic features. Then, the comprehensive features of the analysis result are compared in depth with the dimension attributes in the concept dimension library through a feature association algorithm, and the dimension adaptation rules of the same type of scenes in the historical data are combined for correction to dynamically filter out the attribute dimensions that adapt to the explicit and implicit requirements. Finally, the dimensions with redundant relevance are removed through dimension orthogonality verification, and are integrated into the initial concept dimension. This method has strong adaptability, can cover new scenes and implicit semantic requirements, and has high dimension filtering accuracy.

[0099] Step B13, de-redundancy and priority sorting of the initial concept dimension according to the task constraint parameters, removing the dimensions that do not match the task constraints and retaining the basic and orthogonal core dimensions to obtain the target concept dimension.

[0100] In this embodiment, de-redundancy is an operation of removing repeated or strongly associated dimensions in the initial concept dimension. Priority sorting is a process of hierarchically dividing the importance of the generated task according to the dimensions. The dimensions that do not match the task constraints are attribute dimensions that do not meet the quality or cost requirements. The basic and orthogonal core dimensions are key attribute dimensions that are independent of each other and have no association.

[0101] As an optional implementation, the dimension association analysis module is started, the quality and cost requirements in the task constraint parameters are combined, the threshold of the redundant dimensions in the initial concept dimension is dynamically determined, and the de-redundancy operation is performed on the strongly associated dimensions that exceed the dynamic threshold. Then, the historical dimension adaptation data is invoked, each dimension is given an importance weight that matches the task constraints through a weight assignment algorithm, and the priority sorting is completed according to the weight. Subsequently, the dimension constraint adaptation verification is performed to verify the quality and cost matching degree of each dimension one by one, and the dimensions that do not match are removed. Finally, the orthogonal verification is performed on the remaining dimensions to retain only the basic core dimensions that are independent of each other, and the target concept dimension is integrated. This method has high processing accuracy, can adapt to various task constraints, has strong dimension orthogonality, and improves the adaptation and fitting degree of the target concept dimension and the generated task.

[0102] Exemplarily, in the text-to-image generation scenario, the text prompt word "generate a retro-style city night scene picture" input by the user is extracted, the core semantic features of "retro style, city, night scene" are extracted through the BERT semantic analysis model, the corresponding "retro style, city landscape scene" is determined by combining the image generation quality threshold of 86% (semantic alignment degree ≥ 86%, feature matching degree ≥ 86%) and the inference cost expectation of ≤ 30 seconds in the task constraint parameter, and the analysis result is integrated. Based on the rules of "retro style scene associated style, scene element, color tone, light and shadow" in the concept dimension library, the four initial concept dimensions that are adapted to the analysis result are selected. According to the task constraint parameter, it is determined that there is no redundant dimension by using the cosine similarity algorithm, and the influence priority on the generation quality is sorted as style > scene element > color tone > light and shadow, the "detail texture" dimension that does not match the cost requirement (the adaptation cost is more than 30 seconds) is removed, and the four core dimensions that are basic and orthogonal are reserved, to obtain the target concept dimension.

[0103] Due to the dimension screening through semantic accurate analysis and constraint adaptation, the problems of poor adaptability and redundancy of concept dimensions in text-to-image generation are solved, and the accuracy of the target concept dimension and the efficiency of the subsequent test prompt word generation are improved.

[0104] Based on any of the above embodiments, in the fourth embodiment of the present application, the step S30 comprises steps C11-C13:

[0105] Step C11, according to the preset dimension combination rule, multidimensionally cross-combines the target concept dimension, combines the priority judgment strategy, and generates a dimension combination set arranged according to the exploration priority.

[0106] In this embodiment, the preset dimension combination rule is a core criterion pre-set for regulating the association and collocation between target concept dimensions. The priority judgment strategy is a judgment logic for determining the value of dimension combination exploration. The exploration priority is a hierarchical sequence of dimension combinations arranged according to the exploration value. The dimension combination set is a structured data set integrating all dimension combinations that meet the rules and arranged according to the exploration priority.

[0107] As an optional implementation, the dimension association taboo and allowed range in the preset dimension combination rule are called, the full cross-combination of the target concept dimension is carried out according to the rule, and a preliminary dimension combination pool is generated. Then, the fixed priority weight in the priority judgment strategy is called, the priority weight is divided according to the importance of the basic attributes of the dimensions, the weight of each combination in the dimension combination pool is calculated by accumulation, and the dimension combination set arranged according to the exploration priority is generated by sorting the accumulation results from high to low. At the same time, each combination is labeled with the corresponding weight score and dimension composition information. This method has clear combination logic, fixed sorting rule, simple and short execution process, and guarantees the efficiency of the basic exploration process.

[0108] Step C12, according to the attribute value domain corresponding to different concept dimension combination in the dimension combination set, each dimension combination is given a quantitative scenario attribute value, and a scenario dimension combination is obtained.

[0109] In this embodiment, the attribute value domain corresponding to different concept dimension combination is the preset range interval of the attribute value that can be selected by each type of dimension combination. The quantitative scenario attribute value is a dimension combination attribute parameter that has a specific scenario representation and can be quantitatively defined. The scenario dimension combination is a dimension combination form that has an objectified attribute feature after the dimension combination is given a quantitative scenario attribute value.

[0110] As an optional implementation, the initial attribute value domain corresponding to each dimension combination in the dimension combination set is called. The assignment effect data of the historical similar dimension combination is accessed, and the initial attribute value domain is dynamically calibrated in combination with the exploration priority of the dimension combination, the value domain range is expanded or contracted to match the actual exploration demand. Then, the scenario attribute value with quantitative characteristics and exploration guidance is selected from the calibrated value domain to give each dimension combination, and the historical adaptation effect label is bound to the attribute value. After the assignment is completed, the scenario dimension combination with dynamic calibration label is formed, and the associated rule of value domain calibration is updated synchronously. This method has strong assignment adaptation and can accurately match the exploration demand of the dimension combination, improving the accuracy and practicality of the subsequent test prompt word.

[0111] Step C13, converting the scenario dimension combination and its corresponding attribute value into a natural language description conforming to the text-to-image generation task specification, forming the test prompt word, and generating the task exploration decision corresponding to the test prompt word according to the preset exploration decision rule.

[0112] As an optional implementation, the core requirements of the text-to-image generation task specification are called and the instruction conversion data of the historical similar dimension combination is accessed, the scenario dimension combination and its attribute value are converted into a natural language description with an exploration value label and form a test prompt word in combination with the data calibration conversion rule. Then, the preset exploration decision rule is called and the execution feedback data of the historical decision is associated, the task exploration decision with differentiated exploration priority is generated for the test prompt word, the dynamic association link is established for the prompt word and the decision, and the calibration logic of the conversion decision is updated. This method has strong conversion expression adaptability, the decision guidance is in line with the actual exploration demand, and can cover the conversion scenarios of various dimension combinations.

[0113] Exemplarily, in the text-to-image generation scenario, the target concept dimensions are set as "image style", "core subject", and "color tone", the preset dimension combination rule is to cross-combine one sub-dimension of each dimension, and the priority determination strategy is to score and sort according to the dimension correlation degree of 0-10 points. First, sub-dimensions such as "painting style", "mechanical device", and "warm yellow tone" are extracted from the three types of dimensions, 12 groups of dimension combinations are cross-generated, and after priority scoring, "painting style, mechanical device, warm yellow tone" gets 9.2 points ranked first, and "watercolor style, mechanical device, cold blue tone" gets 8.5 points ranked second, forming a dimension combination set ranked according to the exploration priority. Then, the attribute value domain of each combination is retrieved, the first ranked combination is assigned with the quantitative scene attribute value "painting style weight 0.8, mechanical device precision 0.9, warm yellow tone saturation 70%", and the scene dimension combination is obtained. Finally, the scene dimension combination is converted into the test prompt word "generate an image with painting style, high-precision mechanical device, and 70% saturation warm yellow tone" according to the text-to-image generation task specification, and the task exploration decision "preferentially call the Stable Diffusion model and execute the high-precision generation sub-task" is generated according to the exploration decision rule, and the remaining combinations are also converted and generated according to the process.

[0114] Due to the process of standardized dimension combination, quantitative attribute assignment, and directional decision generation, the problems of no system in test prompt words, disordered exploration priority, and low adaptability of decision and prompt words in text-to-image generation are solved, and the scene adaptation accuracy of test prompt words, the orderliness of task exploration, and the matching degree of prompt words and generation decisions are improved.

[0115] Based on any of the above embodiments, in the fifth embodiment of the present application, the step S40 includes steps D11-D13:

[0116] Step D11, according to the preset sub-task execution rule, sequentially complete the simple sub-tasks corresponding to the test prompt word, and record the tool calling information and execution time consumption of each simple sub-task, to obtain the image result and its corresponding execution process data.

[0117] In this embodiment, the preset sub-task execution rule is a standardized process criterion for guiding the sequential execution of simple sub-tasks. The tool calling information is the type and configuration data of the tool called when executing the sub-task. The execution time consumption is the overall time consumption of completing the sub-task. The image result is the visual content generated by the sub-task. The execution process data is the whole-process recording data integrating the tool calling information and the execution time consumption.

[0118] As an optional implementation, the preset subtask execution rules and historical subtask execution data are called, and according to the parallel adaptation conditions in the rules and the resource consumption law of the historical data, the simple subtasks corresponding to the test prompt words are divided into multiple batches of tasks that can be executed in parallel. The execution processes of the multiple batches of subtasks are started at the same time, and the tool call information and execution time of each subtask are collected in parallel through a synchronous recording module during the execution process. After all batches of tasks are completed, all image results are collected uniformly. Then, the scattered execution data is sorted into standardized execution process data through a data integration algorithm. This method has high overall execution efficiency, can greatly shorten the task cycle, and is suitable for multi-task concurrent scenarios.

[0119] Step D12, the multi-dimensional alignment score of the test prompt words and the corresponding image results in content semantics, feature attributes and scene matching degree is calculated through a preset evaluation algorithm, and the single generation compliance probability of each simple subtask corresponding to the test prompt words is calculated to generate a comprehensive evaluation result.

[0120] In this embodiment, the preset evaluation algorithm is a standardized calculation logic for determining the matching degree of the test prompt words and the image results. The content semantics is the semantic consistency of the prompt words and the image in the core generation requirement. The feature attribute is the matching degree of the attribute specified by the prompt words and the presentation attribute of the image. The scene matching degree is the fitting degree of the image and the scene set by the prompt words. The multi-dimensional alignment score is a matching quantitative value calculated by integrating the above three indicators. The single generation compliance probability is the frequency ratio of the subtask generation result meeting the quality threshold. The comprehensive evaluation result is an evaluation conclusion formed by integrating the multi-dimensional alignment score and the compliance probability.

[0121] As an optional implementation, the preset evaluation algorithm is called and historical evaluation data is accessed, and the weight distribution of content semantics, feature attributes and scene matching degree is dynamically adjusted in combination with the evaluation accuracy of different dimensions in the historical data. Then, the test prompt words and the image results are cross-dimensionally verified in multiple rounds to eliminate the errors of single-dimensional determination and to calculate accurate multi-dimensional alignment scores. Then, the compliance determination standard is corrected in combination with the quality requirement in the current task constraint parameter, and the single generation compliance probability of each simple subtask is calculated. Finally, the alignment score and the compliance probability are deeply associated through a data calibration algorithm to generate a comprehensive evaluation result with a scene adaptation label. This method has comprehensive evaluation dimensions, strong weight adaptation and high result accuracy.

[0122] Step D13, according to the exploration space pruning strategy, it is judged whether the current text-to-image tool has stably mastered all simple subtasks corresponding to the test prompt words, and based on the alignment scores and compliance probabilities of each dimension in the comprehensive evaluation result, targeted learning is carried out to extract the adaptation knowledge between the tool capability and the subtask requirement, and the task exploration decision is generated according to the judgment result.

[0123] In this embodiment, the exploration space pruning strategy is a standardized criterion for reducing ineffective task exploration range and focusing on core exploration direction. Stable mastery is the ability of the tool to continuously meet the quality threshold of the generated task results. The judgment result is a quantitative or qualitative conclusion on the ability state of the tool.

[0124] As an optional implementation, the core criterion of the exploration space pruning strategy is called and the historical ability data of the tool is accessed. The alignment score and the compliance probability of the current comprehensive evaluation result are cross-verified with the historical data, and whether the tool stably masters all simple subtasks is determined according to the data trend. Then, the multi-dimensional correlation learning module is started, the matching differences of each dimension in the evaluation data are bound and analyzed with the performance characteristics of the tool, and the hierarchical adaptive knowledge with the ability change trend is refined. Finally, the task exploration decision with differentiated exploration priority is generated by combining the determination result of the tool mastering the subtask with the ability boundary information in the adaptive knowledge, and the dynamic correlation link of the adaptive knowledge is updated synchronously. This method has high determination accuracy, dynamic adaptability of knowledge, and decision guidance more in line with actual ability.

[0125] Exemplarily, in the text-to-image generation scene, according to the preset subtask execution rule of "serial execution according to the priority of test prompt words", the simple subtasks corresponding to the three test prompt words "generate 1950s retro railway lamp image", "generate 3500K warm color retro scene image", etc. are completed in turn, the Stable Diffusion v1.5 and Midjourney v6 tools are called, the tool configuration parameters (resolution 1024x768, sampling step 20) and execution time (25 seconds, 28 seconds, 26 seconds) are recorded, and three image results and integrated execution process data are obtained. Through the preset evaluation algorithm, the content semantic alignment scores of the three groups of data are 89, 87 and 90 respectively, the feature attribute scores are 88, 86 and 89 respectively, the scene matching scores are 90, 88 and 91 respectively, the average alignment score is 88.5, the single generation compliance probability (quality threshold 85) is 92%, and the comprehensive evaluation result is generated. According to the exploration space pruning strategy, it is determined that the current tool stably masters all simple subtasks (compliance probability ≥ 90%), and the task exploration decision of "starting three-order dimension combination complex task exploration" is generated.

[0126] Due to the ordered execution, multi-dimensional evaluation and accurate determination, the problems of chaotic subtask execution, one-sided evaluation and blind exploration are solved, and the process standardization, evaluation accuracy and exploration effectiveness of text-to-image generation are improved.

[0127] Based on any of the above embodiments, in the embodiment six of the present application, the step S50 includes steps E11-E13:

[0128] Step E11, retrieve the text-to-image tool performance features and the current task execution history, integrate the task exploration decisions and their corresponding adaptation knowledge, tool performance features, and task execution history to obtain tool selection comprehensive reference data.

[0129] In this embodiment, the tool selection comprehensive reference data is a standardized data set formed by integrating the above three types of information to support target tool screening.

[0130] As an optional implementation, first, retrieve the structured data of text-to-image tool performance features and current task execution history according to the preset classification rules. Then extract the core orientation labels in the task exploration decisions and the tool adaptation dimension information in the adaptation knowledge, fill and integrate the four types of information according to the fixed field mapping template, assign preset priority weights to each type of information and label the corresponding association, and finally generate tool selection comprehensive reference data with priority labels, while completing standardized storage of the integrated data according to the preset format. This method has fixed integration logic, simple execution process, and fast overall response speed.

[0131] As another optional implementation, first, analyze the complexity level of the task exploration decisions, and dynamically set the weight parameters of the tool performance features and the task execution history according to the level. Retrieve the dimension adaptation capability data in the tool performance features and the corresponding dimension records in the task execution history, and couple and match the capability data and the records of the same dimension to calculate the matching score of the capability and demand of each tool. Then, combine the dynamic weights to weight and integrate the tool performance score and the historical matching score to form tool selection comprehensive reference data with matching score labels. This method greatly improves the adaptability of the reference data, can accurately match the complexity requirements of the current task, reduces the blindness of tool selection, and improves the efficiency of tool adaptation.

[0132] Step E12, according to the tool selection comprehensive reference data and the task constraint parameters, screen all text-to-image tools, propose tools with quality not meeting the standard and cost exceeding the expectation, and obtain the feasible set of text-to-image tools.

[0133] As an optional implementation, the full-dimension detailed data of the comprehensive reference data is called, and the core requirements of the task constraint parameters are combined to give dynamic weights to the quality standard degree, cost controllability, and exploration decision adaptability. Then, multi-dimensional cross verification is carried out on each text-to-image tool, and the historical quality stability, cost fluctuation range, and adaptation degree of the task exploration decision of the tool are verified synchronously. The comprehensive adaptation score of each tool is calculated by weighting, and the tools with substandard comprehensive scores are eliminated. At the same time, the potential ability of the tool is supplemented and evaluated, and the tools that meet the evaluation standard are included in the candidate range, and finally the text-to-image tool feasible set is formed. This method has comprehensive screening dimensions, strong adaptability, and can mine the potential ability of the tool.

[0134] Step E13, based on the cost-quality trade-off decision algorithm, the target text-to-image tool that can meet the quality requirements of the current task and realize the optimal inference cost is selected from the text-to-image tool feasible set according to the judgment of the task complexity in the task exploration decision.

[0135] In this embodiment, the cost-quality trade-off decision algorithm is a standardized decision logic for balancing image generation quality and inference cost. The optimal inference cost is the state of minimizing inference computing power and time consumption under the premise of meeting the quality requirements.

[0136] As an optional implementation, the cost-quality trade-off decision algorithm is called and the historical execution data of the tool is accessed, and the judgment of the task complexity in the task exploration decision is decomposed into semantic complexity, dimension combination complexity, etc. The trade-off weight of quality and cost is dynamically adjusted according to each subdivision dimension. Then, multi-round quality and cost simulation calculation is carried out on each tool in the text-to-image tool feasible set, and the quality standard stability and cost fluctuation range of the tool under different complexity scenarios are simulated. Then, the comprehensive adaptation score of each tool is calculated by algorithm weighting, and the tool with the highest comprehensive score that meets the quality requirements and the optimal inference cost is selected as the target text-to-image tool. This method has comprehensive trade-off dimensions, can adapt to various task complexities, has high tool screening accuracy, can cover special task requirements, and improves the quality and cost adaptability of the tool.

[0137] Exemplarily, referring to Figure 2 , Figure 2 The schematic diagram is selected for the model of the present application. When the user inputs the first user prompt word "A photo of a snowboard", the system triggers the operation of reading the knowledge module, and in the knowledge module, the system selects the most suitable image generation model according to the user prompt word, and the selected model is used to generate the image corresponding to the user prompt word. <thinking>In the link, the inference is made according to the capability profiles and the raw data, and the conclusion is that "sd-turbo is the fastest in speed, and other models also have very high speed but are relatively slower", and then in <answer>The explicit model in the tag is sd-turbo, then the operation of "generate an image using sd-turbo" is performed, and the corresponding snowboard image is output. When the user inputs the second user prompt word "Generate an image containing 5 hamburgers", the system again starts the flow of reading the knowledge module, and in <thinking>The link confirmed through raw data summary that all models except flow-gro were fundamentally unable to cope with the "five [objects]" type of prompt, flow-gro was the only justifiable choice, followed by <answer>The model is identified as flow-gro in the label. The operation of "generate image using flow-gro" is executed, and the corresponding hamburger image is output. During this process, the strategy module relies on the information of the knowledge module to complete the response and reasoning for different user prompt words and select the most suitable generation model.

[0138] By integrating data, employing tiered filtering, and balancing cost and quality, the problem of blindly selecting text-to-image tools and the imbalance between cost and quality has been solved, thereby improving the accuracy of tool adaptation and the cost-effectiveness of the generated tasks.

[0139] Based on any of the above embodiments, in Embodiment Seven of this application, step S60 includes steps F11 to F13:

[0140] Step F11: Call the dedicated adaptation interface of the target text to image tool, configure the tool running parameters according to the quality and accuracy requirements in the task constraint parameters, execute the complex image generation operation corresponding to the test prompt word, and obtain the initial generated image and its corresponding process data.

[0141] In this embodiment, the tool's runtime parameters are the core configuration information that drives the tool to complete the generation operation. Complex image generation is a high-order image generation task under multi-dimensional combined requirements. The initial generated image is the visual result output by the tool after performing the generation operation. Process data is a complete record of information such as parameter calls, execution time, and resource consumption during the generation operation.

[0142] As an optional implementation, a dedicated adaptation interface for the image tool is invoked, and historical interaction data of the interface is accessed. Combining the quality and accuracy requirements in the task constraints with historical adaptation cases, multiple sets of candidate tool operation parameters are dynamically generated through a parameter iteration algorithm. The candidate parameters are then pre-verified using the interface link and the generation effect is simulated. The parameters with the highest adaptability are selected to complete the tool operation configuration. Subsequently, a test prompt is submitted to initiate a complex image generation operation. During the operation execution phase, full-dimensional process data is acquired through a multi-node data acquisition module. After generation, the initial generated image and process data are correlated and verified to ensure the consistency between data and images. This method has strong parameter adaptability, can handle dynamic quality and accuracy requirements, provides comprehensive process data acquisition, and can adapt to scenarios with fluctuations in the interface link.

[0143] Step F12: By comparing the quality threshold in the task constraint parameters, verify the semantic alignment and feature matching degree between the initially generated image and the test prompt words, and generate a judgment result.

[0144] In this embodiment, the quality threshold is a semantic and feature matching qualification standard that the image generation must meet, as specified in the task constraint parameters. The judgment result is a qualitative or quantitative conclusion on whether the initially generated image meets the quality threshold.

[0145] As an optional implementation, the quality threshold of the task constraint parameter is called and historical verification data is accessed, and the verification accuracy in different dimensions in the historical data is combined to give dynamic verification weights to the semantic alignment degree and the feature matching degree. Then, multi-round cross verification is carried out on the initially generated image, the implicit semantic of the image is extracted through deep semantic mining, and the extracted implicit semantic is compared with the prompt word. Then, the complete features of the image are obtained through feature fullness recognition, and the features of the image are matched with the prompt word, and the feature baseline of the same type of qualified image is introduced for auxiliary verification. The weighted verification result is compared with the quality threshold to generate a preliminary judgment. After attribution analysis of the deviation term, a judgment result with optimization suggestions is generated. This method has comprehensive verification dimensions, can adapt to complex semantic and special feature requirements, has high judgment accuracy, and improves the guidance value of the result for subsequent optimization.

[0146] As another optional implementation, the basic qualified threshold and the optimization potential threshold in the task constraint parameter are first called, and the semantic alignment degree and the feature matching degree of the initially generated image are verified. If both are lower than the basic qualified threshold, a non-qualified judgment result is generated. If they are between the two thresholds, a judgment result with optimization potential is generated and the adjustable parameters are labeled. If it is higher than the optimization potential threshold, a qualified judgment result is generated. This method has more flexible judgment results, can provide optimization direction for part of the generated results close to the standard, and at the same time provides accurate parameter direction for optimization adjustment, improving the overall efficiency of the generation task.

[0147] Step F13, if the judgment result meets the quality threshold, the initially generated image is standardized and output as the target image.

[0148] As an optional implementation, the preset standardization processing rule is called and historical image delivery data is accessed, and a dedicated standardization scheme is dynamically generated for the initially generated image that meets the quality threshold based on the delivery standard adaptation law in different scenarios in the historical data. The dedicated standardization scheme includes the format, resolution, and color calibration parameters adapted to the attributes. Then, the hierarchical processing operation is performed according to the dedicated standardization scheme, the basic format conversion is first completed, and then the core features of the image are accurately calibrated. After processing, the image is verified through multi-dimensional compliance checking and compared with historical delivery cases, and the image is determined as the target image after confirming that it meets the standard. Then, the output is completed according to the dynamic adaptation link of the receiving end. This method has strong adaptability, can meet individualized delivery requirements, has high standardization accuracy, and improves the adaptation and compatibility of the target image and the receiving end.

[0149] Exemplarily, in the text-to-image generation scenario, the REST special adaptation interface of the target text-to-image tool Stable Diffusion v1.5 is called, the quality and precision requirements of "resolution 1024x768, semantic alignment degree ≥ 88%" in the task constraint parameters are configured, the tool running parameters of the sampling step 25 and CFG Scale 7.5 are configured, the complex image generation operation corresponding to the test prompt word "generate an image of a 1950s retro style, cast iron material retro street lamp matched with 3500K warm color tone" is executed, and the initial generated image and process data (execution time 26 seconds, GPU occupancy rate 78%, sampler Eulera) are obtained. According to the quality threshold of 88 points in the task constraint parameters, the semantic alignment degree of the initial generated image and the test prompt word is 90 points, and the feature matching degree is 89 points, and the determination result of "satisfying the quality threshold" is obtained. After standardizing the initial generated image (format conversion to PNG, compression to 2MB, color space sRGB calibration), the initial generated image is output as the target image.

[0150] Due to the accurate parameter configuration, double-dimensional quality checking and standardization processing, the problems of unstable generated image quality and non-uniform output format are solved, and the achievement standard rate and delivery standardization of text-to-image generation are improved.

[0151] Based on any one of the above embodiments, in the eighth embodiment of the present application, refer to Figure 3 , Figure 3 is a flowchart of the eighth embodiment of the control method based on the text-to-image generation model of the present application. After step F12, steps G11-G13 are further included.

[0152] In step G11, if the determination result is not satisfying the quality threshold, the initial generated image is associated with the tool used for this generation, the generation time, the review score, the semantic alignment degree and the feature matching degree, and single generation task data is obtained by integration.

[0153] In this embodiment, the review score is a secondary comprehensive quality score of the initial generated image. The single generation task data is a standardized data set that can be traced and analyzed after integrating all the above information.

[0154] As an optional implementation, the basic data integration rule is invoked, and the data analysis case of the historical non-compliance generation task is accessed, the key data dimensions for the generation failure attribution in the historical case are combined, and the data correlation dimensions are dynamically expanded for the current non-compliance task. Then, the initial generation image is associated with the generation tool, the generation time consumption, the review score, the semantic alignment degree, and the feature matching degree, and the parameter configuration and resource occupation in the generation process are supplemented as extended data, and the deep binding of the data in each dimension is realized through the data fusion algorithm. Then, the attributed label is marked on the integrated data, and the single generation task data with failure attribution guidance is obtained. This method integrates comprehensive dimensions, can adapt to various non-compliance scenarios, has high data correlation accuracy and attribution guidance value, and improves the guidance fit of data for subsequent generation strategy optimization.

[0155] Step G12, through a preset difference analysis algorithm, the specific causes of insufficient semantic alignment degree, feature matching deviation, and time consumption exceeding the standard in the single generation task data are located, and the task data attribution result containing the problem root and tool adaptation short board is output.

[0156] In this embodiment, the preset difference analysis algorithm is a standardized analysis logic for locating the causes of various quality and efficiency problems in the generation task data. The feature matching deviation is the problem that the image presentation features do not match the specified features of the prompt word. The task data attribution result is a standardized analysis conclusion integrating the problem root and tool short board.

[0157] As an optional implementation, the preset difference analysis algorithm is invoked, and the attribution data of the historical generation task is accessed, the single generation task data is divided into semantic features, tool parameters, resource consumption, and other subdivided dimensions, and the direct causes of problems in each dimension are located through multidimensional cross verification. Then, combined with the cause rules of similar problems in historical data, the direct causes are deeply traced to lock the problem root. At the same time, an adaptation model of tool capability and task demand is constructed, and the adaptation short board of the tool in semantic analysis, feature rendering, and algorithm scheduling is quantitatively analyzed. Finally, the attribution result is added with an optimization priority label to form a task data attribution result with precise guidance. This method analyzes comprehensive dimensions, can adapt to various generation problems, has high attribution accuracy, and has optimization priority guidance, which improves the guidance fit of the result for generation strategy iteration.

[0158] As another optional implementation, the semantic alignment, feature matching, and resource allocation data in the single-generation task data are extracted first. The causal chain analysis module is started to verify whether the insufficient semantic alignment is associated with the feature matching deviation, and then whether the feature matching deviation is related to the resource allocation parameters. Subsequently, the problems of each associated node are bound to the tool adaptation short board, and finally the task data attribution result containing the causal transmission link and the corresponding tool short board is output. The attribution result of this method is more in-depth, which can lock the deep transmission root of the problem, improve the effectiveness of strategy optimization, and reduce the resource consumption of repeated adjustment.

[0159] Step G13, according to the task data attribution result, a tool selection strategy, a tool generation precision, and a resource allocation parameter are formulated, and an optimization adjustment strategy for a new round of image generation is integrated to call and configure a corresponding text-to-image tool according to the optimization adjustment strategy to complete the image generation operation of the test prompt word.

[0160] In this embodiment, the tool selection strategy is the core criterion for guiding the selection of the tool for the new round of generation task. The tool generation precision is a key parameter for controlling the image generation quality and semantic fit. The resource allocation parameter is a configuration basis for allocating generation task computing power, time, and other resources. The optimization adjustment strategy is a comprehensive generation optimization scheme integrating tool selection, precision control, and resource allocation.

[0161] As an optional implementation, the full-dimensional detailed information of the task data attribution result is called and accessed into the historical optimization strategy data, and the strategy optimization rules of similar problems in the historical data are combined to generate multiple groups of candidate strategies for the three dimensions of tool selection, generation precision, and resource allocation. Then, the generation compliance rate and resource consumption ratio of each candidate strategy are calculated through a strategy simulation algorithm, and the single-dimensional strategy with the optimal comprehensive benefit is selected. Subsequently, the deep coupling of the three types of strategies is realized through a strategy fusion algorithm to form an optimization adjustment strategy with priority. Then, the strategy is dynamically selected according to the strategy, and the generation precision and resource parameters are configured in stages to execute the image generation operation of the test prompt word. This method has strong strategy adaptability, can cover deep implicit attribution problems, has high resource utilization rate, and improves the quality compliance rate and resource adaptability of the new round of generation.

[0162] Exemplarily, referring to Figure 4 , Figure 4 An algorithm flowchart for the present application is shown. First, a user prompt word (user prompt, i.e., "a rainbow-layered cake on a white styled table against a teal background" input by the user) is input, and the process is carried out in a goal-oriented manner (objective, i.e., the lowest cost or reaching the threshold of the goal) relying on a self-evolution mechanism (self-evolution mechanism, including initialization, exploration, instantiation, prompting, and other links). The initialization stage of the mechanism completes the concept space conversion, and then enters the exploration link through search (search), and then carries out stability search (stability search) and exploration (explore) operations in turn. Subsequently, the process data is transmitted to the decision mechanism (decision mechanism), combined with the information of the memory (memory) and knowledge (knowledge) modules, and the select operation is performed, the corresponding model is called from the model library (model library, including models with human-shaped marks, etc.), and the selected model (use selected model) is used to generate an image (generate image). At the same time, the prompt search record (prompt search record, including "two chairs in watercolor style", "a cat in watercolor style", "three green apples", "a vintage cap in the left bottom", and the corresponding matching degrees of 90%, 60%, 95%, and 80%) is called, N images (N images) are generated and integrated relying on the module (module). Then enter the evaluation (evaluation) link, check whether the final score is greater than the threshold value, if not, add the evolution module (add evolution module) and restart the attempt (restart attempt). At the same time, combined with the high tool capability (high tool capability, including speed and quality balance, processing a specified number of images, and other analysis dimensions), analysis (analysis) and learning (learning) are carried out. Finally, the image that meets the user's prompt word requirement is output (final output), and the whole process is completed.

[0163] Further, with reference to Figure 5 , Figure 5 The figure is an effect comparison diagram of the present application. For different prompts, the generation results of the model of the present application are compared with the effects of Idea2Img, GenArtist and ChatGen: when the prompt is "The concept had a carousel and the spherical ball were the objects of the magic show (this concept contains a carousel and a spherical ball, which are the props of the magic show)", Idea2Img generates a scene containing a carousel, GenArtist generates a scene of decorating ball-shaped objects, ChatGen generates a cone-shaped prop, and the model of the present application generates a magic show scene integrating a carousel and a spherical ball, which is more consistent with the core elements of the prompt. When the prompt is "A bicycle on the left of a man (a bicycle on the left of a man)", Idea2Img generates a picture of a man with a bicycle next to him, GenArtist generates a scene of multiple people and a bicycle, ChatGen generates a single bicycle, and the model of the present application generates a bicycle on the left side of a man, which accurately matches the position description. When the prompt is "A greenscooter is parked next to a blue vintage car (a green scooter is parked next to a blue vintage car)", Idea2Img and GenArtist generate pictures related to a green scooter, and ChatGen generates a low-matching-degree scene of a small scooter combined with a blue car, and the model of the present application generates a picture of a green scooter parked next to a blue vintage car, which matches both the elements and the position. When the prompt is "a suitcase on a surface of a boat (a suitcase on the surface of a boat)", Idea2Img generates a scene of a boat and a suitcase, GenArtist generates a boat and a large box, ChatGen generates a picture of a suitcase on a wooden board, and the model of the present application generates a picture of a suitcase on the surface of a boat, which highly matches the element scene. When the prompt is "A round bagel, nectarines and a breakfast (a round bagel, nectarines and a breakfast)", Idea2Img generates a breakfast combination containing a bagel and nectarines, GenArtist generates scattered breakfast ingredients, ChatGen generates a single food, and the model of the present application generates a complete breakfast combination containing a round bagel and nectarines, which is complete in elements and coordinated in scene.When the prompt word is "Five candles”, Idea2Img generates a decorative scene of multiple candles, GenArtist generates a magnificent multi-candle device, ChatGen generates a small number of candles, and the model of the present application generates an accurate arrangement of five candles, the number of which completely matches the description. Overall, the present application is superior to Idea2Img, GenArtist, and ChatGen in terms of element reduction, position matching, scene fit, and detail accuracy of the prompt word.

[0164] Due to the attribution and precision strategy optimization through data association, the problems of poor dimension attribute matching and low efficiency in generating animation are solved, and the dimension fit and resource utilization of animation generation are improved.

[0165] The present application provides a text-to-image generation device, which comprises at least one processor and a memory connected in communication with the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the control method of the text-to-image generation model in Embodiment I.

[0166] Reference will now be made to the drawings, in which Figure 6 which shows a structural schematic diagram of a text-to-image generation device suitable for implementing the embodiments of the present application. The text-to-image generation device in the embodiments of the present application can include but is not limited to mobile terminals such as mobile phones, notebook computers, image generation terminals, personal digital assistants (PDA, Personal Digital Assistant), tablet computers (PAD, Portable Application Description), portable multimedia players (PMP, Portable MediaPlayer), image generation computing platforms, and the like, as well as fixed terminals such as dedicated AI inference servers, desktop computers, and the like. Figure 6 The text-to-image generation device shown is only an example and should not impose any limitations on the functions and use range of the embodiments of the present application.

[0167] As Figure 6 As shown, the text-to-image generation device can include a processing apparatus 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage apparatus 1003 into a random access memory (RAM) 1004. In the random access memory 1004, various programs and data required for the operation of the text-to-image generation device are also stored. The processing apparatus 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input apparatus 1007 including, for example, a touch screen, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output apparatus 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; the storage apparatus 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication apparatus 1009. The communication apparatus 1009 can allow the text-to-image generation device to communicate with other devices wirelessly or by wire to exchange data. Although the text-to-image generation device with various systems is shown in the figure, it should be understood that all the systems shown are not required to be implemented or possessed. More or fewer systems can be alternatively implemented or possessed.

[0168] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by a communication apparatus, or installed from the storage apparatus 1003, or installed from the read-only memory 1002. When the computer program is executed by the processing apparatus 1001, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are performed.

[0169] The text-to-image generation device provided by the present disclosure adopts the control method of the text-to-image generation model in the above-mentioned embodiments, and can solve the technical problem of low image generation efficiency. Compared with the prior art, the text-to-image generation device provided by the present disclosure has the same beneficial effects as the control method of the text-to-image generation model provided by the above-mentioned embodiments, and other technical features in the text-to-image generation device are the same as the features disclosed in the previous embodiment method, which will not be repeated here.

[0170] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0171] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0172] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to execute the control method of the text-to-image generation model in the above embodiments.

[0173] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, radio frequency (RF), etc., or any suitable combination thereof.

[0174] The aforementioned computer-readable storage medium may be included in a text-to-image generating device; or it may exist independently and not assembled into a text-to-image generating device.

[0175] The computer readable storage medium described above carries one or more programs, when the one or more programs are executed by the text-to-image generation device, cause the text-to-image generation device to: propose a text prompt word based on intrinsic knowledge of a model, match an image generation quality threshold value corresponding to the text prompt word and an inference cost expectation according to a preset index, and integrate generation task constraint parameters; filter and define a set of target concept dimensions conforming to general scene characteristics according to the text prompt word and the corresponding task constraint parameters; multi-dimensionally cross-combine the target concept dimensions, assign specific scene attribute values to the obtained combinations and instantiate, obtain a test prompt word and a task exploration decision; obtain an image result by completing a simple subtask corresponding to the test prompt word, evaluate the alignment degree of the test prompt word and the image result and carry out learning, and obtain adaptive knowledge; based on the task exploration decision and the adaptive knowledge, combine text-to-image tool performance characteristics and current task execution history, and filter a target text-to-image tool from a text-to-image tool feasible set; complete image generation corresponding to the test prompt word through the target text-to-image tool, and output a target image.

[0176] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0177] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0178] The modules involved in the embodiments of the present application can be implemented in the form of software or hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0179] The computer readable storage medium provided by the present application is a computer readable storage medium, which stores computer readable program instructions (i.e. computer program) for executing the control method of the text-to-image generation model, and can solve the technical problem of low image generation efficiency. Compared with the prior art, the computer readable storage medium provided by the present application has the same beneficial effects as the control method of the text-to-image generation model provided by the above embodiments, and will not be described here.

[0180] The above only describes some embodiments of the present application, and does not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the present application specification and drawings, or direct / indirect application in other related technical fields is included in the patent protection scope of the present application.< / answer> < / thinking> < / answer> < / thinking>

Claims

1. A control method of a text-to-image generation model, comprising: The method comprises: The text prompt word based on the internal knowledge of the model matches the image generation quality threshold and the reasoning cost expectation corresponding to the text prompt word according to the preset index, and integrates the generation task constraint parameter; Extract the core semantic features of the text prompt word, combine the quality requirement and cost requirement in the task constraint parameter, determine the scene type corresponding to the text prompt word, and integrate the analysis result; Based on the association rules of different scenes and concept dimensions in the concept dimension library, the initial concept dimension suitable for the analysis result is selected; According to the task constraint parameter, the initial concept dimension is de-redundant and prioritized, the dimensions not matching the task constraint are removed, and the basic and orthogonal core dimensions are retained to obtain the target concept dimension; Multi-dimensional cross combination of the target concept dimension, assigning specific scenario attribute values to the obtained combination and instantiating, obtaining test prompt words and task exploration decisions; According to the preset subtask execution rule, the simple subtasks corresponding to the test prompt words are completed in turn, and the tool calling information and execution time of each simple subtask are recorded to obtain image results and their corresponding execution process data; Through the preset evaluation algorithm, the multi-dimensional alignment scores of the test prompt words and the corresponding image results in content semantics, feature attributes and scene matching degree are calculated, and the single generation compliance probability of each simple subtask corresponding to the test prompt word is calculated to generate a comprehensive evaluation result; According to the exploration space pruning strategy, it is judged whether the current text-to-image tool has mastered all simple subtasks corresponding to the test prompt word, and based on the alignment scores and compliance probabilities of each dimension in the comprehensive evaluation result, targeted learning is carried out to extract the adaptation knowledge between tool capability and subtask demand; Based on the task exploration decision, the adaptation knowledge, combining the performance characteristics of the text-to-image tool and the current task execution history, the target text-to-image tool is selected from the feasible set of the text-to-image tool; The image corresponding to the text prompt word input by the user is completed through the target text-to-image tool, and the target image is output. 2.The control method of a text-to-image generation model of claim 1, wherein, The step of generating the text prompt word based on the internal knowledge of the model, matching the image generation quality threshold and the reasoning cost expectation corresponding to the text prompt word according to the preset index, and integrating the generation task constraint parameter comprises: Receive and verify the text prompt word, filter out invalid and illegal text content, and obtain the text prompt word that passes the verification; According to the text prompt word scene classification rule and historical configuration data in the preset index library, the image generation quality threshold and the reasoning cost expectation suitable for the text prompt word are matched; According to the parameter packaging specification, the image generation quality threshold, the reasoning cost expectation and the corresponding text prompt word are associated and bound, and the task constraint parameter is integrated. 3.The method of claim 1, wherein, The step of multi-dimensional cross combination of the target concept dimension, assigning specific scenario attribute values to the obtained combination and instantiating, obtaining test prompt words and task exploration decisions comprises: According to the preset dimension combination rule, the target concept dimension is multi-dimensionally combined, a priority judgment strategy is combined, and a dimension combination set arranged according to an exploration priority is generated; According to the attribute value domain corresponding to different concept dimension combinations in the dimension combination set, a quantitative scenario attribute value is given to each dimension combination, and a scenario dimension combination is obtained; The scenario dimension combination and the corresponding attribute value are converted into a natural language description conforming to the text-to-image generation task specification, a test prompt word is formed, and a task exploration decision corresponding to the test prompt word is generated according to a preset exploration decision rule. 4.The method of claim 1, wherein, The step of filtering out a target text-to-image tool from the text-to-image tool feasible set based on the task exploration decision and the adaptive knowledge, combining the text-to-image tool performance characteristics and the current task execution history includes: The text-to-image tool performance characteristics and the current task execution history are called, the task exploration decision and the corresponding adaptive knowledge, tool performance characteristics, and task execution history are associated and integrated, and tool selection comprehensive reference data is obtained; According to the tool selection comprehensive reference data and the task constraint parameters, all text-to-image tools are filtered out, tools with quality not meeting the standard and cost exceeding the expectation are proposed, and the text-to-image tool feasible set is obtained; Based on a cost-quality trade-off decision algorithm, the target text-to-image tool that can meet the quality requirements of the current task and achieve the optimal reasoning cost is filtered out from the text-to-image tool feasible set according to the task complexity judgment in the task exploration decision. 5.The method of claim 1, wherein, The step of completing the image generation corresponding to the test prompt word by the target text-to-image tool and outputting a target image includes: A special adaptive interface of the target text-to-image tool is called, tool running parameters are configured according to the quality precision requirement in the task constraint parameters, a complex image generation operation corresponding to the test prompt word is performed, and an initial generated image and the corresponding process data are obtained; The quality threshold in the task constraint parameters is compared, the semantic alignment degree and the feature matching degree of the initial generated image and the test prompt word are checked, and a judgment result is generated; If the judgment result meets the quality threshold, the initial generated image is standardized and output as the target image. 6.The method of claim 5, wherein, After the step of comparing the quality threshold in the task constraint parameters, checking the semantic alignment degree and the feature matching degree of the initial generated image and the test prompt word, and generating a judgment result, the control method of the text-to-image generation model further includes: If the judgment result does not meet the quality threshold, the initial generated image, the tool used for generation, the generation time consumption, the review score, the semantic alignment degree, and the feature matching degree are associated, and single-generation task data is integrated and obtained; Through a preset difference analysis algorithm, the specific causes of insufficient semantic alignment degree, feature matching deviation, and time consumption exceeding the standard in the single-generation task data are located, and a task data attribution result containing the problem root and the tool adaptation short board is output. ​ According to the task data attribution result, a tool selection strategy, tool generation precision and resource allocation parameters are formulated, and an optimized adjustment strategy for a new round of image generation is formed, so as to call and configure a corresponding text-to-image tool according to the optimized adjustment strategy to complete the image generation operation of the test prompt word.

7. A text-to-image generation device, comprising: The text-to-image generation device comprises a memory, a processor and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the control method of the text-to-image generation model according to any one of claims 1 to 6.

8. A storage medium, characterized by The storage medium is a computer readable storage medium, and the storage medium stores a computer program. When the computer program is executed by the processor, the steps of the control method of the text-to-image generation model according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Video generation and screening method and device, equipment and medium

    CN121000948A

  • Method and device, apparatus, vehicle and medium for generating a classification model

    DE102024211705A1