Method and system for automatically generating marketing and distribution multi-modal data

By deeply integrating diffusion models with large-scale language models, and combining domain knowledge and multiple conditional controls, logically consistent and power-compliant multimodal data is generated, solving the problems of data quality and cross-modal alignment difficulties in power operation and maintenance, and realizing efficient construction of training datasets.

CN121834658APending Publication Date: 2026-04-10STATE GRID SHANGHAI MUNICIPAL ELECTRIC POWER CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, the generation of multimodal data for power system marketing and distribution business lacks cross-modal semantic consistency and a deep understanding of business rules, resulting in low-quality generated data and a scarcity of high-quality labeled data, making it difficult to generate training data that conforms to power specifications.

Method used

A method that deeply integrates diffusion model and large language model is adopted. The large language model with domain knowledge enhancement is used to analyze user needs, generate structured image control parameters and text generation outline, and combine multiple fine-grained condition control diffusion model to generate high-fidelity business images. Visual semantic alignment and business logic consistency audit are performed, and iterative correction is carried out until the preset quality standard is met.

Benefits of technology

The generated data is logically consistent and conforms to power industry standards. It can efficiently construct training datasets, solve the problem of scarcity of high-quality training data, and improve the R&D efficiency of intelligent applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834658A_ABST
    Figure CN121834658A_ABST
Patent Text Reader

Abstract

The invention discloses a marketing and distribution multi-modal data automatic generation method and system based on fusion of a diffusion model and a large language model, and belongs to the technical field of artificial intelligence and electric power informatization. Analyzing the natural language demand of a user by using a marketing and distribution field knowledge enhanced large-scale language model, and outputting a structured image control parameter set and a text generation outline; constructing four types of fine-grained control conditions of semantics, structures, styles and numerical values based on the parameter set, and driving a condition control diffusion model to generate a high-fidelity service image; generating a specialized description text by the same language model in combination with the image visual features and the text outline; and finally, performing visual semantic alignment verification and business logic consistency auditing on the generated image-text data, and performing iterative correction based on a diagnosis result until a data pair meeting a preset quality standard is output. The data generated by adopting the method is standard, controllable and safe, a training data set can be efficiently constructed, and intelligent application research and development are accelerated.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of artificial intelligence and electric power informationization, and particularly relates to a multi-modal data generation technology for electric power system marketing and power distribution business, and more particularly to a method and system for generating high-realistic and high-correlation power marketing and distribution business scene data by fusing a diffusion model and a large language model. BACKGROUND

[0002] With the deepening of the construction of smart grids, the marketing and power distribution business integration of electric power systems has become the key to improving power supply service quality and operational efficiency. Artificial intelligence-based marketing and distribution integration applications, such as automatic defect recognition and intelligent inspection report generation, rely on large-scale, high-quality, multi-modal labeled data for model training. However, in actual business, there are core pain points such as the scarcity of high-quality labeled data, the lack of rare fault scene samples, the difficulty of sharing data involving security and privacy, and the weak correlation of existing text, image, and data multi-modal information.

[0003] In the prior art, there are methods of generating power equipment images using generative adversarial networks or generating text reports using template filling, but they all belong to single-modal generation or simple post-correlation, lacking understanding and constraints on cross-modal semantic consistency and deep business rules. In recent years, the development of diffusion models and large language models has provided a new path for data generation, but general models lack domain knowledge and are difficult to generate business data that conforms to electric power specifications and logically self-consistent.

[0004] Therefore, there is an urgent need for a method and system that can automatically generate marketing and distribution multi-modal data that maintains high visual fidelity, strictly conforms to electric power business logic and specifications, and has strong correlation between images and text. SUMMARY

[0005] The purpose of the present application is to overcome the defects of the prior art and provide a marketing and distribution multi-modal data automatic generation method based on the deep fusion of diffusion models and large language models, which can solve the problems of scarcity of high-quality training data and difficulty of cross-modal alignment, and generate data that is standardized, controllable, and safe, which can efficiently construct training data sets and accelerate the research and development of intelligent applications.

[0006] Another purpose of the present application is to provide a marketing and distribution multi-modal data automatic generation system based on the fusion of diffusion models and large language models.

[0007] In a first aspect, the present application provides a marketing and distribution multi-modal data automatic generation method based on the fusion of diffusion models and large language models, comprising the following steps:

[0008] S1, a large language model enhanced with marketing and distribution domain knowledge is used to analyze user natural language requirements and output a structured image control parameter set and a text generation outline;

[0009] S2, based on the image control parameter set, constructing multiple fine-grained control conditions including semantic conditions, structural conditions, style conditions and numerical constraint conditions, driving the condition control diffusion model to generate high-fidelity business images;

[0010] S3, combining the text generation content outline with the visual features of the high-fidelity scene image, inputting to the domain knowledge enhanced large language model to generate professional description text;

[0011] S4, performing visual semantic alignment verification and business logic consistency audit on the generated high-fidelity business image and description text, and iteratively correcting based on the diagnostic results until a data pair meeting the preset quality standard is output.

[0012] The above-mentioned method for automatically generating multi-modal data of marketing and distribution based on the fusion of diffusion model and large language model, in step S1, the image control parameter set at least includes subject object parameters, defect or state parameters, scene context parameters and technical constraint parameters; the text generation content outline at least includes target text type, essential key element description and format specification to be followed.

[0013] The above-mentioned method for automatically generating multi-modal data of marketing and distribution based on the fusion of diffusion model and large language model, in step S2, the construction of the structural condition includes: according to the subject object and technical constraint in the image control parameter set, retrieving the corresponding line drawing from the pre-built power standard component vector graph library, and performing affine transformation and defect area labeling according to the scene context parameter, generating edge graph and semantic segmentation sketch as structural constraint for the diffusion model.

[0014] The above-mentioned method for automatically generating multi-modal data of marketing and distribution based on the fusion of diffusion model and large language model, in step S2, the condition control diffusion model adopts a model integrated with ControlNet architecture, and the denoising network of the model fuses the encoding features of the semantic condition, the structural condition and the style condition at each step.

[0015] The above-mentioned method for automatically generating multi-modal data of marketing and distribution based on the fusion of diffusion model and large language model, in step S2, when training or fine-tuning the condition control diffusion model, a multi-objective weighted loss function including basic reconstruction loss, perception loss, CLIP semantic alignment loss, structural fidelity loss and style loss is used for optimization.

[0016] The method for automatically generating the marketing and distribution multi-modal data based on the diffusion model and the large language model fusion, in step S4, the visual semantic alignment verification adopts fine-grained visual semantic alignment verification, specifically comprising: using a visual question answering model or an open set target detection model, the generated description text is decomposed, and the related area is located in the generated image to verify the authenticity of the assertion, and the alignment accuracy is calculated.

[0017] The method for automatically generating the marketing and distribution multi-modal data based on the diffusion model and the large language model fusion, in step S4, the business logic consistency audit comprises: inputting the high-fidelity business image generated in step S2 and the description text generated in step S3 into a knowledge graph reasoning engine of built-in power field business rules for compliance checking, and / or calling a large language model for expert-level logic auditing.

[0018] The method for automatically generating the marketing and distribution multi-modal data based on the diffusion model and the large language model fusion, in step S4, the iterative correction adopts convergence control based on quality score, and the comprehensive quality score function Q is:

[0019] Q = w1 x visual semantic alignment score + w2 x logic consistency score + w3 x format specification score;

[0020] When Q >= tau or the maximum iteration number K is reached, the iteration is stopped, wherein w1, w2, and w3 are weights, and tau is a quality threshold.

[0021] In a second aspect, the present application provides a system for automatically generating marketing and distribution multi-modal data based on diffusion model and large language model fusion, which is used to realize the above-mentioned method for automatically generating marketing and distribution multi-modal data, comprising:

[0022] An instruction generation module, which is built-in with a large language model fine-tuned by power marketing and distribution field knowledge, is used to receive natural language scene requirements and output structured image control parameter set and text generation content outline;

[0023] An image generation module, which is built-in with a conditional control diffusion model and a control condition constructor, is used to construct multiple control conditions according to the image control parameter set and generate high-fidelity scene images;

[0024] A text generation module is used to fuse the visual features of the high-fidelity scene images and the text generation content outline, and call the large language model to generate professional description text;

[0025] An alignment verification and iterative control module is used to perform fine-grained alignment verification and logic auditing on the generated images and text, diagnose inconsistent types, and control the iterative correction process of the system.

[0026] In a second aspect, the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the automated generation method for multimodal data of operation and maintenance described in the first aspect.

[0027] The technical solution of the automated generation method and system for multimodal data of camp operations based on the fusion of diffusion model and large language model of the present invention has the following significant advantages compared with the prior art:

[0028] (1) High-quality and logically consistent data generation: Through the collaboration of domain knowledge-enhanced LLM and multi-condition controlled diffusion model, the data is highly consistent in semantics and business rules.

[0029] (2) Good industry standard compliance: The generated equipment images are accurate in shape, the text report format is standardized, and it conforms to the technical standards of the power industry;

[0030] (3) Efficiently construct scarce scenario data: It can generate training data for various rare faults and extreme working conditions on demand, solving the long-tail problem of AI model training;

[0031] (4) Safe and controllable: The generation process is based on simulation synthesis and does not rely on sensitive real data, which meets the requirements of data security and privacy protection.

[0032] (5) High degree of automation: forming a closed-loop process of “generation-verification-correction”, which greatly reduces the cost of manual intervention. Attached Figure Description

[0033] Figure 1 This is a schematic diagram of the overall process of an automated method for generating multimodal data for logistics operations according to the present invention;

[0034] Figure 2 This is a schematic diagram illustrating the principle of fine-grained conditional image generation in this invention;

[0035] Figure 3 This is a schematic diagram of the cross-modal alignment and iterative optimization process in this invention;

[0036] Figure 4 This is a framework diagram of an automated multimodal data generation system for operations and distribution according to the present invention. Detailed Implementation

[0037] To enable those skilled in the art to better understand the technical solution of the present invention, its specific embodiments are described in detail below with reference to the accompanying drawings:

[0038] Please see Figures 1 to 3In an embodiment of the present application, a method for automatically generating power supply and distribution multi-modal data based on diffusion model and large language model fusion includes the following steps:

[0039] S1, power supply and distribution knowledge enhancement and structured instruction decomposition step: using a large language model fine-tuned by power supply and distribution field knowledge, receiving user input natural language scene description, and based on internalized domain knowledge, decomposing and outputting a structured cross-modal generation instruction set. The instruction set includes an image control parameter set for precise control of image generation, and a text generation content outline for guiding text generation.

[0040] S2, fine-grained conditional image generation step: based on the image control parameter set output by S1, multiple fine-grained control conditions are constructed and input into the conditional control diffusion model to generate high-fidelity scene images. Specifically, it includes:

[0041] S2.1, four types of control conditions including semantic conditions, structural conditions, style conditions and numerical constraint conditions are constructed. The structural condition is generated by retrieving the line drawing from the pre-built power standard component library and transforming and defect labeling according to the parameters;

[0042] S2.2, a conditional diffusion model integrated with ControlNet is used to fuse the encoding features of all control conditions during the denoising process for image synthesis;

[0043] S2.3, in the model training stage, a multi-objective loss function including basic reconstruction loss, perception loss, CLIP semantic alignment loss, structure fidelity loss and style loss is used for optimization.

[0044] S3, field specialized text generation step: the text generation content outline of S1 and the visual features of the scene image generated by S2 are fused and input into the language model enhanced in the same field to generate professional description text with standard format, accurate terminology and consistent with image content semantics.

[0045] S4, cross-modal alignment and iterative optimization step: the generated high-fidelity scene image and professional description text are checked for bidirectional consistency: including fine-grained assertion verification using visual models, and business logic audit based on rule engine and LLM. If the quality threshold is not met, diagnose the problem type, adjust the corresponding control parameters or generation outline, trigger iterative regeneration until the output meets the requirements of the data pair.

[0046] Please refer to Figure 4 The power supply and distribution multi-modal data automatic generation system for implementing the above method of the present application includes instruction generation module 1, image generation module 2, text generation module 3 and alignment verification and iterative control module 4.

[0047] The instruction generation module 1 internally has a large language model fine-tuned by power supply and distribution field knowledge, is used for receiving a natural language scene demand, and outputs a structured image control parameter set and a text generation content outline; the image generation module 2 internally has a conditional control diffusion model and a control condition constructor, is used for constructing multiple control conditions according to the image control parameter set, and generates a high-fidelity scene image; the text generation module 3 is used for fusing the text generation content outline and the visual features of the high-fidelity scene image, calling the large language model to generate professional description text; and the alignment verification and iterative control module 4 is used for performing fine-grained alignment verification and logical auditing on the generated image and text, diagnosing inconsistent types, and controlling the iterative correction process of the system.

[0048] The embodiment takes generating "10kV pole-mounted circuit breaker mechanism box internal condensation defect" inspection data as an example.

[0049] Embodiment 1: Image synthesis and optimization based on conditional diffusion model

[0050] The embodiment details the image synthesis technical solution based on the conditional diffusion model, and the core is to construct accurate multi-level control conditions and optimize the generation quality through a loss function.

[0051] 1. System initialization and image generation process definition:

[0052] Stable Diffusion v2.1 is used as the basic architecture, and a ControlNet extension module specially designed for power equipment is integrated. The image generation process is defined as a reverse denoising process: given an initial noise image The target image x0 is obtained by T-step progressive denoising. At each time step t, the prediction function of the denoising network ∈ θ is defined as:

[0053] ∈ θ (x t ,t,C)=U-Net θ (x t ,t,Φ(C))

[0054] Where C={C sem , C struct , C style , C num} represents a four-level control condition set specially designed for power equipment, and Φ(·) is a multi-modal condition encoder.

[0055] 2. Mathematical expression and generation process of multi-level control conditions

[0056] In this embodiment, based on the user input "10kV pole-mounted circuit breaker mechanism box internal condensation defect" requirement, after large language model analysis, the structured parameter set P is obtained. Based on this parameter set, the following four types of control conditions are constructed:

[0057] (1) Semantic condition C sem Generation: Convert the structured parameter set P to enhanced prompt words by LLM in the power field:

[0058] C sem = LLM expand (T(P))

[0059] Where T(·) is a pre-set power equipment description template filling function, LLM expand is a specially trained prompt word expansion model used to enrich visual detail vocabulary related to power inspection. In this example, the generated prompt words are: "high-definition professional inspection photos, internal view of ZW32-12 type pole-mounted circuit breaker mechanism box, Meiyu season environment, transparent water droplets with a diameter of 0.5-2mm distributed on the inner wall and corners of the box, density about 4 / cm 2 , humidity meter shows 85% RH, box lighting, visible operating mechanism and secondary wiring."

[0060] (2) Generation of structure condition C struct : This step ensures the accuracy of the generated device geometry. Retrieve the matching ZW32-12 type circuit breaker line drawing model from the power standard component library :

[0061]

[0062] Apply affine transformation A according to the specified inspection viewing angle (30 degrees of depression) in the parameter set:

[0063] M trans = A(M base ,P pose ,P view )

[0064] On the transformed model, according to the defect type and location parameters P defect_{type} (condensation), P defect\_loc (sidewall and corner), generate defect area mask through algorithm:

[0065] M defect = Annotate(M trans ,P defect\_type ,P defect\_loc )

[0066] Finally, through Canny edge detection and semantic segmentation algorithm, the accurate edge map C edge= Canny(M defect ) and semantic segmentation sketch C seg = Segment(M defect ) as inputs of ControlNet.

[0067] (3) Style condition C style generation: To match the style requirement of "professional inspection photo", style description is converted to pre-trained embedding vector through lookup table mapping:

[0068] C style = EmbeddingLookup(P style = "professional inspection", ε)

[0069] Where ε is the LoRA weight library obtained by fine-tuning using a large number of real power inspection photos, which can control the contrast, noise pattern and hue of the generated image.

[0070] (4) Numerical constraint condition C num setting: Extract specific numerical constraints from parameters, such as the size range of condensation droplets [d min = 0.5, d max = 2.0] mm, target color distribution H target (translucent to semi-transparent), and brightness range of the overall image.

[0071] C num = {size_range: [0.5, 2.0], color_hist: H target , brightness: [80, 120]}

[0072] 3. Condition fusion and image generation mechanism: In the diffusion model integrated with ControlNet, control conditions directly affect the generation process through feature fusion. Let the intermediate feature map of the base UNet at the l-th layer be The corresponding feature map output by the ControlNet branch after processing the structure condition is Then the fused feature is:

[0073]

[0074] Where α (l) is a learnable scaling parameter optimized for power equipment image generation, and Conv is a convolutional adaptation layer. At the same time, semantic condition C sem and style condition C style are injected into UNet through cross-attention mechanism.

[0075] 4. Multi-objective loss function design and optimization

[0076] In the model training stage, a multi-objective loss function specially designed to improve the image generation quality of power equipment is adopted:

[0077]

[0078] wherein, is the total loss, λ1, λ2, λ3, λ4, λ5 are the weight coefficients of each loss;

[0079] Each loss function is defined as follows:

[0080] (1) Basic reconstruction loss The mean square error of noise prediction is adopted:

[0081]

[0082] wherein, is the expected value (usually calculated as the average value on the batch data), ∈ is the real Gaussian noise added to the original image, ∈ θ is the noise predicted by the denoising model (such as U-Net); x t is the noisy image at time step t; t is the time step in the diffusion process; C is the conditional information (such as text prompt words), is the square of L2 norm (mean square error).

[0083] (2) Perceptual loss A pre-trained VGG network φ is used to ensure that the generated image is realistic in texture:

[0084]

[0085] wherein, is the expected value (usually calculated as the average value on the batch data), φ l is the feature extraction function of the l-th layer of the pre-trained VGG network (or other feature extraction network), is the image generated by the model, x0 is the real (target) image, ||·||1 is the L1 norm (absolute error), is the sum of the feature differences of the selected multiple network layers (l).

[0086] (3) CLIP semantic alignment loss Ensure that the image content is consistent with the prompt words:

[0087]

[0088] wherein, E I is the image encoder of the CLIP model, E T is the text encoder of the CLIP model, C semFor semantic condition (prompt word), used to describe the expected image content, For the CLIP feature vector of the generated image, E T (C sem For the CLIP text feature vector of the semantic condition, the fractional part is the cosine similarity, which is used to measure the matching degree of the image and the text in the semantic space.

[0089] (4) Structure fidelity loss Adopt multi-scale structural similarity and Dice coefficient to ensure the accuracy of the device shape:

[0090]

[0091] Where s is the scale index, representing different resolution levels in multi-scale calculation; SSIM s is the structural similarity index calculated at scale s, C edge is the edge map provided as a condition, guiding the shape of the generated image, is the mask map obtained by segmenting the generated image, C seg is the target segmentation mask map provided as a condition, Dice(·) is the Dice coefficient, which measures the overlap between two segmentation maps.

[0092] (5) Style loss Match a specific style through Gram matrix:

[0093]

[0094] Where, is the feature map of the generated image at the lth layer of VGG, x ref is the reference style image (e.g., a real photo style of a specific device type), G(·) is the Gram matrix of the given feature map, which is used to capture texture and style information, is the square of the Frobenius norm, which is used to measure the difference between two Gram matrices.

[0095] (6) Numerical constraint loss is realized through additional regularization loss, such as constraint on water droplet size

[0096]

[0097] Where Area(·) is a function that calculates the area of a region; is the condensation area (defect area) identified by segmentation in the generated image; d max is the maximum area threshold of the allowed defect area; d minmin_area_threshold is the minimum area threshold for allowed defect regions; max(0, ·) is the ReLU function that only produces a penalty when the area exceeds the threshold range.

[0098] Through the above fine control and optimization, the system finally generates a high-fidelity circuit breaker mechanism box condensation defect image with a resolution of 2048x1536, which has accurate device structure and clear defect features, and meets the professional inspection standards.

[0099] Embodiment 2: Cross-modal alignment and iterative optimization

[0100] This embodiment details how to perform strict cross-modal alignment verification and automatic iterative optimization after generating images and texts to ensure the business logic consistency of the data pair.

[0101] 1. Fine-grained visual-text alignment verification:

[0102] First, the generated defect record single text description T is semantically parsed to decompose it into a set of atomic assertions For example, a1 = "water droplets on the side wall of the box", a2 = "water droplets about 1mm in diameter". For each assertion a i , use the open set object detection model D open to locate and verify in the generated image :

[0103] Score(a i ) = IoU(B text , B visual ) · S conf

[0104] Where Score(a i ) is the score of the ith assertion in the image-text alignment; B text is the spatial constraint implied by the text assertion; B visual is the corresponding region bounding box detected by the model in the image; IoU(B text , B visual ) is the intersection over union, which measures the degree of overlap between two bounding boxes; S conf is the detection confidence.

[0105] The overall image-text alignment score S align is calculated as:

[0106]

[0107] Where n is the total number of assertions; Score(a i ) is the score of the ith assertion in the image-text alignment; τ align is the alignment threshold, which is set to 0.7. is an indicator function that takes value 1 if the condition holds, and 0 otherwise.

[0108] Meanwhile, consistency verification is performed on attributes such as color and quantity:

[0109]

[0110] where m is the total number of attributes, is the jth attribute in the textual description; is the jth attribute in the visual system; Sim(·) is a similarity function that measures the consistency between textual and visual attributes; S attr is the attribute alignment score, representing the average value of attribute consistency.

[0111] 2. Deep audit of business logic consistency: encode power domain knowledge as a set of rules For example, rule r1: IF defect type is "condensation" AND defect level is "general" THEN the handling recommendation should not contain "immediate power outage". The verification process is performed by a logical reasoning engine:

[0112]

[0113] where V rule is the rule verification result, True if and only if all rules pass, k is the total number of business rules, r i is the ith business rule, Eval is the evaluation function, is the generated image, T is the generated textual description.

[0114] In addition, construct expert audit hints P audit Input LLM for deep audit:

[0115]

[0116] where Concat denotes the concatenation function, is the generated image encoding representation (such as through a CLP encoder or image description model), T is the generated textual description.

[0117] Extract the inconsistent item set from the audit report R audit of the LLM (Large Language Model)

[0118] 3. Automatic execution of iterative correction strategy: based on the verification result, the system automatically diagnoses the type of inconsistency:

[0119]

[0120] where F diagnoseFor diagnostic function, based on each score and inconsistency item to determine the source of the problem, S align For image-text alignment score (from previous alignment verification), S attr For attribute alignment score, V rule For rule verification result, For inconsistency item set. Type is the type of inconsistency diagnosed, including "Image_Detail", "Text_Inaccuracy" and "Logic_Conflict".

[0121] If the diagnosis is "Image_Detail" (e.g. water droplets are not obvious in the image), adjust the structural condition parameters: And regenerate the image;

[0122] Where, is the defect position parameter value at the kth iteration; Δ loc is the position adjustment step size; is the alignment loss The gradient of the position parameter indicates the adjustment direction; is the updated defect position parameter value.

[0123] If the diagnosis is "Text_Inaccuracy" (e.g. the defect level determination in the text is ambiguous), enhance the visual reference weight during text generation: And regenerate the text;

[0124] Where, is the text generation prompt word at the kth iteration, represents the splicing or fusion operation, [VISION_REF] is the visual reference enhancement mark, β is the weight coefficient of the visual reference, is the updated text generation prompt word,

[0125] If the diagnosis is "Logic_Conflict", trigger the demand to be re-analyzed:

[0126] Where, P (k) is the demand analysis result at the kth iteration, LLM clarify is a large language model used to clarify the demand, Inconsistency item set, P (k+1) is the updated clear demand analysis result.

[0127] The system calculates the overall quality score to determine convergence, and the overall quality score formula is as follows:

[0128]

[0129] Wherein, Q is the total quality score; S align is the overall text-image alignment score, S attr is the attribute alignment score, V rule is the rule verification result, w1, w2, w3, w4 are weight coefficients of each score; is the number of inconsistent items in the current iteration; is the maximum allowed number of inconsistent items; is the inconsistency penalty term, the fewer the inconsistent items, the higher the score of this term.

[0130] The iteration termination condition is: Q≥τ quality = 0.9 or the maximum number of iterations K max = 5. That is, when the total quality score reaches the quality threshold τ quality = 0.9 or the maximum number of iterations, the iteration is terminated.

[0131] In the condensation defect generation task of the present embodiment, after the above process, the system outputs a high-quality data pair after an average of 2.8 iterations. The final evaluation shows that the cross-modal alignment accuracy reaches 94.7%, the business rule compliance rate reaches 97.2%, and the artificial evaluation quality score is 4.6 / 5 points, which is significantly better than the traditional generation method.

[0132] Embodiments 1 and 2 fully demonstrate the feasibility and superiority of the technical solution of the present application. By deeply integrating domain knowledge into the generation process and establishing a strict closed-loop verification and correction mechanism, the present application can automatically and large-scale generate multi-modal data that meets the high standards of the power industry, effectively solving the data bottleneck problem in the development of integrated power supply and distribution intelligent applications, and providing strong technical support for the digital transformation of the power industry.

[0133] In summary, the automatic generation method and system of multi-modal data of power supply and distribution of the present application solves the problem of the scarcity of high-quality, strongly logically correlated image and text data in the integrated power supply and distribution AI application, and the generated data is standardized, controllable and safe, which can efficiently construct a training data set and accelerate the research and development of intelligent applications.

[0134] Those skilled in the art in this technical field should realize that the above embodiments are only used to illustrate the present application, and are not used as a limitation on the present application, as long as the changes and modifications of the above described embodiments are within the scope of the spirit of the present application. The above described embodiments will fall within the scope of the claims of the present application.

Claims

1. A method for automatically generating multi-modal data of business and distribution based on diffusion model and large language model fusion, characterized in that, The method comprises the following steps: S1, using a large language model enhanced by power marketing field knowledge to parse user natural language requirements, outputting a structured image control parameter set and a text generation outline; S2, based on the image control parameter set, constructing multiple fine-grained control conditions including semantic conditions, structure conditions, style conditions and numerical constraint conditions to drive the condition control diffusion model to generate high-fidelity business images; S3, combining the text generation content outline with the visual features of the high-fidelity scene image, inputting into the large language model enhanced by the field knowledge to generate professional description text; S4, performing visual semantic alignment verification and business logic consistency audit on the generated high-fidelity business image and description text, and iteratively correcting based on the diagnostic results until a data pair meeting the preset quality standard is output.

2. The method of claim 1, wherein the method is based on a diffusion model and a large language model. In step S1, the image control parameter set at least includes subject object parameters, defect or state parameters, scene context parameters and technical constraint parameters; and the text generation content outline at least includes target text type, essential key element description and format specification to be followed.

3. The method of claim 1, wherein the method is based on a diffusion model and a large language model. In step S2, the construction of the structure condition includes: according to the subject object and technical constraint in the image control parameter set, retrieving the corresponding line drawing from the pre-built power standard component vector graph library, and performing affine transformation and defect area labeling according to the scene context parameters to generate edge graph and semantic segmentation sketch as the structure constraint for the diffusion model.

4. The method of claim 1, wherein the method is based on a diffusion model and a large language model. In step S2, the condition control diffusion model adopts a model integrated with ControlNet architecture, and the denoising network of the model fuses the encoding features of the semantic condition, structure condition and style condition at each input step.

5. The method of claim 1, wherein the method is based on a diffusion model and a large language model fusion. In step S2, when training or fine-tuning the condition control diffusion model, a multi-objective weighted loss function including basic reconstruction loss, perception loss, CLIP semantic alignment loss, structure fidelity loss and style loss is used for optimization.

6. The method of claim 1, wherein the method is based on a diffusion model and a large language model. In step S4, the visual semantic alignment verification adopts fine-grained visual semantic alignment verification, specifically including: using a visual question answering model or an open set target detection model to perform assertion decomposition on the generated description text, and positioning the related area in the generated image to verify the truth of the assertion, and calculating the alignment accuracy.

7. The method of claim 1, wherein the method is based on a diffusion model and a large language model fusion. In step S4, the business logic consistency audit includes: inputting the high-fidelity business image generated in step S2 and the description text generated in step S3 into a knowledge graph reasoning engine of built-in power field business rules for compliance checking, and / or calling a large language model for expert-level logic auditing.

8. The method of claim 1, wherein the method is based on a diffusion model and a large language model fusion. In step S4, the iterative correction adopts convergence control based on quality score, and the comprehensive quality score function Q is: Q = w1 x visual semantic alignment score + w2 x logical consistency score + w3 x format specification score; When Q ≥ τ or the maximum iteration number K is reached, the iteration is stopped, wherein w1, w2 and w3 are weights, and τ is a quality threshold.

9. A system for automatically generating campaign multi-modal data, for implementing the method for automatically generating campaign multi-modal data according to any one of claims 1 to 8, characterized in that, The method comprises: An instruction generation module built-in with a large language model fine-tuned by power marketing field knowledge, used to receive natural language scene requirements and output a structured image control parameter set and a text generation content outline; An image generation module, which is internally provided with a conditional control diffusion model and a control condition constructor, is configured to construct multiple control conditions according to the image control parameter set and generate a high-fidelity scene image; A text generation module is configured to fuse the text generation content outline and the visual features of the high-fidelity scene image, and call the large language model to generate professional description text. An alignment verification and iteration control module is configured to perform fine-grained alignment verification and logical auditing on the generated image and text, diagnose the inconsistency type, and control the iteration correction process of the system.

10. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the steps of the operation and maintenance multi-modal data automatic generation method according to any one of claims 1 to 8.