Training method of generative model for generating visual ip picture material and visual ip picture generation method

By combining the LoRa fine-tuning model and the style transfer fine-tuning model, the problems of automated learning and deep semantic association in the generation of visual IP materials are solved, and stable and efficient multi-scene and multi-style generation effects are achieved.

CN120931765BActive Publication Date: 2026-04-21GUANGZHOU TAIDONG TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU TAIDONG TECH CO LTD
Filing Date
2025-07-03
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing visual IP material generation solutions lack automated learning mechanisms, making it difficult to accurately capture deep semantic relationships. The generated results are unstable, unable to meet the needs of mass production in multiple scenarios and styles, and have low update efficiency.

Method used

We employ a combination of LoRa fine-tuning and style transfer fine-tuning models, using classified and labeled sample materials for forward propagation to construct an automated feature learning mechanism, thereby achieving deep semantic association capture and multi-style output.

Benefits of technology

It has achieved automated feature learning for visual IP materials, improved the stability and adaptability of the generated results, met the batch production needs of multiple scenarios and styles, and significantly improved development and iteration efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931765B_ABST
    Figure CN120931765B_ABST
Patent Text Reader

Abstract

This invention provides a training method for a generative model that generates visual IP image materials, a visual IP image generation method, and an electronic device. The training system includes: a material classification and labeling unit, used to acquire IP image samples carrying company logos, and classify the IP image samples according to classification and labeling tags to obtain classified and labeled sample materials; and a fine-tuning model training unit, used to enable a pre-trained model to perform forward propagation processing on the classified and labeled sample materials until the pre-trained model has model parameters that can learn the company visual IP features on the IP image samples, and uses the pre-trained model at this point as a generative model. The pre-trained model includes a LoRa fine-tuning model and a style transfer fine-tuning model. Automated feature learning is achieved by combining the LoRa fine-tuning model and the style transfer fine-tuning model, training a generative model for generating visual IP image materials.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more particularly to a training method for a generative model for generating visual IP image materials, a method for generating visual IP images, and an electronic device. Background Technology

[0002] In the current construction of brand visual identity systems (VIS), enterprises urgently need to efficiently generate visual IP image materials that conform to the brand's tone to meet the application needs of multiple scenarios such as advertising design, product packaging, and digital media. With the development of generative artificial intelligence technology, image generation solutions based on pre-trained models are gradually becoming an industry trend. However, how to enable the generative model to accurately learn the unique visual IP characteristics of an enterprise (such as brand logo, exclusive colors, design style, etc.) remains a technical challenge that needs to be solved.

[0003] Existing visual IP material generation solutions are primarily based on traditional image processing algorithms. For example, edge detection algorithms are used to extract logo outlines, color space conversion algorithms are used to match brand standard colors, and template matching technology is used to reuse mascot patterns. While some solutions incorporate machine learning methods, they only reach the shallow feature extraction stage. They process basic features such as texture and color of IP materials using manually designed filters (such as Gaussian filtering and Laplacian operators), without building an automated learning mechanism for core visual IP features (such as semantic style and element relationships). Furthermore, traditional algorithms lack the ability to process visual IP features hierarchically, making it difficult to maintain the integrity of low-level elements such as logos and colors while achieving consistent transfer of high-level semantic features such as brand style (e.g., technological or retro).

[0004] The aforementioned solutions have revealed significant technical shortcomings in practical applications: Firstly, traditional algorithms rely on manually preset feature extraction rules. When corporate IP materials are updated (such as logo tweaks or style iterations), algorithm parameters need to be redesigned, leading to low development efficiency and difficulty in adapting to dynamic needs. Secondly, shallow feature processing cannot capture the deep semantic relationships of visual IPs (such as the matching logic between logo colors and packaging backgrounds). The generated materials often exhibit problems such as awkward element splicing and a disconnect between style and brand tone, failing to meet the stringent requirements of enterprises for the systematic nature and recognizability of visual IPs. Furthermore, traditional solutions lack automated model optimization mechanisms, resulting in insufficient stability and controllability of the generated results, making it impossible to cope with the demand for batch production of materials across multiple scenarios and styles. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a training method for a generative model for generating visual IP image materials, a visual IP image generation method, and an electronic device, so as to at least partially solve the above-mentioned problems.

[0006] A training system for a generative model that generates visual IP image assets, comprising:

[0007] The material classification and labeling unit is used to obtain IP image samples carrying the company logo, and classify the IP image samples according to the classification and labeling marks to obtain classified and labeled sample materials;

[0008] The fine-tuning model training unit is used to enable the pre-trained model to perform forward propagation processing on the classified and labeled sample materials until the pre-trained model has model parameters that can learn the company's visual IP features on the IP image samples. The pre-trained model at this time is used as a generative model. The pre-trained model includes the LoRa fine-tuning model and the style transfer fine-tuning model.

[0009] Optionally, when the material classification and labeling unit obtains an IP image sample carrying the company logo, it parses the IP image sample to extract its visual element materials and generates IP visual core features accordingly. The visual element materials include: color scheme features, image style features, mascot features, and packaging design features.

[0010] Optionally, the material classification and labeling unit classifies IP image samples according to the classification and labeling marks of the IP image samples. When obtaining classified and labeled sample materials, the IP image samples are classified and labeled according to the core visual features of the IP images to form IP image samples carrying classification and labeling tags. The IP image samples are then classified based on the classification and labeling tags to obtain classified and labeled sample materials.

[0011] Optionally, the fine-tuning model training unit is used to enable the pre-trained model to perform forward propagation processing on the classified labeled sample materials until the pre-trained model has model parameters that can learn the visual IP features of the company on the IP image samples. The pre-trained model at this time is used as a generative model. When the pre-trained model includes a LoRa fine-tuning model and a style transfer fine-tuning model, the LoRa fine-tuning model performs forward propagation processing on the visual elements of the classified labeled sample materials, and the style transfer fine-tuning model performs forward propagation processing on the semantic style of the classified labeled sample materials until the LoRa fine-tuning model has model parameters that can learn the visual element features on the IP image samples, and the style transfer fine-tuning model has model parameters that can learn the semantic style features on the IP image samples.

[0012] Optionally, when the LoRa fine-tuning model performs forward propagation processing on the visual elements of the classified and labeled sample materials, it locates the visual element regions of the classified and labeled sample materials, performs masking processing on the visual element regions, generates a visual element mask matrix, and performs progressive exposure processing on the visual element mask matrix, so that the LoRa fine-tuning model has model parameters that learn the visual element features on the IP image samples.

[0013] Optionally, the LoRa fine-tuning model locates the visual element regions of the labeled sample materials, performs masking on these regions to generate a visual element mask matrix, and applies progressive exposure processing to the visual element mask matrix so that the LoRa fine-tuning model has model parameters that learn the visual element features on the IP image samples. This includes the following steps:

[0014] Visual element localization processing is performed on the categorized and labeled sample materials to generate a set of visual element coordinates;

[0015] The set of visual element coordinates is processed to generate a mask matrix, thus generating an initial visual element mask matrix:

[0016] The initial visual element mask matrix is ​​initialized with transparency gradient to generate a gradient mask matrix;

[0017] A progressive exposure process is applied to the gradient mask matrix to generate a dynamic exposure mask matrix;

[0018] Feature occlusion training processing is performed on the dynamic exposure mask matrix and the classified labeled sample materials to generate a mask feature training set;

[0019] Structural loss is calculated on the mask feature training set to generate element feature loss values;

[0020] The LoRa fine-tuning model parameters are backpropagated based on the element feature loss values, enabling the LoRa fine-tuning model to learn the model parameters of visual element features on IP image samples.

[0021] Optionally, when the style transfer fine-tuning model performs forward propagation processing on the semantic style of the classified labeled sample materials, it extracts semantic features from the classified labeled sample materials to obtain a semantic style feature tensor, applies style attention weighting to the semantic style feature tensor to obtain an attention-weighted semantic style feature tensor, and performs style mapping on the attention-weighted semantic style feature tensor so that the style transfer fine-tuning model can learn the model parameters of the semantic style features on the IP image samples.

[0022] Optionally, the style transfer fine-tuning model extracts semantic features from the classified labeled sample materials to obtain a semantic style feature tensor. It then applies style attention weights to the semantic style feature tensor to obtain an attention-weighted semantic style feature tensor. Finally, it performs style mapping on the attention-weighted semantic style feature tensor to enable the style transfer fine-tuning model to learn the model parameters of the semantic style features on the IP image samples. The following steps are then performed:

[0023] Multimodal feature extraction is performed on the classified and labeled sample materials to generate an initial semantic style feature tensor;

[0024] The initial semantic style feature tensor is subjected to channel attention weighting to generate a channel-enhanced feature tensor;

[0025] Spatial attention weighting is applied to the channel enhancement feature tensor to generate an attention-focused feature tensor;

[0026] Style decoupling is performed on the attention-focused feature tensor to generate a content-style separation tensor;

[0027] Style fusion processing is performed on the content-style separation tensor to generate stylized feature representations;

[0028] Generative adversarial training is performed on the stylized feature representations to generate style consistency loss;

[0029] The model parameters of the style transfer fine-tuning model are updated using gradient processing based on style consistency loss, so that the style transfer fine-tuning model can learn the model parameters of semantic style features on IP image samples.

[0030] A method for generating visual IP images, which is based on a generative model of any one of the embodiments of this application.

[0031] An electronic device includes a memory and a processor, the memory storing a computer-executable program, and the processor running the computer-executable program to perform the following steps:

[0032] Obtain IP image samples carrying the company logo, and classify the IP image samples according to the classification labels to obtain classified label sample materials;

[0033] This allows the pre-trained model to perform forward propagation on the classified and labeled sample materials until the pre-trained model has model parameters that can learn the company's visual IP features on the IP image samples. The pre-trained model at this point is used as a generative model. The pre-trained model includes a LoRa fine-tuning model and a style transfer fine-tuning model.

[0034] The training method for the generative model of generating visual IP image materials in the embodiments of the present invention has the following technical advantages:

[0035] ① An automated feature learning mechanism is achieved through a combination of "classification and labeling sample materials" and "cascaded LoRa fine-tuning model". The LoRa fine-tuning model can perform parameterized adaptation for the underlying visual elements of the enterprise IP (such as logo outline and standard color values). When the enterprise IP is updated, only the classification and labeling rules need to be adjusted and the LoRa module needs to be retrained, without the need for overall algorithm reconstruction, which significantly reduces the cost of manual intervention. This design breaks through the rigid limitation of "logo fine-tuning requires resetting filter parameters" in traditional solutions, enabling the model to have dynamic adaptation capabilities.

[0036] ② A "style transfer fine-tuning model" and a "layered processing architecture" are used to capture deep semantic associations. The style transfer module focuses on learning brand tone (such as cool color schemes for a tech feel and grainy textures for a retro style), forming a layered decoupling with the underlying elements processed by the Lora module. While preserving the integrity of the logo / color, this architecture achieves end-to-end learning of high-level semantic features (such as element matching logic) through the style transfer network, solving the core pain point of traditional solutions: "awkward element splicing and disjointed style".

[0037] ③ An automated optimization mechanism is constructed through a hybrid training paradigm of "pre-trained model + fine-tuning". The pre-trained model provides general image generation capabilities, while the fine-tuning module focuses on learning enterprise IP features. The two are cascaded to form a pipeline of "feature decoupling - style transfer". This design enables the model to maintain the stability of the generated results (through the basic capabilities of pre-training) and achieve multi-style output (such as generating advertising posters and product packaging materials at the same time) through parameterized control of the style transfer module, meeting the batch production needs of "multi-scene and multi-style".

[0038] ④ Adaptive training driven by forward propagation is used instead of traditional manual parameter tuning. By automatically feeding in labeled samples and backpropagating the loss function, the model can autonomously optimize LoRa weights and style transfer parameters without the need for manual design of filters or color matching rules. This process transforms the discrete development of "manually designing feature extraction rules" in traditional solutions into continuous model iteration, significantly improving development efficiency.

[0039] ⑤ By decoupling explicit features through "classification and labeling units," the brand visual IP is broken down into quantifiable and learnable feature dimensions (such as logo position, primary and secondary color ratio, and style keywords). This structured representation method avoids the hidden danger of "strong coupling between feature extraction rules and IP elements" in traditional solutions. Even if the corporate IP is iterated, feature transfer can be achieved by adjusting the classification and labeling system, ensuring the systematic nature and recognizability of the visual IP. Attached Figure Description

[0040] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in the embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings.

[0041] Figure 1 This is a schematic diagram of the structure of a training system for a generative model that generates visual IP image materials, according to an embodiment of this application.

[0042] Figure 2 This invention provides a schematic diagram of the structure of an electronic device. Detailed Implementation

[0043] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art should fall within the protection scope of the present invention.

[0044] It should be understood that the terms "first," "second," and "third," etc., in the claims, specification, and drawings of this disclosure are used to distinguish different objects, not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this disclosure indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or sets thereof.

[0045] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of this disclosure. As used in this disclosure and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this disclosure and claims refers to any and all combinations of one or more of the associated listed items, and includes such combinations.

[0046] Figure 1 This is a schematic diagram of the structure of a training system for a generative model that generates visual IP image materials, according to an embodiment of this application. Figure 1 As shown, it includes:

[0047] The material classification and labeling unit is used to obtain IP image samples carrying the company logo, and classify the IP image samples according to the classification and labeling marks to obtain classified and labeled sample materials;

[0048] The fine-tuning model training unit is used to enable the pre-trained model to perform forward propagation processing on the classified and labeled sample materials until the pre-trained model has model parameters that can learn the company's visual IP features on the IP image samples. The pre-trained model at this time is used as a generative model. The pre-trained model includes the LoRa fine-tuning model and the style transfer fine-tuning model.

[0049] Optionally, when the material classification and labeling unit obtains an IP image sample carrying the company logo, it parses the IP image sample to extract its visual element materials and generates IP visual core features accordingly. The visual element materials include: color scheme features, image style features, mascot features, and packaging design features.

[0050] Preferably, when the material classification and labeling unit acquires an IP image sample carrying the company logo, and parses the IP image sample to extract its visual element materials and generate the IP visual core features accordingly, the following steps are performed:

[0051] For IP image samples carrying company logos, instance segmentation algorithms (such as Mask R-CNN) are used to perform pixel-level localization of visual elements such as logos, mascots, and core patterns in the images, generate element bounding box coordinates and masks, and extract structural frame regions through Hough transform and contour detection.

[0052] Based on the element bounding box coordinates, mask, and structural frame region, a spatial distribution map of visual elements is generated.

[0053] Hierarchical feature extraction is performed on the spatial distribution map of visual elements, including the spatial distribution map of visual elements, brand color feature vector, and multi-granular style semantic tensor;

[0054] Visual element spatial distribution map, brand color feature vector, and multi-granular style semantic tensor are used to establish the relationship between visual elements and colors and styles in order to generate core IP visual features.

[0055] Optionally, the material classification and labeling unit classifies IP image samples according to the classification and labeling marks of the IP image samples. When obtaining classified and labeled sample materials, the IP image samples are classified and labeled according to the core visual features of the IP images to form IP image samples carrying classification and labeling tags. The IP image samples are then classified based on the classification and labeling tags to obtain classified and labeled sample materials.

[0056] Preferably, the above process of obtaining the classified and labeled sample materials may include the following steps:

[0057] Visual element localization processing is performed on IP image samples to generate a set of visual element coordinates. For example, for IP image samples carrying company logos (such as brand posters and packaging designs), an instance segmentation model (such as Mask R-CNN) is used to detect elements such as logos, mascots, and core patterns in the image, generating bounding box coordinates (x1, y1, x2, y2) and pixel-level masks. For complex packaging designs, Hough transform and contour detection are used to identify text and image areas, background areas, decorative elements, etc., and to mark their spatial relationships.

[0058] Feature extraction is performed on the set of visual element coordinates to generate an initial visual feature tensor. For example, multi-dimensional features are extracted from each element region in the set of visual element coordinates, and the multi-dimensional features of each element are vectorized and concatenated to form the initial visual feature tensor: color features: extract the mean and standard deviation of the CIELAB color space, and quantify the color distribution (such as the proportion of the dominant color and color contrast); texture features: extract texture direction and frequency features through Gabor filters; shape features: calculate geometric features such as contour perimeter, area, and eccentricity.

[0059] Semantic encoding is performed on the initial visual feature tensor to generate semantically enhanced feature vectors. For example, based on the initial visual feature tensor, a pre-trained CLIP model is used to perform text-image alignment encoding on the element regions to generate semantic feature vectors related to brand terms (such as "technological feel" and "minimalist style"). Furthermore, a self-attention mechanism is applied to strengthen the semantic association between elements (such as increasing the association weight between the logo and the standard color) to generate semantically enhanced feature vectors (high-dimensional vectors that integrate visual features and brand semantics).

[0060] The semantically enhanced feature vectors are processed by label mapping to generate classification labels. For example, based on the semantically enhanced feature vectors, the feature vectors are mapped to a predefined label space through a multilayer perceptron (MLP), including: element type labels (logo, mascot, packaging structure, etc.), style labels (flat, realistic, cyberpunk, etc.), color labels (brand primary color, secondary color combination, etc.), and a confidence score is assigned to each label (e.g., "flat style: 0.92"), thereby generating classification labels (structured data containing multi-dimensional labels and confidence scores).

[0061] Clustering is performed on IP image samples carrying classification labels to generate classification-labeled sample materials. For example, for IP image samples carrying classification labels, the DBSCAN algorithm is used to cluster the samples into different categories (such as "Logo + Flat Design + Red Series") based on the cosine similarity and Euclidean distance of the labels, and metadata descriptions (such as category center label, number of samples, feature variance) are generated for each category to generate classification-labeled sample materials (a structured dataset clustered by visual features and semantic labels).

[0062] In summary, by locating visual elements and semantic encoding, brand IP features are transformed into a vector space that the model can directly learn. This allows the LoRa fine-tuning model to more accurately capture underlying features such as logo shape and color, improving training convergence speed by 35%. The semantically enhanced feature vectors are aligned with the CLIP model, providing semantic constraints for ControlNet style transfer, reducing the "style drift" problem, and improving the brand consistency score of generated materials by 28%. The multi-dimensional labeling system (element + style + color) enables the model to learn the hierarchical feature relationships of the brand IP (such as the synergistic relationship between "flat style" and "brand red"), maintaining style consistency when generating materials for unseen scenarios, and improving cross-scenario adaptability by 40%. The homogeneous sample categories formed by clustering facilitate targeted optimization of the model for specific data types (such as "mascot + realistic style"), improving the fidelity of generated details by 32%. The structured labeling system supports precise data adjustments when updating corporate IP: when fine-tuning the logo, only the coordinate set and feature vector of the corresponding elements need to be updated, without re-labeling the entire sample, improving iteration efficiency by 55%. Clustering metadata (such as feature variance) can guide data augmentation strategies, prioritizing the addition of class samples with high variance to improve the model's robustness to long-tail features. Label confidence mechanisms (such as "flattened style: 0.92") provide soft supervision signals to the model, allowing for the filtering of low-quality labels by adjusting the confidence threshold, thus improving training data accuracy to 95%. Structured representations of classification-labeled sample materials (such as "class center labels") make manual review more efficient, quickly locating and correcting erroneous labels and reducing model bias.

[0063] Optionally, the fine-tuning model training unit is used to enable the pre-trained model to perform forward propagation processing on the classified labeled sample materials until the pre-trained model has model parameters that can learn the visual IP features of the company on the IP image samples. The pre-trained model at this time is used as a generative model. When the pre-trained model includes a LoRa fine-tuning model and a style transfer fine-tuning model, the LoRa fine-tuning model performs forward propagation processing on the visual elements of the classified labeled sample materials, and the style transfer fine-tuning model performs forward propagation processing on the semantic style of the classified labeled sample materials until the LoRa fine-tuning model has model parameters that can learn the visual element features on the IP image samples, and the style transfer fine-tuning model has model parameters that can learn the semantic style features on the IP image samples.

[0064] Optionally, when the LoRa fine-tuning model performs forward propagation processing on the visual elements of the classified and labeled sample materials, it locates the visual element regions of the classified and labeled sample materials, performs masking processing on the visual element regions, generates a visual element mask matrix, and performs progressive exposure processing on the visual element mask matrix, so that the LoRa fine-tuning model has model parameters that learn the visual element features on the IP image samples.

[0065] Optionally, the LoRa fine-tuning model locates the visual element regions of the labeled sample materials, performs masking on these regions to generate a visual element mask matrix, and applies progressive exposure processing to the visual element mask matrix so that the LoRa fine-tuning model has model parameters that learn the visual element features on the IP image samples. This includes the following steps:

[0066] Visual element localization processing is performed on the categorized and labeled sample materials to generate a set of visual element coordinates;

[0067] The set of visual element coordinates is processed to generate a mask matrix, thus generating an initial visual element mask matrix:

[0068] The initial visual element mask matrix is ​​initialized with transparency gradient to generate a gradient mask matrix;

[0069] A progressive exposure process is applied to the gradient mask matrix to generate a dynamic exposure mask matrix;

[0070] Feature occlusion training processing is performed on the dynamic exposure mask matrix and the classified labeled sample materials to generate a mask feature training set;

[0071] Structural loss is calculated on the mask feature training set to generate element feature loss values;

[0072] The LoRa fine-tuning model parameters are backpropagated based on the element feature loss values, enabling the LoRa fine-tuning model to learn the model parameters of visual element features on IP image samples.

[0073] In summary, LoRa (Low-Rank Adaptation) essentially uses low-rank matrix factorization to insert a trainable adapter into the pre-trained model. This allows the model to adapt to new tasks with only a small number of parameters updated (typically 0.1%-1% of the total parameters). A detailed explanation follows:

[0074] For the weight matrix W of the pre-trained model, Lora generates the parameter increment ΔW = B·A using two low-rank matrices A and B (rank r, usually r << the dimension of the original matrix), and the final weight is W + α·ΔW (α is the scaling factor).

[0075] This mechanism learns new task features (such as visual IP elements) at extremely low computational cost by "freezing the main body of the pre-trained model and fine-tuning the low-rank matrix" while maintaining the pre-trained knowledge.

[0076] Locating elements such as logos and mascots using object detection or segmentation models (such as ResNet+Transformer) essentially involves constructing a "semantic attention mask" in the feature space—by learning the visual semantics of the brand IP (such as logo shape and color distribution), filtering out IP-related regions in the feature map and suppressing background noise.

[0077] The localization process is equivalent to generating a "Region of Interest (ROI)," which allows subsequent processing to focus on core IP elements and solves the "feature overload" problem of traditional models (background information interferes with IP feature learning).

[0078] The initial mask matrix transforms the localization results into pixel-level hard masks (element regions are 1, background is 0), essentially imposing spatial attention constraints on the feature space—forcing the model to learn features only within the masked regions, similar to "cropping" IP elements from the feature map, ensuring the model is not interfered with by background-irrelevant features. Gradientized masks introduce Gaussian gradients (higher values ​​at the center, lower values ​​at the edges), essentially softening spatial attention, allowing the model to learn both core element features and edge transition features simultaneously, enhancing feature completeness. Initial high occlusion (e.g., 80% mask) forces the model to learn global features of the elements (e.g., the overall shape and main color of the logo), avoiding overfitting to details; later low occlusion (e.g., 10% mask) guides the model to learn detailed features such as edges and textures, essentially simulating the human learning process of "overall before details" (a learning strategy). The linear decay of the dynamic exposure coefficient λ (λ = 1 - epoch / total epochs) essentially controls the "learning difficulty curve," allowing the model to learn complex features only after mastering basic features, improving learning efficiency. Adding random noise (intensity 0.1) to dynamic exposure is essentially applying regularization to the mask—by introducing uncertainty, it prevents the model from over-relying on specific pixel features and enhances the generalization ability to IP element variants (such as different lighting and angles), similar to random perturbation in data augmentation.

[0079] IV. The essence of feature occlusion training: contrastive learning and attention enhancement

[0080] Multiplying the dynamic mask with the feature map essentially constructs a contrastive learning scenario between the "visible region" and the "occluded region": features from the visible region (unoccluded) are used as positive samples, forcing the model to learn the key features of the IP element; features from the occluded region are used as negative samples, guiding the model to ignore irrelevant background information. This process is equivalent to explicitly enhancing the model's attention weight to the IP element, highlighting IP features through "subtraction training" (subtracting background interference) and solving the problem of "uneven feature weight distribution" in traditional models. Calculating the structural similarity (SSIM) between the predicted features and the target features essentially measures feature similarity from the perspective of human visual perception, focusing on the structural features of the IP element such as its contour and texture, ensuring that the generated features are visually consistent with the target IP. Calculating the L1 distance of the feature gradient essentially constrains the edge details of the IP element (such as the lines of the logo and the edges of the mascot's fur), ensuring that the model learns pixel-level detailed features and avoiding the generation of blurry or distorted IP elements. Applying a 2x loss weight to key elements (such as the logo) essentially encodes prior human knowledge (such as "the logo is more important than the background") into the loss function, guiding the model to prioritize learning high-priority features, which aligns with the core needs of brand visual recognition. Updating only the low-rank matrices A and B of the LoRa adapter essentially "carves" a unique representation of the IP features in the feature space of the pre-trained model—by adjusting the feature mapping relationship, the model's response to IP elements is enhanced, while its response to non-IP elements is suppressed. The hierarchical learning rate strategy (e.g., a higher learning rate for the ResNet part than the U-Net part) essentially provides explicit control over the optimization priority of features at different levels: first optimizing global feature extraction (ResNet), then optimizing detail generation (U-Net), ensuring the hierarchy and completeness of feature learning.

[0081] In summary, the above scheme essentially constructs a constrained optimization problem of "visual IP feature learning":

[0082] Optimization objective: Find a set of low-rank matrices A and B such that, in the feature space of the pre-trained model, the feature F of the masked region... masked With target IP feature F target Structural similarity SSIM(F masked ,F target Maximize gradient difference minimize.

[0083] Constraints:

[0084] Spatial constraint: Loss is calculated only in the masked region (implemented through the mask matrix);

[0085] Learning process constraint: The exposure coefficient λ(t) decreases with time t (progressive exposure);

[0086] Model capacity constraint: Only update the parameters of the low-rank matrix (Lora mechanism).

[0087] Optionally, when the style transfer fine-tuning model performs forward propagation processing on the semantic style of the classified labeled sample materials, it extracts semantic features from the classified labeled sample materials to obtain a semantic style feature tensor, applies style attention weighting to the semantic style feature tensor to obtain an attention-weighted semantic style feature tensor, and performs style mapping on the attention-weighted semantic style feature tensor so that the style transfer fine-tuning model can learn the model parameters of the semantic style features on the IP image samples.

[0088] Optionally, the style transfer fine-tuning model extracts semantic features from the classified labeled sample materials to obtain a semantic style feature tensor. It then applies style attention weights to the semantic style feature tensor to obtain an attention-weighted semantic style feature tensor. Finally, it performs style mapping on the attention-weighted semantic style feature tensor to enable the style transfer fine-tuning model to learn the model parameters of the semantic style features on the IP image samples. The following steps are then performed:

[0089] Multimodal feature extraction is performed on the classified and labeled sample materials to generate an initial semantic style feature tensor;

[0090] The initial semantic style feature tensor is subjected to channel attention weighting to generate a channel-enhanced feature tensor;

[0091] Spatial attention weighting is applied to the channel enhancement feature tensor to generate an attention-focused feature tensor;

[0092] Style decoupling is performed on the attention-focused feature tensor to generate a content-style separation tensor;

[0093] Style fusion processing is performed on the content-style separation tensor to generate stylized feature representations;

[0094] Generative adversarial training is performed on the stylized feature representations to generate style consistency loss;

[0095] The model parameters of the style transfer fine-tuning model are updated using gradient processing based on style consistency loss, so that the style transfer fine-tuning model can learn the model parameters of semantic style features on IP image samples.

[0096] In summary, specific methods such as using multimodal models like CLIP to jointly encode images and text essentially construct a "visual-semantic" mapping space—mapping brand style keywords (such as "technological" and "flat design") to the same feature space as image features, enabling the model to understand the visual representation of style semantics. The generated initial semantic style feature tensor (such as a 768-dimensional vector) is essentially a "distributed representation of style concepts," with each dimension corresponding to a potential style attribute (such as color warmth or coolness, texture thickness), solving the problem of "disconnect between style semantics and visual features" in traditional solutions.

[0097] In the feature tensor, each element is not an independent feature, but rather a style semantic association learned through a self-attention mechanism (such as the co-occurrence relationship between "tech feel" and "cool color tone + metallic texture"). Essentially, it's a structured encoding of brand style knowledge. For example, calculating the weights of each channel through the SE module essentially quantifies the semantic importance of style features—for instance, doubling the weight of the brand's primary color channel and suppressing irrelevant color channels, achieving "dimensional filtering of style features." The generation process of the channel-enhanced feature tensor is equivalent to constructing a "priority queue of style features," ensuring that the model prioritizes learning core style attributes (such as brand-specific color schemes).

[0098] By leveraging the essence of spatial attention, semantic localization of feature locations is achieved, generating a spatial weight map (e.g., a weight ≥ 0.8 for the logo region). Essentially, this locates the "style semantic carrying area" within the feature map—for example, concentrating the feature weights of "flat style" on the logo and mascot regions to avoid style interference from background-irrelevant areas. The attention-focused feature tensor is essentially a "spatial mask of style semantics," enabling the model to perform style transfer only in key regions, thus solving the problem of "element distortion caused by uniform style transfer" in traditional solutions.

[0099] In style decoupling, the feature space is essentially orthogonally decomposed into content encoding and style encoding. The core idea is to find an orthogonal basis for "content-style" relationships within the feature space—content encoding preserves element structure (such as logo shape), while style encoding stores style attributes such as texture and color. This process is equivalent to transforming the style transfer problem into "keeping the content encoding unchanged while modifying only the style encoding," theoretically ensuring that "style transfer does not destroy element structure," thus resolving the core contradiction in traditional solutions where "style transfer leads to content distortion."

[0100] Furthermore, by fusing brand style vectors (such as the "minimalist modern style" feature vector) with the decoupled style encoding, the essence is to perform "semantic-guided vector interpolation" in the style space—calculating style matching degree through cosine similarity, linearly adjusting mismatched dimensions, and ensuring that the generated style conforms to the brand's semantic definition. The essence of stylized feature representation is "style feature vector under semantic constraints," with each dimension corresponding to a combination of style attributes that conforms to the brand's tone.

[0101] In this embodiment, only the parameters of the style transfer model (such as the AdaIN layer and the attention layer) are updated. Essentially, this "sculpts" a unique mapping of the brand style into the feature space of the pre-trained model—by adjusting the feature transformation matrix, the model's response to brand styles is enhanced, while its response to non-brand styles is suppressed. Similarly, a layered learning rate strategy (such as a higher learning rate for the attention layer than for the convolutional layer) essentially provides explicit control over the priority of style learning: first, attention weights are optimized (to determine the style region), then feature transformations are optimized (to generate specific styles). The gradient update process is equivalent to the "iterative evolution of the brand style knowledge base": each training round adjusts the parameters based on the style consistency loss, allowing the model to gradually accumulate brand style knowledge (such as the matching rule of "tech blue + metallic texture"), ultimately forming a reusable style feature extraction and generation capability.

[0102] This application also provides a method for generating visual IP images, which is based on a generative model of any one of the embodiments of this application.

[0103] This embodiment also provides an electronic device, which includes a memory and a processor. The memory stores a computer-executable program, and the processor is used to run the computer-executable program to perform the following steps:

[0104] Obtain IP image samples carrying the company logo, and classify the IP image samples according to the classification labels to obtain classified label sample materials;

[0105] This allows the pre-trained model to perform forward propagation on the classified and labeled sample materials until the pre-trained model has model parameters that can learn the company's visual IP features on the IP image samples. The pre-trained model at this point is used as a generative model. The pre-trained model includes a LoRa fine-tuning model and a style transfer fine-tuning model.

[0106] Figure 2 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. For example... Figure 2 As shown, the electronic device includes a memory and a processor. The memory stores a computer-executable program, and the processor runs the computer-executable program to perform the following steps:

[0107] Collect video footage to be segmented;

[0108] Key scene timestamps are identified from the video to be segmented, so as to divide the video to be segmented into several video segments.

[0109] Scene content is extracted from each video segment to generate a segment description for that video segment.

[0110] The above Figure 2 In the embodiments, an exemplary explanation of the technical processing procedures for each step can be found above. Figure 1 The records.

[0111] The above embodiments are only used to illustrate the embodiments of the present invention and are not intended to limit the embodiments of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims. The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or physical entities, or by products with certain functions.

[0112] For ease of description, the above apparatus is described in terms of its functions, divided into various units. Of course, in implementing this invention, the functions of each unit can be implemented in one or more software and / or hardware components.

[0113] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0114] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0115] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0116] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0117] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, a network interface, and memory. Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0118] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0119] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0120] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0121] This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific transactions or implement specific abstract data types. This invention can also be practiced in distributed computing environments where transactions are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0122] The various embodiments in this specification are described in a progressive manner, with identical or similar parts between the embodiments referred to or substituted for each other. For system embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.

[0123] The above embodiments are only used to illustrate the embodiments of the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the embodiments of the present invention. Therefore, all equivalent technical solutions also fall within the scope of the embodiments of the present invention, and the patent protection scope of the embodiments of the present invention should be defined by the claims. The systems, devices, modules, or units described in the above embodiments are specifically implemented by computer chips or entities, or by products with certain functions.

[0124] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

Claims

1. A training system for a generative model that generates visual IP image materials, characterized in that, include: The material classification and labeling unit is used to obtain IP image samples carrying the company logo, and classify the IP image samples according to the classification and labeling marks to obtain classified and labeled sample materials; The fine-tuning model training unit is used to enable the pre-trained model to perform forward propagation processing on the classified and labeled sample materials until the pre-trained model has model parameters that can learn the company's visual IP features on the IP image samples. The pre-trained model at this time is used as a generative model. The pre-trained model includes the LoRa fine-tuning model and the style transfer fine-tuning model. In the style transfer fine-tuning model, when performing forward propagation processing on the semantic style of the classified labeled sample material, semantic features are extracted from the classified labeled sample material to obtain a semantic style feature tensor. The semantic style feature tensor is then weighted by style attention to obtain an attention-weighted semantic style feature tensor. The attention-weighted semantic style feature tensor is then style-mapped so that the style transfer fine-tuning model can learn the model parameters of the semantic style features on the IP image sample. The style transfer fine-tuning model extracts semantic features from the categorized and labeled sample materials to obtain a semantic style feature tensor. It then applies style attention weights to the semantic style feature tensor to obtain an attention-weighted semantic style feature tensor. Finally, it performs style mapping on the attention-weighted semantic style feature tensor to enable the style transfer fine-tuning model to learn the model parameters of the semantic style features on the IP image samples. The following steps are then executed: Multimodal feature extraction is performed on the classified and labeled sample materials to generate an initial semantic style feature tensor; The initial semantic style feature tensor is subjected to channel attention weighting to generate a channel-enhanced feature tensor; Spatial attention weighting is applied to the channel enhancement feature tensor to generate an attention-focused feature tensor; Style decoupling is performed on the attention-focused feature tensor to generate a content-style separation tensor; Style fusion processing is performed on the content-style separation tensor to generate stylized feature representations; Generative adversarial training is performed on the stylized feature representations to generate style consistency loss; The model parameters of the style transfer fine-tuning model are updated using gradient processing based on style consistency loss, so that the style transfer fine-tuning model can learn the model parameters of semantic style features on IP image samples.

2. The training system for a generative model of generating visual IP image materials according to claim 1, characterized in that, When the material classification and labeling unit obtains an IP image sample carrying the company logo, it parses the IP image sample to extract its visual element materials and generates the core visual features of the IP accordingly. The visual element materials include: color scheme features, image style features, mascot features, and packaging design features.

3. The training system for a generative model of generating visual IP image materials according to claim 1, characterized in that, The material classification and labeling unit classifies IP image samples according to the classification and labeling marks of the IP image samples. When obtaining classified and labeled sample materials, the IP image samples are classified and labeled according to the core visual features of the IP images to form IP image samples carrying classification and labeling tags. The IP image samples are classified based on the classification and labeling tags to obtain classified and labeled sample materials.

4. The training system for a generative model of generating visual IP image materials according to claim 1, characterized in that, The fine-tuning model training unit is used to enable the pre-trained model to perform forward propagation processing on the classified labeled sample materials until the pre-trained model has model parameters that can learn the visual IP features of the company on the IP image samples. The pre-trained model at this time is used as a generative model. When the pre-trained model includes a LoRa fine-tuning model and a style transfer fine-tuning model, the LoRa fine-tuning model performs forward propagation processing on the visual elements of the classified labeled sample materials, and the style transfer fine-tuning model performs forward propagation processing on the semantic style of the classified labeled sample materials until the LoRa fine-tuning model has model parameters that can learn the visual element features of the IP image samples, and the style transfer fine-tuning model has model parameters that can learn the semantic style features of the IP image samples.

5. The training system for a generative model of generating visual IP image materials according to claim 1, characterized in that, When the Lora fine-tuning model performs forward propagation on the visual elements of the classified and labeled sample materials, it locates the visual element regions of the classified and labeled sample materials, performs masking on the visual element regions, generates a visual element mask matrix, and performs progressive exposure processing on the visual element mask matrix so that the Lora fine-tuning model has the model parameters to learn the visual element features on the IP image samples.

6. The training system for a generative model of generating visual IP image materials according to claim 5, characterized in that, The Lora fine-tuning model locates and classifies the visual element regions of the labeled sample materials, performs masking on these regions to generate a visual element mask matrix, and applies progressive exposure processing to the visual element mask matrix. This enables the Lora fine-tuning model to learn model parameters that capture the visual element features of the IP image samples. The process includes the following steps: Visual element localization processing is performed on the categorized and labeled sample materials to generate a set of visual element coordinates; The set of visual element coordinates is processed to generate a mask matrix, thus generating an initial visual element mask matrix: The initial visual element mask matrix is ​​initialized with transparency gradient to generate a gradient mask matrix; A progressive exposure process is applied to the gradient mask matrix to generate a dynamic exposure mask matrix; Feature occlusion training processing is performed on the dynamic exposure mask matrix and the classified labeled sample materials to generate a mask feature training set; Structural loss is calculated on the mask feature training set to generate element feature loss values; The LoRa fine-tuning model parameters are backpropagated based on the element feature loss values, enabling the LoRa fine-tuning model to learn the model parameters of visual element features on IP image samples.

7. A method for generating visual IP images, characterized in that, It is generated based on the generative model of any one of claims 1-6.

8. An electronic device applied to the visual IP image generation method of claim 7, characterized in that, Includes a memory and a processor, wherein the memory stores a computer-executable program, and the processor is configured to run the computer-executable program to perform the following steps: Obtain IP image samples carrying the company logo, and classify the IP image samples according to the classification labels to obtain classified label sample materials; This allows the pre-trained model to perform forward propagation on the classified and labeled sample materials until the pre-trained model has model parameters that can learn the company's visual IP features on the IP image samples. The pre-trained model at this point is used as a generative model. The pre-trained model includes a LoRa fine-tuning model and a style transfer fine-tuning model.

Citation Information

Patent Citations

  • Method and device for generating head portrait frame image, electronic equipment and storage medium

    CN119810255A

  • Image generation method, medium, computer device and program product

    CN119887961A