Interactive painting generation method, system and device based on artificial intelligence
The interactive painting generation method aligns and fuses multi-modal data to improve semantic accuracy and detail in generated images by using a unified semantic space and iterative correction, addressing inconsistencies and enhancing user intent alignment.
Patent Information
- Application Number
- CN202510398905.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art is difficult to fully utilize the complementary information of multimodal data in image generation, and lacks an effective feature alignment mechanism, resulting in the generation of images that are inconsistent with the user's intention, lacks real-time interaction capabilities, and cannot meet personalized needs.
By obtaining multimodal data (text, speech, sketch), feature alignment and weighting fusion are performed, unified semantic spatial features are generated, semantic inconsistencies are detected using the visual-text attention mechanism, conflict exponential matrix is generated, alarm marks are triggered, and images are generated are generated by iterative corrections.
The semantic accuracy and detail richness of image generation are achieved, ensuring that the image is consistent with the user's intentions, having real-time interaction capabilities, and meeting personalized needs.
Smart Images

Figure CN120318354A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and specifically to an interactive painting generation method, system and device based on artificial intelligence. Background Art
[0002] With the rapid development of artificial intelligence technology, image generation technology has gradually become a research hotspot. Traditional image generation methods mainly rely on single-modal data (such as text or sketches), and generate images through generative adversarial networks (GANs) or diffusion models (Diffusion Model). However, these methods have the following limitations: traditional methods usually only rely on single-modal data (such as text or sketches), and it is difficult to make full use of the complementary information of multi-modal data (such as text, speech, sketches), resulting in insufficient semantic accuracy and detail richness of the generated images; in the process of multi-modal data fusion, the existing technology lacks an effective feature alignment mechanism, and it is difficult to ensure the consistency of different modal data in the semantic space, resulting in the generated images not conforming to the user's intention. Summary of the Invention
[0003] Aiming at the deficiencies of the existing technology, the present invention provides an interactive painting generation method, system and device based on artificial intelligence.
[0004] In order to achieve the above purpose, the technical solution of the present invention is as follows:
[0005] An interactive painting generation method based on artificial intelligence, including the following steps:
[0006] Step 1: Obtain the original multi-modal data, and after processing the original multi-modal data, obtain a formatted feature dataset containing multi-modal features; the original multi-modal data includes the collected original text data, original speech data, and original sketch data;
[0007] Step 2: Align the multi-modal features in the formatted feature dataset, and assign weights according to the importance of different modalities. After weighted fusion of the multi-modal features, perform normalization processing to obtain unified semantic space features;
[0008] Step 3: Locate the regions where the image and the text description are inconsistent through a visual-text attention mechanism according to the unified semantic space features, and quantify the semantic inconsistency degree between the image and the text description to obtain a conflict index matrix. Determine whether the conflict index matrix triggers an alarm mark according to a dynamic threshold;
[0009] Step 4: Iteratively correct the generated image based on the unified semantic space features, the conflict index matrix, and the alarm mark.
[0010] Preferably, the present invention also discloses an interactive painting generation system based on artificial intelligence, which implements the above-mentioned interactive painting generation method based on artificial intelligence, including:
[0011] A data collection and preprocessing module, responsible for collecting raw multi-modal data and performing preliminary processing to generate a formatted feature data set;
[0012] A multi-modal feature alignment and fusion module, which aligns and weights and fuses multi-modal features to generate unified semantic space features;
[0013] A conflict detection and alarm module, which detects the semantic inconsistency between the image and the text description, generates a conflict index matrix, and triggers an alarm according to a dynamic threshold;
[0014] An image generation and iterative correction module, which generates and iteratively corrects images based on the unified semantic space features, the conflict index matrix, and the alarm mark;
[0015] A system control and scheduling module, responsible for the scheduling and resource management of the entire system to ensure the collaborative work of each module.
[0016] Preferably, the present invention also discloses an interactive painting generation device based on artificial intelligence, which applies the above-mentioned interactive painting generation method based on artificial intelligence.
[0017] Compared with the prior art, the beneficial effects of the present invention are:
[0018] 1. By fusing multi-modal data such as text, voice, and sketches, making full use of the complementary information of different modalities, the generated images are more in line with the user's intention in terms of semantics and details;
[0019] 2. Through feature alignment and normalization processing, ensure the consistency of different modality data in the unified semantic space, and avoid the generated images not conforming to the user's intention;
[0020] 3. Through the visual-text attention mechanism and the conflict index matrix, accurately locate and quantify the semantic inconsistency between the image and the text description, and combine the dynamic threshold to trigger the alarm mark to timely correct the error area. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The disclosure of the present invention will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them:
[0022] Figure 1 is the step flow chart of the present invention;
[0023] Figure 2 is the data flow diagram in Step 1 of the present invention;
[0024] Figure 3 It is the specific flowchart of Step 2 of the present invention;
[0025] Figure 4 It is the specific flowchart of Step 3 of the present invention;
[0026] Figure 5 It is the sketch of the initial image static horseshoe in Embodiment 1 of the present invention;
[0027] Figure 6 It is the sketch of the corrected image dynamic horseshoe in Embodiment 1 of the present invention. Detailed implementation manners
[0028] It is easy to understand that according to the technical solution of the present invention, without changing the essence of the present invention, those of ordinary skill in the art can propose various structural ways and implementation manners that can be mutually replaced. Therefore, the following detailed implementation manners and the accompanying drawings are only exemplary descriptions of the technical solution of the present invention, and should not be regarded as all of the present invention or as a limitation or restriction on the technical solution of the present invention.
[0029] Application overview
[0030] As described above, the image generation technology of artificial intelligence is one of the important developments in the fields of computer vision and deep learning in recent years. It uses a deep neural network model to generate realistic or artistic images.
[0031] The image generation technology based on artificial intelligence mainly relies on single-modal data (such as text or sketch), and has the following defects:
[0032] Modal singularity: It cannot make full use of the complementary information of multi-modal data (such as text, voice, sketch), resulting in insufficient semantic accuracy and detail richness of the generated images.
[0033] Insufficient semantic alignment: Lack of an effective cross-modal alignment mechanism, the generated images do not match the user's intention, and it is difficult to meet complex requirements.
[0034] Missing conflict detection: It cannot dynamically detect and correct the semantic inconsistency between the image and the text description, affecting the image quality.
[0035] Lack of interactivity: Lack of real-time interaction ability, unable to dynamically adjust the generated images according to user feedback, and it is difficult to meet personalized requirements.
[0036] In view of the above-mentioned defects in the prior art, the basic concept of this application is to solve the above-mentioned defects through multi-modal data fusion, semantic alignment, conflict detection and iterative correction. Obtain multi-modal data such as text, speech, and sketches, perform feature alignment and weighted fusion, generate unified semantic space features, and improve the semantic accuracy and detail richness of images. Use the visual-text attention mechanism to detect the semantic inconsistency between the image and the text description, generate a conflict index matrix and combine it with a dynamic threshold to trigger an alarm mark to ensure that the image is consistent with the user's intention. Figure 1 consistent. Based on the unified semantic space features, conflict index matrix and alarm marks, iteratively correct and generate images, and dynamically adjust in combination with user feedback to meet personalized needs. Through modules such as data collection and preprocessing, multi-modal feature alignment and fusion, conflict detection and alarm, and image generation and iterative correction, efficient and intelligent image generation is realized.
[0037] After introducing the basic concept of the present invention, the embodiments of the present invention will be specifically introduced below with reference to the accompanying drawings.
[0038] Embodiment 1:
[0039] As Figure 1 shown, the present invention discloses an interactive painting generation method based on artificial intelligence, including the following steps:
[0040] Step 1: Obtain the original multi-modal data, and obtain a formatted feature data set containing multi-modal features after processing the original multi-modal data; the original multi-modal data includes the collected original text data, original speech data, and original sketch data.
[0041] As Figure 2 shown, it is the data flow diagram in Step 1;
[0042] Obtaining a formatted feature data set after processing the original multi-modal data includes:
[0043] Text feature extraction:
[0044] Use the BERT-base model for semantic encoding, that is: T = BERT base (t) ∈ R 768 ; where t is the string of the original text input, BERT base is a 12-layer Transformer encoder, and T is the output 768-dimensional semantic vector.
[0045] Example:
[0046] User input: t = "A running horse, with its four legs stretched and its mane flying";
[0047] Output vector: T = [0.23, -0.56,..., 0.78], which is a 768-dimensional semantic vector.
[0048] Speech feature extraction:
[0049] Convert speech to text: S text = Whisper(s audio ); where s audio is the speech waveform data, Whisper is a speech-to-text model that supports multilingual recognition, and S text is the output result of speech-to-text;
[0050] Multi-feature fusion: Among them, BiLSTM is a bidirectional long short-term memory network (Bidirectional LSTM) used to extract the semantic features of the text and output a 256-dimensional vector; Prosody is a prosody feature extractor that extracts acoustic features such as fundamental frequency (Pitch) and energy (Energy) from the audio and outputs a 256-dimensional vector; is a vector concatenation operation that concatenates two 256-dimensional vectors into a 512-dimensional vector, and S is the output 512-dimensional speech fusion vector.
[0051] Example:
[0052] A 3-second audio file with a sampling rate of 16 kHz;
[0053] User speech input: s audio = "The horse runs very fast and its hooves leave the ground" (recording for 3 seconds);
[0054] Whisper transcription result: S text = = "The horse runs very fast and its hooves leave the ground";
[0055] 256-dimensional vector output by BiLSTM: S BiLSTM = [0.12, -0.34,..., 0.91];
[0056] Prosody outputs a 256-dimensional vector: S Prosody = [45.6Hz, 120dB,..., 0.78];
[0057] 512-dimensional speech fusion vector output after concatenation:
[0058] Sketch feature extraction:
[0059] Stroke encoding based on a graph convolutional network (GCN), that is: I = GCN(Stroke2Graph(I raw )) ∈ R 256 ; where I rawIndicates the strokes drawn by the user through the digital tablet, which is a sequence of coordinates recorded by the digital tablet and consists of N coordinate points, denoted as Indicates the digital tablet coordinate sequence; Stroke2Graph is to convert the original stroke data I raw into a converter of graph structure. The nodes in the graph represent the turning points of the strokes, that is, the points where the stroke direction changes significantly. The edges in the graph connect adjacent nodes and have a direction attribute, indicating the drawing order of the strokes; GCN is a 3-layer graph convolutional network used for feature extraction and encoding of graph structure data. The dimension of the hidden layer is 128, and the finally output I is a 256-dimensional sketch vector.
[0060] Example:
[0061] User hand-drawn input: I raw = {(0,0),(12,5),...,(100,80)}, (a total of 50 coordinate points).
[0062] Output vector: I = [-0.45,0.67,...,0.02].
[0063] The formatted feature dataset obtained after processing the original multi-modal data also includes:
[0064] Aligning the data of text, voice and sketch in time through the timestamps of the hardware on the input device: Figure 3 D = Align(T,S,I), and it satisfies |t
[0065] - t text - t sketch | < 50ms; where t text is the start time of text input, and t sketch is the start time of sketch input.
[0066] Example: The voice description is completed at t text = 1000ms, and the sketch is completed at t sketch = 1020ms. At this time, the condition of time alignment is satisfied.
[0067] In addition, the average time difference between the corresponding time points of the text and the sketch can also be calculated to ensure that the delay between the two is less than 50 milliseconds, that is, the synchronization accuracy High synchronization accuracy is the key to ensuring that multi-modal data (such as text and sketch) can be matched in real time during the interactive painting generation process, avoiding the generated image not matching the user's intention due to the time difference.
[0068] Example: The voice description is completed at t text = 1000ms, and the sketch is completed at t sketch = 1020ms, then:
[0069] Since 20 ms < 50 ms, it is determined that the synchronization accuracy between the text and the sketch meets the requirements.
[0070] The cosine similarity calculation method can also be used: semantic similarity This metric is used to measure the semantic similarity between the text and the sketch. T represents the text feature vector, and S text is the output result of speech-to-text corresponding to the sketch. The calculation result is required to be greater than 0.85. High semantic similarity ensures that the generated image is highly consistent semantically with the text description input by the user, avoids semantic deviation, and improves the quality and accuracy of image generation.
[0071] Example: T = [0.9, 0.8, 0.7], representing the semantic features of "blue sky, white clouds, and green grassland".
[0072] S text = [0.85, 0.75, 0.65], representing the semantic features corresponding to the sketch.
[0073] According to the cosine similarity formula:
[0074] Since 0.92 > 0.85, it is determined that the semantic fidelity between the text and the sketch meets the requirements.
[0075] Step 2: Align the multi-modal features in the formatted feature dataset, assign weights according to the importance of different modalities, and perform normalization processing after weighted fusion of the multi-modal features to obtain unified semantic space features.
[0076] The implementation process is as Figure 3 shown, specifically:
[0077] Inherit the output of Step 1: D = {T ∈ R 768 , S ∈ R 512 , I ∈ R 256};
[0078] Example data:
[0079] Text feature: T = [0.23, -0.56,..., 0.78] ("horse running on the grassland");
[0080] Voice feature: S = [0.12, -0.34,..., 0.91] ("the horse's hooves should be in the air");
[0081] Sketch feature: I = [-0.45, 0.67,..., 0.02] (static horse hoof drawing).
[0082] Use the extended CLIP architecture for feature alignment:
[0083]
[0084] Among them, Q = W q T, where T is a 768-dimensional semantic vector of text features, and Q is a 512-dimensional query vector after projection;
[0085] K = W k [S; I], where [S; I] is a concatenated vector of speech feature S and sketch feature I, with a dimension of 768, and K is a 512-dimensional key vector after projection;
[0086] V = W v [S; I], and V is a 512-dimensional value vector after projection;
[0087] W * is a learnable parameter matrix that projects features of different modalities into a unified 512-dimensional semantic space, including query projection matrix W q 、key projection matrix W k , value projection matrix W v ;
[0088] W q ∈R 512×768 Specifically, it is a text feature projection matrix that projects 768-dimensional text features into 512 dimensions;
[0089] Example: Q = W q T = [0.12, -0.23,..., 0.45] ∈ R 512 ;
[0090] W k ∈R 512×768 Specifically, it is a speech and sketch feature projection matrix that projects 768-dimensional text features into 512 dimensions;
[0091] Example: K = W k [S; I] = [0.34, 0.56,...,-0.78] ∈ R 512 ;
[0092] W v ∈R 512×768 Specifically, it is a speech and sketch feature projection matrix that projects 768-dimensional text features into 512 dimensions;
[0093] Example: V = W v [S; I] = [0.45, -0.12,..., 0.67] ∈ R 512 ;
[0094] d is a scaling factor used to scale the attention weights to prevent gradient explosion; in the formula is used to normalize the attention scores.
[0095] The CrossModalEncoder is an encoder for cross-modal feature alignment, and F is the cross-modal alignment feature.
[0096] The Softmax function normalizes the attention scores into a probability distribution, representing the importance of different modalities.
[0097] Through QK T Calculate the similarity between the text and the speech / sketch features, and then perform weighted fusion.
[0098] Example:
[0099] Scenario: The user describes "a running horse" and hand-draws a static sketch of a horseshoe.
[0100] Calculate the attention weights:
[0101] Q = W q T = [0.12, -0.23,..., 0.45];
[0102] K = W k [S; I] = [0.34, 0.56,...,-0.78];
[0103]
[0104] Weighted fusion of features:
[0105] V = W v [S; I] = [0.45, -0.12,…, 0.67];
[0106] F = Attn·V = 0.6·V speech +0.4·V sketch , where V speech and V sketch are the parts of the speech feature and the sketch feature in V respectively.
[0107] According to the cross-modal alignment feature F obtained in the above steps, calculate the modality importance through the MLP:
[0108] α, β, γ = MLP(F), and α + β + γ = 1; where MLP is a multi-layer perceptron for calculating the modality weights; α, β, γ are the text weight, the speech weight, and the sketch weight in sequence;
[0109] Calculate the fused feature according to the assigned weights: F fused = αT + βS + γI, F fused is the fused feature.
[0110] Example:
[0111] Input cross-modal alignment feature: F = [0.12, -0.23, ..., 0.45];
[0112] Calculate modal weights:
[0113] α = 0.4, β = 0.5, γ = 0.1;
[0114] Feature fusion:
[0115] F fused = αT + βS + γI = 0.4·T + 0.5·S + 0.1·I.
[0116] After the multi-modal feature weighted fusion feature is normalized, the unified semantic space feature is:
[0117] F aligned = LayerNorm(F fused ) ∈ R 1024 , LayerNorm is the layer normalization operation, and F aligned is the unified semantic space feature.
[0118] Step 3: According to the unified semantic space feature, locate the regions where the image and the text description are inconsistent through the visual-text attention mechanism, and quantify the semantic inconsistency degree between the image and the text description to obtain the conflict index matrix. Determine whether the conflict index matrix triggers an alarm mark according to the dynamic threshold;
[0119] The implementation process is as Figure 4 shown, specifically:
[0120] Receive the unified semantic space feature F aligned output in Step 2 as the input for conflict detection.
[0121] Example:
[0122] F = [0.12, -0.23, …, 0.45], representing the fusion feature of the user's description of "running horse" and the hand-drawn sketch.
[0123] Calculate the attention weights of the image feature F img and the text feature F text :
[0124]
[0125] where AttnMap is the attention mapping function to calculate the similarity between the image and the text features; n is the total number of feature segments, indicating that the image is divided into n regions; the value of H represents the similarity between the image and the text description, the higher the value, the more consistent, and the lower the value, the more conflicting; is the image feature F imgThe $i$-th segment of, representing the feature vector of a certain region in the image; is the text feature $F$ text The $i$-th segment of, representing the feature vector of a certain keyword in the text description.
[0126] Example:
[0127] Scenario: The user describes "a running horse" and hand-draws a static sketch of a horse's hoof.
[0128] Attention heatmap calculation:
[0129] Image feature: $F$ img = [0.9, -0.2, 0.5];
[0130] Text feature: $F$ txt = [0.3, 0.1, -0.7];
[0131] Attention weight: AttnMap($F$ img , $F$ txt ) = -0.48;
[0132] Heatmap: $H$ = [0.25, 0.75, 0.40], representing the similarity of different regions.
[0133] According to the empirical thresholds $\tau$ low and $\tau$ high , identify the conflict regions:
[0134] $M$ mask = II($H \lt \tau$ low $\cup H \gt \tau$ high ), where $\tau$ low is the low similarity threshold, identifying regions with too low similarity; $\tau$ high is the high similarity threshold, identifying regions with too high similarity; II is the indicator function, which is 1 when the condition is met and 0 otherwise, that is, if $H$ i < 0.3 or $H$ i > 0.7, then the rest is 0; $M$ mask is the conflict mask.
[0135] Example:
[0136] Scenario: The user describes "a running horse" and hand-draws a static sketch of a horse's hoof.
[0137] According to the heatmap calculated above: $H$ = [0.25, 0.75, 0.40].
[0138] The low similarity threshold $\tau$ low is set to 0.3, and the high similarity threshold $\tau$ high is set to 0.7;
[0139] The conflict mask indicates that there are conflicts in the 1st and 3rd regions.
[0140] Calculate the image feature F img and the text feature F text to obtain the cosine similarity, and convert it into a conflict index C:
[0141] where n is the number of feature dimensions, is the k-th dimension of the image feature F img , is the k-th dimension of the text feature F text ;
[0142] The conflict index C is used to measure the semantic inconsistency between the image and the text description. The larger its value, the more significant the difference between the image and the text description; the smaller the value, the more consistent they are.
[0143] is the cosine similarity between the image feature F img and the text feature F text . The value range of the cosine similarity is [-1, 1]. The closer the value is to 1, the more similar they are; the closer the value is to -1, the less similar they are.
[0144] Example: Suppose the cosine similarities between the image and text features in three dimensions are 0.8, 0.6, and 0.4 respectively. Then the conflict index is calculated as follows:
[0145]
[0146] The conflict index C = 0.4, indicating that there is a certain degree of inconsistency between the image and the text description.
[0147] According to the modality divergence threshold η, determine whether to trigger an alarm:
[0148] Alarm flag
[0149] Example: Set η = 0.65;
[0150] Suppose the text description input by the user is "Draw a landscape painting with a blue sky, white clouds, and green grassland", and the cosine similarities between the image features generated by the system and the text features in three dimensions are 0.7, 0.5, and 0.3 respectively. Then the conflict index is calculated as follows:
[0151]
[0152] At this time, since 0.5 < 0.65, then Alert = False; the system determines that the image and the text description are basically consistent and no alarm needs to be triggered.
[0153] Also assume that the conflict index C = 1.48. Since 1.48 > 0.65, then:
[0154] Alert = True; The system triggers an alarm, indicating that the generated image needs to be corrected.
[0155] Through the visual-text attention mechanism and the conflict index matrix, the areas where the image and the text description are inconsistent can be accurately located and targeted corrections can be made. This significantly improves the consistency between the generated image and the user's description and reduces semantic deviation.
[0156] Step 4: Iteratively correct the generated image based on the unified semantic space features, the conflict index matrix, and the alarm flag.
[0157] The specific implementation is as follows:
[0158] Receive the conflict detection results output in Step 3. The conflict detection results include:
[0159] Unified semantic space features F aligned , conflict index matrix C ∈ R H×W , alarm flag Alert ∈ {True, False}.
[0160] Based on the unified semantic space features F aligned and the conflict index matrix C ∈ R H×W Iteratively correct the generated image. The correction formula is: O (t+1) = UNet[O (t) , F aligned ⊙ (1 + λC)]; where O (t) is the image generated in the t-th iteration, λ is the configurable conflict correction strength, and ⊙ is element-wise multiplication.
[0161] Example:
[0162] Unified semantic space features F aligned = [0.12, -0.23,..., 0.45];
[0163] Conflict index matrix
[0164] Alarm flag Alert = True;
[0165] The conflict correction strength λ is configured to 0.5.
[0166] Iteration process:
[0167] Initial image: O (0) = Stable Diffusion(F aligned );
[0168] The first correction: O(1) = UNet(O (0) F aligned ⊙(1 + ×C));
[0169] Second Amendment: O (2) = UNet(O (1) F aligned ⊙(1 + ×C));
[0170] Third Amendment: O (3) = UNet(O (2) F aligned ⊙(1 + ×C));
[0171] Initial image O (0) : Generate a sketch of a static horseshoe;
[0172] Corrected image O (3) : Generate a sketch of a dynamic horseshoe.
[0173] Overlay the annotation layer of the conflict area on the generated image to mark the conflict area in the image, and make these areas more prominent visually by overlaying a semi-transparent layer. In this way, users can intuitively identify possible problems or inconsistencies in the image, thereby enhancing the interpretability of the image.
[0174] By overlaying the annotation layer of the conflict area, possible problems or inconsistencies in the image can be intuitively identified. This visualization method enhances the interpretability of the image and helps users better understand the semantic content of the generated image. According to the dynamic threshold and the conflict correction intensity λ, the correction strategy can be flexibly adjusted to ensure the optimal performance of the generated image in different scenarios. The adaptability and robustness of the scheme are improved.
[0175] The generation of the contradiction annotation layer is achieved through the following formula:
[0176] O final = O (t) + δ·Overlay(M mask ), where O (t) is the image generated in the t-th iteration and is the basis of the generated image. M mask is the conflict mask obtained in step three and is a binary matrix used to mark the conflict area in the image. δ is a configurable transparency factor, and Overlay(M mask ) represents the operation of overlaying the conflict area identification M mask onto the image, and usually uses a semi-transparent color (such as red) to mark the conflict area.
[0177] Example: Set: transparency factor δ = ±0.3, t = 3.
[0178] Through the conflict detection in Step 3, generate Mark the conflict areas in the image, where 1 represents the conflict area and 0 represents the non-conflict area.
[0179] Overlay the conflict areas (areas with a value of 1) in M with a red semi-transparent layer mask onto the image O (3) .
[0180] Generate the final image O containing the contradictory annotation layer through the formula O final = O (3) + δ·Overlay(M mask ) final .
[0181] Record the quality evaluation metrics during the generation process, including the conflict index matrix C and the peak signal-to-noise ratio PSNR. Through these metrics, the quality of the generated image can be quantitatively evaluated and data support can be provided for subsequent analysis and optimization.
[0182] The quality evaluation metrics of the semantic calibration log are implemented through the following formula:
[0183]
[0184] The conflict index matrix C reflects the degree of inconsistency between the generated image and the target semantics. The larger the value of C, the higher the degree of inconsistency. In the formula, C is part of the denominator, indicating that the higher the conflict index, the lower the quality score Score.
[0185] The peak signal-to-noise ratio PSNR is used to evaluate the quality of the generated image O (m) , where MAX I is the maximum value of the image pixels and MSE is the mean square error. The peak signal-to-noise ratio PSNR is used to evaluate the quality of the generated image. The larger the value of PSNR, the higher the image quality. In the formula, PSNR is part of the numerator, indicating that the higher the image quality, the higher the quality score Score.
[0186] ∈ is a small constant to prevent division by zero, and its value is ∈ = 1×10 -5 . When the value of the conflict index matrix C approaches 0, ∈ can prevent the denominator from being 0 and ensure the stability of the formula.
[0187] Example 2:
[0188] This example provides an AI-based interactive painting generation system that implements the AI-based interactive painting generation method mentioned in Example 1, including:
[0189] The data acquisition and preprocessing module is responsible for collecting raw multi-modal data, performing preliminary processing, and generating a formatted feature dataset; using natural language processing (NLP) technology, the BERT model is used to extract the semantic features of text; the Whisper model is used to convert speech into text, and then the BiLSTM and Prosody are used to extract speech features; the graph convolutional network (GCN) is used to extract the features of sketches; the multi-modal data is aligned by timestamp to ensure that the synchronization accuracy is less than 50ms, and the cosine similarity is used to ensure semantic consistency; ensuring the integrity and consistency of multi-modal data, providing high-quality input for subsequent feature alignment and fusion.
[0190] The multi-modal feature alignment and fusion module aligns and weighted fuses multi-modal features to generate unified semantic space features; uses an extended CLIP architecture for cross-modal feature alignment to ensure the consistency of different modal data in the unified semantic space; calculates the weights of different modalities through a multi-layer perceptron (MLP), and performs weighted fusion according to the importance of the modalities; performs weighted fusion on multi-modal features and performs layer normalization processing to generate unified semantic space features; makes full use of the complementary information of different modalities to improve the semantic accuracy and detail richness of the generated images.
[0191] The conflict detection and alert module detects the semantic inconsistency between the image and the text description, generates a conflict index matrix, and triggers an alert according to a dynamic threshold; through the visual-text attention mechanism, locates the regions in the image that are inconsistent with the text description, and generates a conflict index matrix; sets a dynamic threshold according to the overall conflict degree of the image and user requirements to determine whether to trigger an alert; if the conflict index of a certain region exceeds the dynamic threshold, an alert mark is generated to prompt the user to make corrections; accurately locates and quantifies the semantic inconsistency between the image and the text description, timely corrects the error regions, and improves the quality of the generated images.
[0192] The image generation and iterative correction module generates and iteratively corrects images based on the unified semantic space features, conflict index matrix, and alert marks; uses a generative adversarial network (GAN) or a diffusion model (Diffusion Model) to generate an initial image; based on the unified semantic space features, conflict index matrix, and alert marks, uses a UNet model to correct the image. The annotation layer of the conflict region is superimposed on the generated image, and the conflict region is highlighted through a semi-transparent layer. Through iterative correction, the generated image better conforms to the user's intention, improving the image quality and user satisfaction.
[0193] The system control and scheduling module is responsible for the scheduling and resource management of the entire system, ensuring the coordinated operation of each module. According to user requirements and data flow, it dynamically allocates computing resources to ensure the efficient operation of each module. It monitors the system resource usage, optimizes resource allocation, and improves the overall system performance. Through message queues and event-driven mechanisms, it ensures the smooth data flow and task flow between each module, achieving the efficient coordination of the system. It improves the stability and response speed of the system, ensuring the fluency of the interactive painting generation process and the user experience.
[0194] Application scenarios:
[0195] Artists can generate images that conform to their creative intentions through voice, text, or sketch input, and continuously improve their works through iterative correction.
[0196] Students can generate images through multimodal input to assist in learning and understanding complex concepts.
[0197] Designers can quickly generate design drafts through sketches or text descriptions, and optimize design details through system correction.
[0198] Example 3:
[0199] This example provides an interactive painting generation device based on artificial intelligence, which applies the interactive painting generation method based on artificial intelligence mentioned in Example 1, including a processor and a memory;
[0200] Among them, the processor can be a central processing unit (CPU) or other forms of processing units with data processing capabilities and / or instruction execution capabilities;
[0201] The memory can be a computer program product, and the computer program product includes various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disks, flash memory, etc.
[0202] In order to implement the steps described in the above "Interactive Painting Generation Method Based on Artificial Intelligence", one or more computer program modules can be stored on the above computer-readable storage medium, and the processor can run the interactive painting generation method based on artificial intelligence mentioned in Example 1.
[0203] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0204] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0205] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0206] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0207] The technical scope of the present invention is not limited to the content described above. Those skilled in the art can make various deformations and modifications to the above embodiments without departing from the technical idea of the present invention, and these deformations and modifications should all fall within the protection scope of the present invention.
Claims
1. An interactive painting generation method based on artificial intelligence, characterized in that: It includes the following steps: Obtain the original multimodal data, and after processing the original multimodal data, obtain a formatted feature dataset containing multimodal features; the original multimodal data includes the collected original text data, original speech data, and original sketch data; Align the multimodal features in the formatted feature dataset, assign weights according to the importance of different modalities, and perform normalization processing after weighted fusion of the multimodal features to obtain unified semantic space features; Locate the regions where the image and text descriptions are inconsistent through a visual-text attention mechanism based on the unified semantic space features, quantify the semantic inconsistency degree between the image and text descriptions to obtain a conflict index matrix, and determine whether to trigger an alarm mark according to a dynamic threshold; Iteratively correct and generate an image based on the unified semantic space features, conflict index matrix, and alarm mark.
2. The interactive painting generation method based on artificial intelligence according to claim 1, wherein: The obtaining of the formatted feature dataset after processing the original multimodal data includes: Text feature extraction: Semantic encoding is performed using the BERT-base model, i.e., T = BERT base (t) ∈ R 768 ; where t is the string of the original text input, and BERT base is a 12-layer Transformer encoder, and T is the output 768-dimensional semantic vector; Speech feature extraction: Convert speech to text: S text = Whisper(s audio ); where s audio is the speech waveform data, Whisper is the speech-to-text model, and S text is the output result of speech-to-text; Multi-feature fusion: Among them, BiLSTM is a bidirectional LSTM network that outputs a 256-dimensional vector; Prosody is a prosody feature extractor that outputs a 256-dimensional vector; is a vector concatenation operation, and S is the output 512-dimensional speech fusion vector; Sketch feature extraction: Stroke encoding based on the Graph Convolutional Network (GCN), i.e., I = GCN(Stroke2Graph(I raw )) ∈ R 256 ; where I raw represents the strokes drawn by the user through the digitizing tablet, which consists of N coordinate points, denoted as represents the digitizing tablet coordinate sequence; Stroke2Graph is a converter from strokes to graph structures, GCN is a 3-layer graph convolutional network with a hidden layer dimension of 128, and the finally output I is a 256-dimensional sketch vector.
3. The interactive painting generation method based on artificial intelligence according to claim 1, characterized in that: The obtaining of the formatted feature dataset after processing the original multimodal data further includes: Align the data of three modalities, namely text, speech, and sketch, in time through the timestamps of the hardware on the input device: D = Align(T, S, I), and it satisfies |t text - t sketch | < 50 ms; where t text is the start time of text input, and t sketch is the start time of sketch input.
4. The interactive painting generation method based on artificial intelligence according to claim 2, wherein: The process of aligning the multimodal features in the formatted feature dataset is: Adopt an extended CLIP architecture for feature alignment: Among them, Q = W q T, where T is a 768-dimensional semantic vector of text features, and Q is a 512-dimensional query vector after projection; K = W k [S; I], where [S; I] is the concatenated vector of the speech feature S and the sketch feature I, with a dimension of 768, and K is the projected key vector with a dimension of 512; V = W v [S; I], where V is a 512-dimensional value vector after projection; W * is a learnable parameter matrix that projects features of different modalities into a unified 512-dimensional semantic space, including the query projection matrix W q , the key projection matrix W k , and the value projection matrix W v ; W q ∈R 512×768 Specifically, it is a text feature projection matrix that projects 768-dimensional text features onto 512 dimensions; W k ∈R 512×768 Specifically, they are the voice and sketch feature projection matrices, which project 768-dimensional text features onto 512 dimensions; W v ∈R 512×768 Specifically, it is the voice and sketch feature projection matrix that projects the 768-dimensional text features onto 512 dimensions; d is a scaling factor, CrossModalEncoder is an encoder for cross-modal feature alignment, and F is cross-modal aligned features.
5. The interactive painting generation method based on artificial intelligence according to claim 4, wherein: The assigning of weights according to the importance of different modalities includes: Calculate the modality importance through MLP: α, β, γ = MLP(F), and α + β + γ = 1; where MLP is a multi-layer perceptron for calculating modality weights; α, β, γ are the text weight, speech weight, and sketch weight in sequence; Calculate the fused feature according to the assigned weights: F fused = αT + βS + γI; The obtaining of the unified semantic space features after weighted fusion of the multimodal features and then performing normalization processing is: F aligned = LayerNorm(F fused ) ∈ R 1024 , where LayerNorm is the layer normalization operation and F aligned is the feature in the unified semantic space.
6. The method for generating an interactive painting based on artificial intelligence according to claim 5, characterized in that: The specific steps of step three are: Receive the unified semantic space feature F output in Step 2 aligned as the input for conflict detection; Calculate the image feature F img and the text feature F text with the attention weight: Among them, AttnMap is the attention mapping function, which calculates the similarity between the image and the text features; n is the total number of feature segments, and H is the attention heat map, representing the similarity between the image and the text description. is the i-th segment of the image feature F img is the i-th segment of the text feature F text Based on the empirical thresholds τ low and τ high , identify the conflict regions: M mask = II(H < τ low ∪ H > τ high ), where τ low is the low similarity threshold, τ high is the high similarity threshold, II is the indicator function, which satisfies the condition of 1, otherwise 0; M mask is the conflict mask; Calculate the image feature F img and the text feature F text for cosine similarity, and convert it into a conflict index C: where n is the number of feature dimensions is the k-th dimension of the image feature F img ; is the k-th dimension of the text feature F text ; Judge whether to trigger an alarm according to the modality divergence threshold η:
7. The method for generating an interactive painting based on artificial intelligence according to claim 6, wherein: The iteratively correcting and generating an image based on the unified semantic space features, conflict index matrix, and alarm mark includes: Receive the conflict detection result output in step three, where the conflict detection result includes the unified semantic space feature F aligned , the conflict index matrix C ∈ R H×W , and the alert flag Alert ∈ {True, False}; Based on the unified semantic space feature F aligned and the conflict index matrix C ∈ R H×W Iteratively correct to generate an image, and the correction formula is: O (t+1) = UNet[O (t) , F aligned ⊙ (1 + λC)]; where, O (t) is the image generated in the t-th iteration, λ is the configurable conflict correction intensity, and ⊙ is the element-wise multiplication.
8. The method for generating an interactive painting based on artificial intelligence according to claim 7, wherein: The iteratively correcting and generating an image based on the unified semantic space features, conflict index matrix, and alarm mark further includes: Overlay an annotation layer of the conflict region on the generated image, and the annotation formula is: O final = O (t) + δ · Overlay(M mask ), where O (t) is the image generated in the t-th iteration, M mask is the conflict mask obtained in step three, and δ is a configurable transparency factor; Record the quality evaluation index during the generation process, and the quality evaluation formula is: where C is the conflict index matrix, PSNR is the peak signal-to-noise ratio, and ∈ is a small constant to prevent division by zero, with a value of ∈ = 1×10 -5 .
9. An AI-based interactive painting generation system, characterized in that: Implement the artificial intelligence-based interactive painting generation method as shown in any one of claims 1 to 8, including: A data collection and preprocessing module, responsible for collecting the original multimodal data and performing preliminary processing to generate a formatted feature dataset; A multimodal feature alignment and fusion module, which aligns and weighted fuses the multimodal features to generate unified semantic space features; A conflict detection and alarm module, which detects the semantic inconsistency between the image and text descriptions, generates a conflict index matrix, and triggers an alarm according to a dynamic threshold; An image generation and iterative correction module, which generates and iteratively corrects an image based on the unified semantic space features, conflict index matrix, and alarm mark; A system control and scheduling module, responsible for the scheduling and resource management of the entire system to ensure the collaborative work of each module.
10. An interactive painting generation device based on artificial intelligence, characterized in that: It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, which implements the artificial intelligence-based interactive painting generation method as shown in any one of claims 1 to 8 when the processor executes the program.
Citation Information
Cited By
Multi-modal content compliance auditing method and system
CN120611053A
3D art model intelligent generation system and method based on multi-source data fusion
CN120997449A
Man-machine collaborative creation method and system based on multi-modal fusion
CN121030691A
Man-machine collaborative creation method and system based on multi-modal fusion
CN121030691B
Government affair intelligent writing method and system based on multi-Agent collaboration and long memory technology
CN121327156A