Visual harmonious color matching generation method based on image attention mechanism and knowledge base
Through the large language model, the prompt words are expanded and combined with the image attention mechanism and knowledge base are generated to generate a harmonious visual color scheme, which solves the novelty and harmony problems of color palette generation in the existing technology, and realizes the efficient generation of adaptive visual design.
Patent Information
- Application Number
- CN202510468144.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-25
AI Technical Summary
Existing color generation methods lack novelty when generating color palettes, difficult to ensure harmonious beauty, and difficult for users to find suitable reference images, resulting in limitations in the adaptability and flexibility of the generated color palette.
The prompt words entered by users are extended through a large language model, combined with the image attention mechanism and knowledge base, and automatically generate a harmonious and consistent color scheme that meets user needs, including image retrieval, color extraction and optimization steps to ensure that the generated palette matches the visual scene.
The generated color palette not only adapts to different visual design needs, but also finds a balance between ensuring data readability and aesthetic coordination, improving generation efficiency and accuracy.
Smart Images

Figure CN120372002A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of color matching generation, and in particular to a method for generating visual harmonious color matching based on an image attention mechanism and a knowledge base. Background Art
[0002] Color plays a crucial role in data visualization and artistic creation. A well-designed color matching scheme can not only enhance the readability of data but also improve visual attractiveness. However, designing an effective color matching scheme is a time-consuming and challenging process for both ordinary users and professional designers. Existing color generation methods still face many challenges in practical applications. Color generation methods based on text descriptions usually rely on predefined color combinations or rule-based algorithms, resulting in a lack of novelty in the generated color palettes, with relatively single and predictable results. Although color matching scheme generation algorithms based on images can provide rich visual references, in actual use, it is often difficult for users to find suitable reference images. Especially in the case of a lack of clear design concepts, screening suitable pictures is not only time-consuming and laborious but also difficult to ensure the accuracy and aesthetic coordination of the final color matching. In addition, traditional automatic color generation methods usually focus on the distinguishability between colors while ignoring the overall harmonious beauty, making it difficult to meet the requirements of complex data visualization scenarios. At the same time, these methods lack a deep understanding of user intentions, resulting in limitations in the adaptability and flexibility of automatically generated color palettes. In view of this, the present invention proposes a method for generating a color palette based on an image attention mechanism and a knowledge base, which automatically expands the prompt words input by the user and combines the image attention mechanism to generate a color palette that meets the user's needs. Summary of the Invention
[0003] The purpose of the present invention is to propose a method for generating visual harmonious color matching based on an image attention mechanism and a knowledge base to solve the problems raised in the background art. The present invention automatically provides a high-quality color matching scheme for visual charts through simple text input.
[0004] To achieve the above object, the technical solution adopted by the present invention is as follows:
[0005] A method for generating visual harmonious color matching based on an image attention mechanism and a knowledge base automatically generates a harmonious and user-demand-compliant color matching scheme by expanding the prompt words input by the user through a large language model and optimizing in combination with the image attention mechanism and the knowledge base, including the following steps:
[0006] S1. Input the initial prompt word T as the input for generating the color matching scheme;
[0007] S2. Use the large language model to expand the prompt word T to generate a more detailed description P;
[0008] S3. Retrieve the reference image I from the image database D according to the extended description P r ;
[0009] S4. Based on the extended description P and the retrieved reference image I r , use the generation model G to generate the final target image I c ;
[0010] S5. Using the attention mechanism, extract the color region R related to the description P from the generated image I c . The attention mechanism calculates the attention map M between the image I c and the description P, identifies the region in the image that is most relevant to the text description, and upsamples the attention map to extract the color region R;
[0011] S6. Extract colors from the color region R through a clustering algorithm. During the clustering process, combine the attention scores and the hue values of the colors, and preferentially select the colors in the high-attention regions as representative colors to generate a candidate color set;
[0012] S7. Use the knowledge base to optimize the generated candidate color set, combine color theory, color psychology, and color matching rules, recommend background colors that are harmonious with the color palette, and provide the basis for the recommended background colors;
[0013] S8. Combining the visual context, use the large language model to intelligently perform color allocation by combining the attention scores, visual chart information, and initial prompt words to generate the final visual color matching scheme.
[0014] Preferably, the extension process in S2 follows a Markov decision process. Based on the initial prompt word T, at each subsequent step t, based on the current description P t make a decision d t ; The decision involves multiple key dimensions, including theme, artistic style, object, perspective and composition, and emotion and narrative depth; The large language model executes the decision-making process internally, dynamically generates d t , and gives the current description state P t to iteratively update the prompt word until the prompt word reaches the final state that is complete and meets the requirements.
[0015] Preferably, the retrieval process in S3 can be formally described as:
[0016]
[0017] where D represents the image database; I rIndicates the reference image that best matches the semantics of the extended description P; Sim(P, I) is measured by calculating the cosine similarity between the embedding vectors of the text description P and the image I, and the image I with the highest similarity is selected. r As the reference image, it provides intuitive color guidance for the image generation process.
[0018] Preferably, the generation process described in S4 is expressed as:
[0019] I c = G(P, I r )
[0020] The generation model G uses a diffusion model to generate high-quality images that conform to the description P by gradually denoising, while ensuring that its color features are coordinated with I r Keep in harmony.
[0021] Preferably, the specific content of S5 includes the following:
[0022] S5.1. Calculate the attention map M of the image I c and the description P to identify the text-related regions in the image. The attention process is expressed as:
[0023]
[0024] Attention = M · V
[0025] where Q = W Q (T), K = W K (T), V = W V (I); W Q , W K , W V represent the learned projection matrices; d is the dimension of K;
[0026] S5.2. Upsample the attention map M to match the resolution of the input image I c ; The upsampling process is expressed as:
[0027] R i = Upsample(M i ) × I c
[0028] where M i represents the feature map of the i-th layer in the attention map, used to represent the feature response intensity of that layer; I c is the generated image; R represents the color extraction region, that is, the part of the image I c that is most relevant to the prompt word P.
[0029] Preferably, the specific content of S6 includes the following:
[0030] S6.1. Map the attention scores of color regions and the hue values of colors to the polar coordinate system; in the polar coordinate system, map the colors with high attention scores to the outer edge and distribute the colors in the regions with low attention scores at the center.
[0031] S6.2. Perform fine-grained clustering in the high-attention region. Use the K-means clustering algorithm to cluster the colors according to the Euclidean distance, and preferentially select the color with the highest attention score in each cluster center as the representative color.
[0032] Preferably, the S7 specifically includes the following content:
[0033] S7.1. Utilize the information on color theory, color psychology, and color matching rules in the knowledge base K to retrieve the most relevant color theory K s , as the basis for optimizing the color palette; the knowledge base retrieval process is expressed as:
[0034]
[0035] where Sim(C, K i ) calculates the cosine similarity between the color set C and the color matching rule K stored in the knowledge base K i , and selects the most relevant color theory K s as the optimization basis;
[0036] S7.2. Combine the visualization type V specified by the user to adjust the palette colors C to meet specific visualization requirements; the optimization process is expressed as:
[0037] C’ = Optimize(C, K s , V)
[0038] S7.3. Recommend a harmonious background color B based on the optimized color palette C’, and explain the rationality of this color selection through the large language model in combination with the color theory K s .
[0039] Compared with the prior art, the present invention provides a visualization harmonious color matching generation method based on an image attention mechanism and a knowledge base, having the following beneficial effects:
[0040] Compared with the prior art, the present invention introduces an image attention mechanism and knowledge base optimization in the process of generating a color scheme, enabling the generated color scheme to not only adapt to different visualization design requirements but also enhance aesthetic coordination while ensuring data readability. By expanding the input text, this method can fully explore the user's color matching intention, and after retrieving the most relevant reference images, provide constraints on harmonious colors for the generation model, thus making the overall color of the generated picture more harmonious. Subsequently, the present invention uses the attention mechanism to accurately extract color regions highly relevant to the text description and improves the accuracy and rationality of color selection through a clustering algorithm. Combining the color theory of the knowledge base and visualization context information, the present invention can intelligently optimize the color scheme and recommend background colors that meet specific visualization scenarios, ensuring that the final scheme not only meets the data expression requirements but also has high visual appeal. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings involved in the embodiments are briefly introduced below. Obviously, the accompanying drawings in the following description are only schematic illustrations of some embodiments of the present invention, and those skilled in the art can construct other forms of drawings based on these drawings without creative efforts.
[0042] Figure 1 Schematic diagram of the process of the present invention;
[0043] Figure 2 Schematic diagram of the implementation process;
[0044] Figure 3 Image matched during the image retrieval process;
[0045] Figure 4 Image generated during the image generation process;
[0046] Figure 5 Attention feature map extracted during the color extraction process;
[0047] Figure 6 Result diagram of palette optimization and allocation. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present invention.
[0049] Please refer to the attached Figure 1 , and the present invention generates a visual harmony color scheme according to the following steps:
[0050] Step 1: Input the prompt T;
[0051] Step 2: Use a large language model to expand the prompt T to generate a more detailed description P;
[0052] Step 3: Retrieve the reference image I from the image database according to the expanded description P r ;
[0053] Step 4: Based on the expanded description P and the reference image I r , use a generative model to generate a specific image I c ;
[0054] Step 5: Use the attention mechanism to extract the color region R related to the description P from the generated image I c ;
[0055] Step 6: Extract colors from the relevant region R through a clustering algorithm;
[0056] Step 7: Use the knowledge base to optimize the color palette, recommend background colors harmonious with the color palette, and provide reasons for the recommended background colors;
[0057] Step 8: Combine the visualization context to generate a visualization color scheme for the visualization chart.
[0058] The present invention will be further described in detail below by taking the generation of a visualization color scheme from a certain user input as an example.
[0059] Example 1:
[0060] Refer to the appendix Figure 1 , the present invention proposes a visualization harmonious color matching generation method based on an image attention mechanism and a knowledge base, and the visualization harmonious color matching is generated according to the following steps:
[0061] S1. Input the initial prompt T as the input for generating the color palette.
[0062] S2. Use a large language model to expand the prompt T to generate a more detailed description P. The prompt expansion process follows a Markov decision process, which starts from the initial prompt T provided by the user. In each subsequent step t, based on the current description P t make a decision d t to enrich or modify the prompt. This decision design involves multiple key dimensions, including theme, artistic style, object, perspective, composition, and emotional and narrative depth. The large language model executes this decision-making process internally, dynamically generates d t , and based on the current description P t iteratively updates the prompt to generate a new prompt P t+1 . Each decision d tOptimally selected by the large language model through its extensive knowledge base and creativity to ensure that the prompt gradually becomes complete, precise, and meets the user's needs. This process can be formally described as a sequence:
[0063] {P0,d1,P1,d2,…,P T-1 ,d T}
[0064] where T represents the total number of iterations.
[0065] S3. Retrieve the reference image I from the image database D according to the extended description P r . The retrieval process can be formally described as:
[0066] I r = argmax Sim(P, I)
[0067] I ∈ D
[0068] where D represents the image database; I r is the reference image that best matches the semantics of the extended description P. Sim(P, I) is measured by calculating the cosine similarity between the embedding vectors of the text description P and the image I, and the image I r with the highest similarity is selected as the reference image. This reference image not only provides an intuitive color guide for the subsequent image generation process in S4 but also ensures that the generated color palette is more optimized in terms of color coordination and diversity, thus improving the quality of the final color scheme.
[0069] S4. Based on the extended description P and the retrieved reference image I r , use the generation model G to generate the final target image I c . The generation process can be expressed as:
[0070] I c = G(P, I r )
[0071] where G is the generation model; P is the extended description; I r is the reference image. The generation model uses a diffusion model to generate high-quality images that conform to the description P by gradually denoising, while ensuring that its color characteristics are coordinated with I r to enhance the overall harmony and beauty of the color scheme.
[0072] S5. Use the attention mechanism to extract the color region R related to the description P from the generated image I c . The attention mechanism identifies the text-related regions in the image by calculating the attention map M between the image I c and the description P. The attention process can be expressed as:
[0073]
[0074] Attention = M·V
[0075] where Q = W Q (T), K = W K (T), V = W V (I); W Q , W K , W V denotes the learned projection matrix, and d is the dimension of K. Subsequently, the attention map M is processed to extract the region most relevant to the prompt. Specifically, M i represents the feature map of the i-th layer in the attention map, which is used to represent the feature response intensity of that i-th layer. To match the resolution of the input image I c , the attention map M i is upsampled. The upsampling process can be expressed as:
[0076] R i = Upsample(M i )×I c
[0077] Finally, R represents the color extraction region, that is, the part of the image I c most relevant to the prompt P.
[0078] S6. Select the color extraction region R in S5 and perform color extraction through a clustering algorithm. During the color clustering process, classification is not only based on the characteristics of the color itself, but also the attention scores are combined to improve the accuracy of color extraction. Specifically, first calculate the attention scores of each color region in the image and map them together with the hue values of the colors into the polar coordinate system. In the polar coordinate system, the colors with higher attention scores are mapped to the outer edge, while the colors in the low-attention regions are distributed in the center. Based on this geometric structure, finer-grained clustering is performed in the high-attention regions, and the K-means algorithm is used to group the colors according to the Euclidean distance, so that the colors with similar hues and similar attention scores are grouped into one class. At the same time, the color with the highest attention score in each cluster center is preferentially selected as the representative color, ensuring that the color palette not only has color diversity, but also makes the generated color scheme more in line with the prompt and focus of interest input by the user.
[0079] S7. Use the knowledge base to optimize the color palette, recommend a background color harmonious with the color palette, and provide the reasons for the recommended background color. First, utilize the information on color theory, color psychology, and color matching rules in the knowledge base K to retrieve the most relevant color theory K s , which serves as the basis for optimizing the color palette. The knowledge base retrieval process can be expressed as:
[0080]
[0081] Among them, Sim(C, K i ) calculates the cosine similarity between the color set C and the color matching rules K stored in the knowledge base K i and selects the most relevant color theory K s as the optimization basis. Then, in combination with the visualization type V specified by the user, the palette colors C are adjusted to meet specific visualization requirements. The optimization process can be expressed as:
[0082] C’ = Optimize(C, K s , V)
[0083] Finally, a harmonious background color B is recommended based on the optimized palette C’, and the rationality of this color selection is explained through the large language model in combination with the color theory K s .
[0084] S8. Combining the visualization context, using the large language model in combination with the attention score, visualization chart information, and initial prompt words, intelligently perform color allocation to generate the final visualization color matching scheme. During the color matching allocation process, the large language model will comprehensively consider the attention score in S6 and the optimized palette C’ in S7 to ensure that the colors in the high-attention areas are preferentially allocated. At the same time, in combination with the initial prompt words and the structure and hierarchical information of the visualization chart, ensure that the color matching scheme can accurately express the user's intention, highlight key information, maintain the overall harmonious beauty, thereby improving the visualization effect and user experience.
[0085] Example 2:
[0086] Please refer to Figure 2 , which is based on Example 1 but is different in that the visualization harmonious color matching generation method proposed by the present invention is described below in combination with specific examples. The steps are as follows:
[0087] Step 1: Input the initial prompt word "autumn". The initially input prompt word is concise and describes the theme direction of the color matching scheme that the user hopes to generate.
[0088] Step 2: Use the large language model to expand the prompt word "autumn". The detailed description after expansion is output as "dark red maple leaves, golden wheat ears, sunset-like hues".
[0089] Step 3: According to the keywords "dark red", "golden", and "sunset-like hues" included in the expanded prompt word, retrieve semantically consistent images in the image database, and the matching results are as follows Figure 3As shown. The main colors in this image are warm red and golden yellow, and the light atmosphere is soft, which is highly consistent with the description of the prompt in terms of color mood and light and shadow atmosphere.
[0090] Step 4: Based on the extended description and the reference image, generate the final target image through a diffusion model. As Figure 4 shown, the generated image is a typical autumn natural scene, depicting a large number of red and golden maple leaves, contrasting with the golden wheat fields and soft sunlight in the distance. The overall color tone is highly consistent with "dark red maple leaves, golden wheat ears, and a sunset-like color tone", further enhancing the emotional atmosphere and visual guidance of the color scheme.
[0091] Step 5: Use the attention mechanism to extract the color regions in the generated image that are most relevant to the extended description. Through the attention feature map output by the attention mechanism (as Figure 5 shown), the extraction regions are concentrated in the areas of the leaves and the fallen leaves on the ground, and these areas are bright in color, showing orange, golden yellow, and red.
[0092] Step 6: Use the knowledge base to optimize the extracted color set. The knowledge base recommends that the background color coordinated with the color palette is warm beige, because beige belongs to the warm color system in color psychology with golden yellow, orange, and red, and can effectively improve the comfort and visual attraction of the overall color scheme.
[0093] Step 7: Combine the visual context and intelligently perform color allocation to finally generate a harmonious color scheme suitable for data charts (as Figure 6 shown). The final output color scheme includes: red, dark brown, golden yellow, gray, off-white, and the recommended background color, ensuring that the overall visual effect not only highlights key information but also meets the requirements of aesthetic coordination, achieving the purpose of being clear, harmonious, and attracting users' attention.
[0094] The above is only a further explanation of the present invention and is not intended to limit this patent. Equivalent implementations without departing from the spirit and scope of the inventive concept of the present invention shall be included within the scope of the claims of this patent.
Claims
1. A method for generating visual harmonious color matching based on an image attention mechanism and a knowledge base, characterized in that Expand the prompt of the user input through a large language model, and combine image attention mechanism and knowledge base optimization to automatically generate a harmonious color scheme that meets the user's needs, including the following steps: S1. Input the initial prompt T as the input for color scheme generation; S2. Use the large language model to expand the prompt T to generate a more detailed description P; S3. Retrieve the reference image I from the image database D according to the extended description P r ; S4. Based on the extended description P and the retrieved reference image I r , use the generative model G to generate the final target image I c ; S5. Using the attention mechanism, extract the color region R related to the description P from the generated image I c Among them, the attention mechanism calculates the attention map M between the image I c And the description P, identify the region in the image that is most relevant to the text description, and upsample the attention map to extract the color region R; S6. Extract colors from the color region R through a clustering algorithm. During the clustering process, combine the attention score and the hue value of the color, and preferentially select the colors in the high-attention region as representative colors to generate a candidate color set; S7. Optimize the generated candidate color set using the knowledge base, combine color theory, color psychology and color matching rules, recommend background colors that are harmonious with the color palette, and provide the basis for the recommended background colors; S8. Combine the visual context, and use the large language model to intelligently allocate colors by combining the attention score, visual chart information and the initial prompt to generate the final visual color scheme.
2. The visualization harmonious color matching generation method based on an image attention mechanism and a knowledge base according to claim 1, wherein The expansion process described in S2 follows a Markov decision process. Based on the initial prompt T, at each subsequent step t, based on the current description P t a decision d is made t ; the decision involves multiple key dimensions, including theme, artistic style, object, perspective and composition, and emotion and narrative depth; the large language model executes the decision-making process internally and dynamically generates d t and gives the current description state P t The prompt is iteratively updated until the prompt reaches a complete and satisfactory final state.
3. The visualization harmonious color matching generation method based on an image attention mechanism and a knowledge base according to claim 1, wherein The retrieval process described in S3 can be formally described as: Among them, D represents an image database; I r represents the reference image that best matches the semantics of the extended description P; Sim(P, I) is measured by calculating the cosine similarity between the embedding vectors of the text description P and the image I, and the image I with the highest similarity r is selected as the reference image to provide intuitive color guidance for the image generation process.
4. The visualization harmonious color matching generation method based on an image attention mechanism and a knowledge base according to claim 1, characterized in that The generation process described in S4 is expressed as: I c = G(P, I r ) The generative model G uses a diffusion model to generate high-quality images that conform to the description P by gradually denoising, while ensuring that its color features are coordinated with I r Keep coordinated.
5. The visualization harmonious color matching generation method based on an image attention mechanism and a knowledge base according to claim 1, characterized in that, The specific content of S5 includes the following: S5.
1. Calculate image I c and the attention map M of the description P to identify the text-related regions in the image. The attention process is expressed as: Attention = M·V where Q = W Q (T), K = W K (T), V = W V (I); W Q , W K , W V denotes the learned projection matrix; d is the dimension of K; S5.
2. Upsample the attention map M to match the resolution of the input image I c The upsampling process is expressed as: R i = Upsample(M i ) × I c Among them, M i represents the feature map of the i-th layer in the attention map, which is used to represent the feature response intensity of that layer; I c is the generated image; R represents the color extraction area, that is, the part of the image I c that is most relevant to the prompt word P.
6. The visualization-based harmonious color matching generation method based on an image attention mechanism and a knowledge base according to claim 1, characterized in that The specific content of S6 includes the following: S6.
1. Map the attention score of the color region and the hue value of the color to the polar coordinate system; in the polar coordinate system, map the colors with high attention scores to the outer edge, and distribute the colors in the region with low attention scores in the center; S6.
2. Perform fine-grained clustering in the high-attention region, and use the K-means clustering algorithm to cluster the colors according to the Euclidean distance. Preferentially select the color with the highest attention score in each cluster center as the representative color.
7. The visualization harmonious color matching generation method based on an image attention mechanism and a knowledge base according to claim 1, characterized in that The specific content of S7 includes the following: S7.
1. Retrieve the most relevant color theory K from the information about color theory, color psychology, and color matching rules in knowledge base K s , as the basis for optimizing the color palette; the knowledge base retrieval process is expressed as: Among them, Sim(C, K i ) calculates the cosine similarity between the color set C and the color matching rules K stored in the knowledge base K i , and selects the most relevant color theory K s as the optimization basis; S7.
2. Combine the specified visualization type V of the user to adjust the palette color C to meet specific visualization requirements; the optimization process is expressed as: C’ = Optimize(C, K s , V) S7.
3. Recommend a harmonious background color B based on the optimized color palette C', and explain the rationality of this color selection through a large language model combined with color theory K s Explain the rationality of this color selection.