Plane design element importance detection method, system and equipment based on weak supervision training and medium

By adopting a weakly supervised training method in graphic design, combining the division of global and local grids and multi-scale sequence prediction model, the problems of high training costs and low generalization capabilities in the existing technology are solved, and efficient and accurate detection of the importance of design elements is achieved.

CN120014348AActive Publication Date: 2025-05-16XIDIAN UNIV +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510094035.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-16
Estimated Expiration
2045-01-21

AI Technical Summary

Technical Problem

The prior art has problems such as high training costs, low generalization ability and limited applicable fields in the detection of the importance of graphic design elements, and it is difficult to effectively capture the relationship and relative importance between design elements.

Method used

Using a weakly supervised training method, the training process is optimized to achieve efficient design element importance detection by introducing the division of global and local grids of graphic design, and the combination of multi-scale sequence prediction models.

Benefits of technology

It reduces the dependence on a large amount of labeled data, reduces training costs, improves generalization ability and application feasibility in the field of graphic design, and can more accurately capture the hierarchical relationships and relative importance of design elements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014348A_ABST
    Figure CN120014348A_ABST
Patent Text Reader

Abstract

The invention discloses a plane design element importance detection method, system and equipment based on weak supervision training and a medium. The method comprises the following steps: manually marking a data set of a global grid sequence; dividing the graphic design into a global grid and a local grid through weak supervision training; obtaining a prediction sequence and a relative weight by using the local sequence prediction model and the global sequence prediction model; the relative weight of the text and the visual feature Vt is obtained through a weight adaptive model; calculating importance indexes of the graphic design elements; the system, the equipment and the medium are used for implementing the method. By introducing a strategy of combining division of a global grid and a local grid of plane design and a multi-scale sequence prediction model, a training process is optimized, and efficient sequence prediction is realized; the importance detection of the design elements can be carried out in an efficient, accurate and low-cost mode, and a more feasible solution is provided for the wide field of plane design.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross-technical field of artificial intelligence and creative media design, and in particular to a method, system, device and medium for detecting the importance of graphic design elements based on weak supervision training. Background Art

[0002] In today's digital age, the integration of graphic design and artificial intelligence has become an important driving force for continuous innovation in the field of creative media design. With the rapid development of artificial intelligence technology, especially the emergence of generative models, designers have gained more inspiration and creative possibilities in graphic design. Artificial intelligence technology can not only assist designers in complex graphic processing, but also provide real-time feedback and suggestions required for creative media design by learning a large number of design cases and trends. In this context, the importance detection of design elements has become a core technology of artificial intelligence in creative media design. The main goal of this technology is to analyze and understand the importance of different design elements through deep learning algorithms, so as to use these elements more intelligently in the graphic design process. There are many types of design elements, including color, shape, text, etc., and their interaction directly affects the overall effect of the design work. Through importance detection, designers can more accurately grasp the position of each element in the overall layout, thereby improving the expressiveness and attractiveness of the design. The importance detection of design elements has a wide range of value in practical applications. First of all, it can improve design efficiency, allowing designers to find and emphasize key design elements more quickly, so as to create more in-depth and attractive works in a limited time. Secondly, it helps to improve the user experience of the design. By reasonably allocating the importance of each element, it makes it easier for the audience to understand the design intention and improve the communication effect. Most importantly, the importance detection of design elements promotes the development of creative media design in a smarter direction that is more in line with human aesthetic needs, and provides designers with more creative possibilities and space. Therefore, the importance detection of design elements is not only a hot topic, but also one of the key technologies to promote the continuous progress of creative media design.

[0003] In recent years, saliency detection of images and graphics has been studied in depth, and a variety of different detection methods have emerged. Common saliency detection methods include convolutional neural network methods based on deep learning, methods based on graph theory, and frequency domain analysis. However, these methods generally have some shortcomings. Some of them use supervised learning, which leads to high cost of training data set annotation, limiting their feasibility in practical applications. In addition, some methods are limited by specific data domains and are difficult to adapt to the broad field of graphic design, and their generalization ability is relatively low. There is a clear difference between importance detection and saliency detection. Saliency detection focuses on identifying eye-catching areas in an image, usually by highlighting features such as color, texture, or edges. However, importance detection pays more attention to the hierarchical relationship and relative importance of graphic design elements, involving the understanding and analysis of the overall design structure. Traditional methods have failed to meet this demand well because they tend to focus on saliency and ignore the relative importance between elements, resulting in unsatisfactory application results in the field of graphic design. Therefore, there is an urgent need for a low-cost and high-generalization method for importance detection of graphic design elements to make up for the shortcomings of existing methods. Such a method should be able to better capture the relationship between design elements and reduce dependence on large amounts of labeled data, while also having sufficient generalization capabilities to adapt to the requirements of different fields and design styles, providing more effective support for the field of graphic design.

[0004] The patent application document with publication number CN109741293A discloses an image saliency detection method and device. The patent application detects and identifies salient areas in an image through supervised learning technology. However, due to the high training cost and the limitation of specific data domain, its generalization ability is relatively low and it is difficult to adapt to a wide range of graphic design fields, which makes it difficult to obtain consistent results in practical applications.

[0005] The patent application document with publication number CN111008558A discloses an image importance detection method based on character relationships. The patent application infers the importance of characters by learning to build relationships between characters and between characters and events in the image. Although this method can be used for importance detection in natural images, its practical application feasibility in the design field is significantly limited due to technical reasons such as high training cost, low generalization ability and limited coverage. Summary of the invention

[0006] In order to overcome the deficiencies of the above-mentioned prior art, the purpose of the present invention is to provide a method, system, device and medium for detecting the importance of graphic design elements based on weakly supervised training. By introducing a strategy combining the division of the global grid and the local grid of graphic design and a multi-scale sequence prediction model, the training process is optimized and efficient sequence prediction is achieved. The importance of design elements can be detected in an efficient, accurate and low-cost manner, providing a more feasible solution for a wide range of graphic design fields.

[0007] In order to achieve the above object, the technical solution adopted by the present invention is:

[0008] A method for detecting the importance of graphic design elements based on weakly supervised training comprises the following steps:

[0009] Step 1: Divide the obtained multiple graphic designs into K×K grids, and manually annotate the design order of each cell content in the grid at the grid level to obtain a dataset with annotations based on the global grid order, including graphic designs, K×K global grids, and global grid order;

[0010] Step 2, using the connectivity of the components, extract the elements of the graphic design in the dataset with global grid order annotation obtained in step 1 to obtain text elements and visual elements;

[0011] Step 3: Build a multi-scale sequence prediction model, including a local sequence prediction model and a global sequence prediction model. The input of the multi-scale sequence prediction model is the plane design in the dataset with global grid sequence annotation obtained in step 1, and the output is the predicted sequence of K×K global grid. S t is a one-hot vector, representing the prediction result at time t;

[0012] Step 4: Use the dataset with annotations based on the global grid sequence obtained in step 1 to train the multi-scale sequence prediction model constructed in step 3. During the training process, the K×K global grids of the dataset with annotations based on the global grid sequence obtained in step 1 are divided into M×M local grids of fixed size, and each local grid is input into the local sequence prediction model to obtain the prediction sequence of each local grid; the prediction sequence of the local grid and the plane design are input into the global sequence prediction model to obtain the prediction sequence of the global grid, and the cross quotient loss is calculated using the prediction sequence of the global grid and the global grid sequence in the dataset with annotations based on the global grid sequence obtained in step 1, and back propagation is performed to update the parameters of the multi-scale sequence prediction model in step 3, during which the visual feature V of each position is obtained. t and text features T t The relative weight of

[0013] Step 5: Use the text elements and visual elements obtained in step 2 and the global grid obtained in step 4 to predict the visual features V at each position in the sequence result. t and text features T t The relative weights of the elements are used to calculate the element importance index of the graphic design and obtain the final saliency map.

[0014] The specific method of step 2 includes:

[0015] Step 2-1, extracting initial elements from the plane design in the dataset with global grid order annotation obtained in step 1 using the connectivity of the components;

[0016] Step 2-2, judging the initial elements extracted in step 2-1, if the overlapping ratio of the bounding boxes of the elements exceeds a preset value α, merging the adjacent elements to obtain a new element set;

[0017] Step 2-3, using optical character recognition (OCR) to detect all elements in the new element set in step 2-2, if there is text, it is marked as a text element, otherwise it is a visual element.

[0018] The specific method of step 3 includes:

[0019] Step 3-1, build a local sequence prediction model; the local sequence prediction model includes an encoder and a long short-term memory network LSTM. The input is the K×K global grids in the dataset with global grid sequence annotations obtained in step 1, and the output is the predicted sequence of local grids. Where 0≤i≤K×K; it is a one-hot vector, representing the prediction result of the i-th global grid at the t-th time;

[0020] Step 3-2, build a global sequence prediction model; the global sequence prediction model includes an encoder, a long short-term memory network LSTM, a weight adaptive model, a text representation model and a visual representation model. The input is the K×K global grids in the dataset with global grid sequence annotations obtained in step 1, and the predicted sequence of the local grids output in step 3-1. The output K×K global grid prediction sequence S t It is a one-hot vector, which indicates the prediction result at time t.

[0021] The specific method of step 4 includes:

[0022] Step 4-1, input each global grid of the data set with global grid order annotation obtained in step 1 into the encoder of the local sequence prediction model constructed in step 3-1 to obtain feature representation, input the feature representation and the all-zero vector as START into the long short-term memory network LSTM of the local sequence prediction model constructed in step 3-1, and obtain the probability O of each local grid ranking first i1 , the one with the highest probability is ranked as 1;

[0023] Step 4-2: The probability O output from step 4-1 i1 The feature representation of the local grid with the highest probability of ranking first is input into the long short-term memory network LSTM of the local sequence prediction model constructed in step 3-1 to obtain the probability of each small grid ranking second, and the one with the highest probability is ranked 2; repeat this step to obtain the prediction sequence of all local grids Where 0≤i≤K×K, O it is a one-hot vector, representing the prediction result of the i-th global grid at the t-th time;

[0024] Step 4-3, input the plane design in the dataset with global grid order annotation obtained in step 1 into the encoder of the global sequence prediction model to obtain the feature representation of the plane design, and then input the feature representation of the plane design and the all-zero vector as START into the long short-term memory network LSTM of the global sequence prediction model to obtain the probability S1 of each global grid being ranked first; the one with the largest probability is ranked 1, and the hidden state of the long short-term memory network LSTM output is h t-1 ;

[0025] Step 4-4, calculate the feature representation of the global grid text element and visual element ranked first in step 4-3, and calculate the text feature T through the text representation model t : Use word2vec to map each word to a dimensional word embedding vector, sum the vectors of all words, and input the summation result into the multi-layer perceptron to obtain an ω-dimensional element vector; finally, construct an h×w×ω text feature T by assigning all pixels inside each text element using the element-level vector and setting all remaining elements to zero t , h and w are the height and width of the global grid;

[0026] Calculate the visual features V through the visual representation model t:First, delete the text elements in the image and fill the text pixels with the background color of the graphic; then, use an image encoder based on the pre-trained classification network VGG16 to extract an image feature from the generated image, and add a global average pooling layer and two fully connected layers on top of the last convolutional layer of the classification network VGG16 to output an ω-dimensional image vector; finally, construct the h×w×ω visual representation V in the same way as the text representation t , and set other pixels to zero;

[0027] Step 4-5, based on the text feature T obtained in step 4-4 t and visual features V t And the hidden state h output by step 4-3 t-1 , the text feature T of each position is obtained through the weight adaptive model t The relative weight M t , as follows:

[0028] M t =f(h t-1 ,V t ,T t )

[0029] Among them, h t-1 is the hidden layer state at time t-1, M t has the effect of weighting the contribution of visual and textual representations to the content representation, where M t ∈[0, 1] h×w ; f is a fully connected layer network;

[0030] Step 4-6, the text feature T obtained in step 4-4 t and visual features V t By using the relative weight M obtained in steps 4-5 t Weighted, we get the feature representation C of the global grid t ; As follows:

[0031] C t =D c (M t )⊙T t +(1-D c (M t ))⊙V t

[0032] Among them, ⊙ is element-wise multiplication, D c (·) The function is to repeat M along the feature channel t c times;

[0033] Step 4-7: The feature representation C of the global grid obtained in step 4-6 tThe prediction sequence of all local grids output from step 4-2 and the probability S1 of each global grid ranking first output from step 4-3 are input into the global sequence prediction model to obtain the second global grid. Repeat steps 4-4 to 4-6 to obtain the prediction sequence of all global grids. S t is a one-hot vector representing the prediction result at time t. The cross quotient loss is calculated using the global grid prediction sequence and the global grid sequence in the dataset with annotations based on the global grid sequence obtained in step 1, and back-propagation is performed to update the multi-scale sequence prediction model parameters in step 3.

[0034] The specific method of step 5 includes:

[0035] Step 5-1, mapping the text element or visual element obtained in step 2 to the global grid by calculating the overlap ratio, where the overlap ratio is the intersection area between the global grid and the element divided by the minimum area between the global grid and the element; when the overlap ratio is greater than β, one global grid belongs to one element, and for an element that has not been assigned any global grid, when the center position of the element is within the global grid, the element is set to belong to this global grid, and each element is associated with multiple global grids, and a mapping relationship between each text element and visual element and the global grid is obtained;

[0036] Step 5-2, based on the visual elements or text elements obtained in step 2, according to the mapping relationship obtained in step 5-1, by combining the text features T of each position obtained in step 4-5 t The relative weight M t The saliency values ​​of all elements are normalized to [0, 1] by adding them together to approximate their saliency values, that is, each saliency value is divided by the maximum saliency value, and the result is convolved with a Gaussian filter to obtain the final saliency map.

[0037] The global grid refers to each image block that is meshed into a K×K grid for the plane design in the dataset with global grid order annotations obtained in step 1;

[0038] The local grid refers to each image block that is obtained by dividing the global grid into M×M grids.

[0039] The present invention also provides a graphic design element importance detection system based on weak supervision training, comprising:

[0040] The dataset acquisition module with global grid order annotation is used to divide the acquired multiple graphic designs into K×K grids, and manually annotate the design order of each cell content in the grid at the grid level to obtain a dataset with global grid order annotation, including graphic designs, K×K global grids and global grid order;

[0041] An element extraction module is used to extract elements of a graphic design from a dataset with global grid order annotations using the connectivity of components to obtain text elements and visual elements;

[0042] Multi-scale sequence prediction model building module, used to build a multi-scale sequence prediction model, including a local sequence prediction model and a global sequence prediction model. The input of the multi-scale sequence prediction model is a plane design in a dataset with global grid sequence annotations, and the output is a predicted sequence of a K×K global grid. S t is a one-hot vector, representing the prediction result at time t;

[0043] The multi-scale sequence prediction model training module is used to train the multi-scale sequence prediction model using a dataset with annotations based on the global grid sequence. During the training process, the K×K global grids of the dataset with annotations based on the global grid sequence are divided into M×M local grids of a fixed size, and each local grid is input into the local sequence prediction model to obtain the prediction sequence of each local grid; the prediction sequence of the local grid and the plane design are input into the global sequence prediction model to obtain the prediction sequence of the global grid, and the cross quotient loss is calculated using the prediction sequence of the global grid and the global grid sequence in the dataset with annotations based on the global grid sequence, and back-propagation is performed to update the parameters of the multi-scale sequence prediction model, during which the visual feature V of each position is obtained. t and text features T t The relative weight of

[0044] The element importance detection module is used to predict the visual features V of each position in the sequence results using text elements, visual elements and the global grid. t and text features T t The relative weights of the elements are used to calculate the element importance index of the graphic design and obtain the final saliency map.

[0045] The present invention also provides a device for detecting the importance of graphic design elements based on weak supervision training, comprising:

[0046] Memory: a computer program storing the above-mentioned graphic design element importance detection method based on weak supervision training, which is a computer-readable device;

[0047] Processor: used to implement the method for detecting the importance of graphic design elements based on weakly supervised training when executing the computer program.

[0048] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the method for detecting the importance of graphic design elements based on weakly supervised training.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] 1. The present invention introduces a graphic design element importance detection method based on weakly supervised training through a data set with global grid order annotations constructed in step 1 and a training method in step 4, which can reduce the dependence on a large amount of labeled data, reduce training costs, and improve generalization ability and feasibility of application in the field of graphic design.

[0051] 2. The present invention combines the local sequence prediction of step 3 with the global sequence prediction, and adopts a step-by-step processing strategy of local grids and global grids, so that the hierarchical relationship and relative importance of design elements can be captured more accurately, thereby enhancing the generalization ability and design adaptability of the model.

[0052] In summary, the present invention reduces the annotation cost and improves the training efficiency by introducing a graphic design element importance detection method based on weakly supervised training. It also enhances the generalization ability of the model by combining local and global sequence predictions, and can more accurately capture the hierarchical relationship and relative importance of design elements, with high adaptability and feasibility.

[0053] The present invention significantly improves the efficiency and accuracy of graphic design analysis by introducing a graphic design element importance detection method based on weakly supervised training. Compared with traditional methods, this technology does not require detailed element-level annotation when constructing training data, which greatly reduces the training cost. At the same time, by adopting a step-by-step processing strategy of local grids and global grids and a multi-scale sequence prediction model, it realizes efficient sequence prediction of graphic design elements. This not only enables the model to have a faster computing speed and a shorter training time, but also combines visual features V t and text features T t The relative weight calculation improves the accuracy of importance detection and makes it more generalizable in practical applications, providing a new, efficient and low-cost solution for automated analysis in the field of graphic design. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 It is a schematic diagram of the workflow of the present invention.

[0055] Figure 2 It is a schematic diagram of the overall structure of the multi-scale sequence prediction model of the present invention.

[0056] Figure 3It is a structural schematic diagram of the local sequence prediction model of the present invention.

[0057] Figure 4 It is a structural schematic diagram of the global sequence prediction model of the present invention.

[0058] Figure 5 It is a schematic diagram of the relative importance score of grid-level image-text predicted by the multi-scale sequence prediction model in an embodiment of the present invention.

[0059] Figure 6 It is a schematic diagram of the relative importance of element-level image-text predicted by a multi-scale sequence prediction model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0060] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] The existing importance detection technology for graphic design elements is often subject to expensive and cumbersome supervised training methods, which require a large number of labeled data sets for model learning. This process not only requires huge manpower and time costs, but may also be affected by problems such as labeling errors and inconsistent labels, which limits the model's adaptability to real-world scenarios. In addition, traditional supervised training methods usually require huge computing resources, which makes the training process consume huge energy and computing costs. These problems together lead to an increase in training costs and a decrease in the generalization ability of the model, making the importance detection of graphic design elements face a series of challenges in practical applications. In response to the above problems, the present invention proposes a graphic design element importance detection method based on weak supervised training to address the shortcomings of the prior art. By utilizing weak supervised training, the present invention can more effectively utilize unlabeled data, reduce the requirements for data labeling, and thus significantly reduce training costs. This method can not only improve the generalization ability of the model, but also accelerate the training speed, making the importance detection of graphic design elements more efficient. Compared with traditional methods, the multi-scale sequence prediction model of the present invention can achieve higher accuracy in a short time, while having lower costs and stronger generalization ability, which brings significant advantages to the application in the field of graphic design.

[0062] like Figure 1 As shown, a method for detecting the importance of graphic design elements based on weakly supervised training includes the following steps:

[0063] Step 1: Divide the obtained multiple graphic designs into K×K grids, and manually annotate the design order of each cell content in the grid at the grid level to obtain a dataset with annotations based on the global grid order, including graphic designs, K×K global grids, and global grid order;

[0064] Step 2, using the connectivity of the components, extract the elements of the graphic design in the dataset with global grid order annotation obtained in step 1 to obtain text elements and visual elements;

[0065] The specific method of step 2 includes:

[0066] Step 2-1, extracting initial elements from the plane design in the dataset with global grid order annotation obtained in step 1 using the connectivity of the components;

[0067] Step 2-2, judging the initial elements extracted in step 2-1, if the overlapping ratio of the bounding boxes of the elements exceeds a preset value of 0.3, merging the adjacent elements to obtain a new element set;

[0068] Step 2-3, using optical character recognition (OCR) to detect all elements in the new element set in step 2-2, if there is text, it is marked as a text element, otherwise it is a visual element.

[0069] Step 3, such as Figure 2 As shown in Figure 1, a multi-scale sequence prediction model is constructed, including a local sequence prediction model and a global sequence prediction model. The input of the multi-scale sequence prediction model is the plane design in the dataset with global grid sequence annotation obtained in step 1, and the output is the predicted sequence of K×K global grids. S t is a one-hot vector, representing the prediction result at time t;

[0070] The specific method of step 3 includes:

[0071] Step 3-1, such as Figure 3 As shown in Figure 1, a local sequence prediction model is constructed; the local sequence prediction model includes an encoder and a long short-term memory network LSTM. The input is the K×K global grids in the dataset with global grid sequence annotations obtained in step 1, and the output is the predicted sequence of local grids. Where 0≤i≤K×K; it is a one-hot vector, which represents the prediction result of the i-th global grid at the t-th time. The local sequence prediction model is used to predict the order of the local grid.

[0072] The global grid refers to each image block that is meshed into a K×K grid for the plane design in the dataset with global grid order annotations obtained in step 1;

[0073] The local grid refers to each image block that is obtained by dividing the global grid into M×M grids.

[0074] like Figure 4As shown, step 3-2, build a global sequence prediction model; the global sequence prediction model includes an encoder, a long short-term memory network LSTM, a weight adaptive model, a text representation model and a visual representation model, and the input is the K×K global grids in the dataset with global grid sequence annotations obtained in step 1, and the prediction sequence of the local grid output in step 3-1, and the output K×K global grid prediction sequence S t is a one-hot vector, representing the prediction result at time t; the global sequence prediction model is used to predict the order of the global grid.

[0075] Step 4: Use the dataset with annotations based on global grid order obtained in step 1 to train the multi-scale sequence prediction model constructed in step 3. During the training process, the K×K global grids of the dataset with annotations based on global grid order obtained in step 1 are divided into M×M local grids of fixed size, and each local grid is input into the local sequence prediction model constructed in step 3-1 to obtain the prediction sequence of each local grid; the prediction sequence of the local grid and the plane design are input into the global sequence prediction model constructed in step 3-2 to obtain the prediction sequence of the global grid, and the cross quotient loss is calculated using the prediction sequence of the global grid and the global grid sequence in the dataset with annotations based on global grid order obtained in step 1, and back propagation is performed to update the parameters of the multi-scale sequence prediction model in step 3. During this period, the visual feature V of each position will be obtained. t and text features T t The relative weight of

[0076] The specific method of step 4 includes:

[0077] Step 4-1, input each global grid of the data set with global grid order annotation obtained in step 1 into the encoder of the local sequence prediction model constructed in step 3-1 to obtain feature representation, input the feature representation and the all-zero vector as START into the long short-term memory network LSTM of the local sequence prediction model constructed in step 3-1, and obtain the probability O of each local grid being ranked first i1 , the one with the highest probability is ranked as 1;

[0078] Step 4-2: The probability O output from step 4-1 i1 The feature representation of the local grid with the highest probability of ranking first is input into the long short-term memory network LSTM of the local sequence prediction model constructed in step 3-1 to obtain the probability of each small grid ranking second, and the one with the highest probability is ranked 2; repeat this step to obtain the prediction sequence of all local grids Where 0≤i≤K×K, O itis a one-hot vector, representing the prediction result of the i-th global grid at the t-th time; Figure 3 shown.

[0079] Step 4-3, input the plane design in the dataset with global grid order annotation obtained in step 1 into the encoder of the global sequence prediction model to obtain the feature representation of the plane design, and then input the feature representation of the plane design and the all-zero vector as START into the long short-term memory network LSTM of the global sequence prediction model to obtain the probability S1 of each global grid being ranked first; the one with the largest probability is ranked 1, and the hidden state of the long short-term memory network LSTM output is h t-1 ;

[0080] Step 4-4, calculate the feature representation of the global grid text element and visual element ranked first in step 4-3, and calculate the text feature T through the text representation model t : Use word2vec to map each word to a dimensional word embedding vector, sum the vectors of all words, and input the summation result into a multi-layer perceptron to obtain a 50-dimensional element vector; finally, construct an h×w×50 text feature T by assigning all pixels inside each text element using the element-wise vector and setting all remaining elements to zero t , h and w are the height and width of the global grid;

[0081] Calculate the visual features V through the visual representation model t :First, delete the text elements in the image and fill the text pixels with the background color of the graphic; then, use an image encoder based on the pre-trained classification network VGG16 to extract an image feature from the generated image, and add a global average pooling layer and two fully connected layers on top of the last convolutional layer of the classification network VGG16 to output an ω-dimensional image vector; finally, construct the h×w×50 visual representation V in the same way as the text representation t , and set other pixels to zero;

[0082] Step 4-5, based on the text feature T obtained in step 4-4 t and visual features V t And the hidden state h output by step 4-3 t-1 , the text feature T of each position is obtained through the weight adaptive model t The relative weight M t , as follows:

[0083] M t =f(h t-1 ,V t ,T t )

[0084] Among them, h t-1 is the hidden layer state at time t-1, M t has the effect of weighting the contribution of visual and textual representations to the content representation, where M t ∈[0, 1] h×w ; f is a fully connected layer network;

[0085] Step 4-6, the text feature T obtained in step 4-4 t and visual features V t By using the relative weight M obtained in steps 4-5 t Weighted, we get the feature representation C of the global grid t ; As follows:

[0086] C t =D c (M t )⊙T t +(1-D c (M t ))⊙V t

[0087] Among them, ⊙ is element-wise multiplication, D c (·) The function is to repeat M along the feature channel t c times;

[0088] Step 4-7: The feature representation C of the global grid obtained in step 4-6 t The prediction sequence of all local grids output from step 4-2 and the probability S1 of each global grid ranking first output from step 4-3 are input into the global sequence prediction model to obtain the second global grid. Repeat steps 4-4 to 4-6 to obtain the prediction sequence of all global grids. S t is a one-hot vector representing the prediction result at time t. The cross quotient loss is calculated using the global grid prediction sequence and the global grid sequence in the dataset with annotations based on the global grid sequence obtained in step 1, and back-propagation is performed to update the multi-scale sequence prediction model parameters in step 3.

[0089] Step 5: Use the text elements and visual elements obtained in step 2 and the global grid obtained in step 4 to predict the visual features V at each position in the sequence result. t and text features T t The relative weights of the elements are used to calculate the element importance index of the graphic design and obtain the final saliency map.

[0090] The specific method of step 5 includes:

[0091] Step 5-1, mapping the text element or visual element obtained in step 2 to the global grid by calculating the overlap ratio, where the overlap ratio is the intersection area between the global grid and the element divided by the minimum area of ​​the global grid and the element; when the overlap ratio is greater than 0.5, one global grid belongs to one element, and for an element that has not been assigned any global grid, when the center position of the element is within the global grid, the element is set to belong to this global grid, and each element is associated with multiple global grids, and a mapping relationship between each text element and visual element and the global grid is obtained;

[0092] Step 5-2, based on the visual elements or text elements obtained in step 2, according to the mapping relationship obtained in step 5-1, by combining the text features T of each position obtained in step 4-5 t The relative weight M t The saliency values ​​of all elements are normalized to [0, 1] by adding them together to approximate their saliency values, that is, each saliency value is divided by the maximum saliency value, and the result is convolved with a Gaussian filter to obtain the final saliency map.

[0093] Example

[0094] The embodiment of the present invention discloses a method for detecting the importance of graphic design elements based on weakly supervised training. The method is applied to the problem of detecting the importance of elements in the graphic design process, and provides a method for detecting the importance of design elements that minimizes the model training cost, avoids data set limitations, and is generalized to other data sets.

[0095] In order to evaluate the importance detection method of design elements based on weakly supervised training, the present invention experiments the algorithm from two aspects, namely, from the element perspective and from the grid perspective, and each aspect is described below.

[0096] From the grid perspective. Specifically, given a cell, the present invention accumulates the weight maps of its visual and textual representations to obtain visual and textual weights, which are then normalized and summed to 1 to obtain the importance scores of the visual and textual elements in the cell.

[0097] The experimental results show that Figure 5 As shown in Figure 2, predicting the image and text in the first cell are almost equally important in understanding the content in the cell. In the sixth cell, the text plays a more important role due to the ambiguity of the meaning of the visual element.

[0098] From the perspective of elements. Given a graphic design, calculate an element design importance heat map. Because the model of the present invention predicts a text feature T at each position t The relative weight M t, for each cell in the global-level grid on a design, all local maps need to be aggregated to generate a global saliency map for the design. Specifically, for a visual / textual element, it is first associated with many cells according to the number of overlaps. The detailed association process is achieved by calculating the overlap ratio, which is the intersection area between the cell and the element divided by the minimum area of ​​the cell and the element. When the overlap ratio is greater than 0.5, a cell belongs to an element. For those elements that do not have any designated cell, a cell belonging to the element is set when the center position of the element is within the cell. Therefore, each element is associated with many cells with predicted sequences. Then, its saliency value is approximated by adding the visual / textual weights of the parts of the related cells in the element. Finally, all element-level saliency values ​​are normalized to [0,1] (i.e., each saliency value is divided by the maximum saliency value), and the result is convolved with a Gaussian filter to obtain the final saliency map.

[0099] The experimental results show that Figure 6 As shown in the figure, the present invention can locate the parts of a design that are important for the viewer to understand its message. For example, as shown in the second column, the present invention's method shows that the handwritten text (left) and the illustration of a person at a crosswalk (center) are equally important for understanding the message that the design is intended to convey: cross safely.

[0100] The core content of this invention is to propose a method for detecting the importance of graphic design elements based on weakly supervised training, aiming to solve the technical problems of low training cost and high generalization ability. The method includes manually annotating a data set of global grid order, dividing the graphic design into global and local grids through weakly supervised training, and using local sequence prediction models and global sequence prediction models to obtain prediction sequences and relative weights. Visual feature V t and text features T t The relative weights of V are obtained through the weight adaptive model, and finally the importance index of graphic design elements is calculated. The specific steps include a multi-scale sequence prediction model based on the long short-term memory network LSTM, a visual feature V t and text features T t Extraction, weight adaptive model, and saliency map generation. Compared with the traditional method, the present invention adopts a method based on weak supervision training, which has the advantages of low training cost and strong generalization ability, and also shows obvious superiority in operation speed, training time and accuracy. This method provides an efficient, accurate and economical solution for the importance detection of graphic design elements, and is expected to be widely used in the field of graphic design.

[0101] The inventive points that need to be protected in the scheme of the present invention are: the network model designed in the scheme and the structures of each submodule in the model, as well as the implementation process, method and steps.

[0102] The present invention also provides a graphic design element importance detection system based on weak supervision training, comprising:

[0103] The module for acquiring a dataset with annotations based on global grid order is used to implement K×K grid division of the multiple graphic designs acquired in step 1, and manually annotate the design order of each cell content in the grid at the grid level, thereby obtaining a dataset with annotations based on global grid order, including graphic design, K×K global grids, and global grid order;

[0104] An element extraction module is used to extract the elements of the graphic design in the data set with global grid order annotation obtained in step 1 by using the connectivity of the components in step 2 to obtain text elements and visual elements;

[0105] The multi-scale sequence prediction model construction module is used to implement the multi-scale sequence prediction model constructed in step 3, including a local sequence prediction model and a global sequence prediction model. The input of the multi-scale sequence prediction model is the plane design in the dataset with global grid sequence annotation obtained in step 1, and the output is the prediction sequence of K×K global grid. S t is a one-hot vector, representing the prediction result at time t;

[0106] The multi-scale sequence prediction model training module is used to implement the training of the multi-scale sequence prediction model constructed in step 3 using the data set with annotations based on the global grid sequence obtained in step 1 in step 4. During the training process, the K×K global grids of the data set with annotations based on the global grid sequence obtained in step 1 are divided into M×M local grids of fixed size, and each local grid is input into the local sequence prediction model to obtain the prediction sequence of each local grid; the prediction sequence of the local grid and the plane design are input into the global sequence prediction model to obtain the prediction sequence of the global grid, and the cross quotient loss is calculated using the prediction sequence of the global grid and the global grid sequence in the data set with annotations based on the global grid sequence obtained in step 1, and back-propagation is performed to update the multi-scale sequence prediction model parameters in step 3, during which the visual feature V of each position is obtained. t and text features T t The relative weight of

[0107] The element importance detection module is used to implement the visual feature V of each position in the text element and visual element obtained in step 2 and the global grid prediction sequence result obtained in step 4 in step 5. t and text features T t The relative weights of the elements are used to calculate the element importance index of the graphic design and obtain the final saliency map.

[0108] The present invention also provides a device for detecting the importance of graphic design elements based on weak supervision training, comprising:

[0109] Memory: a computer program storing the above-mentioned graphic design element importance detection method based on weak supervision training, which is a computer-readable device;

[0110] Processor: used to implement the method for detecting the importance of graphic design elements based on weakly supervised training when executing the computer program.

[0111] The present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it can implement the method for detecting the importance of graphic design elements based on weakly supervised training.

[0112] The same and similar parts between the various embodiments in this specification can be referred to each other. The above-mentioned embodiments of the present invention do not constitute a limitation on the protection scope of the present invention.

Claims

1. A method for detecting the importance of graphic design elements based on weakly supervised training, characterized in that: The following steps are involved: Step 1: Divide the acquired multiple graphic designs into K×K grids, and manually annotate the design order of each cell content in the grid at the grid level to obtain a dataset with annotations based on the global grid order, including graphic designs, K×K global grids, and global grid order; Step 2, using the connectivity of the components, extract the elements of the graphic design in the dataset with global grid order annotation obtained in step 1 to obtain text elements and visual elements; Step 3: Build a multi-scale sequence prediction model, including a local sequence prediction model and a global sequence prediction model. The input of the multi-scale sequence prediction model is the plane design in the dataset with global grid sequence annotation obtained in step 1, and the output is the predicted sequence of K×K global grid. S t is a one-hot vector, representing the prediction result at time t; Step 4: Use the dataset with annotations based on the global grid sequence obtained in step 1 to train the multi-scale sequence prediction model constructed in step 3. During the training process, the K×K global grids of the dataset with annotations based on the global grid sequence obtained in step 1 are divided into M×M local grids of fixed size, and each local grid is input into the local sequence prediction model to obtain the prediction sequence of each local grid; the prediction sequence of the local grid and the plane design are input into the global sequence prediction model to obtain the prediction sequence of the global grid, and the cross quotient loss is calculated using the prediction sequence of the global grid and the global grid sequence in the dataset with annotations based on the global grid sequence obtained in step 1, and back propagation is performed to update the parameters of the multi-scale sequence prediction model in step 3, during which the visual feature V of each position is obtained. t and text features T t The relative weight of Step 5: Use the text elements and visual elements obtained in step 2 and the global grid obtained in step 4 to predict the visual features V at each position in the sequence result. t and text features T t The relative weights of the elements are used to calculate the element importance index of the graphic design and obtain the final saliency map.

2. According to claim 1, a method for detecting the importance of graphic design elements based on weakly supervised training is characterized in that: The specific method of step 2 includes: Step 2-1, extracting initial elements from the plane design in the dataset with global grid order annotation obtained in step 1 using the connectivity of the components; Step 2-2, judging the initial elements extracted in step 2-1, if the overlapping ratio of the bounding boxes of the elements exceeds a preset value α, merging the adjacent elements to obtain a new element set; Step 2-3, using optical character recognition (OCR) to detect all elements in the new element set in step 2-2, if there is text, it is marked as a text element, otherwise it is a visual element.

3. The method for detecting the importance of graphic design elements based on weakly supervised training according to claim 1, characterized in that: The specific method of step 3 includes: Step 3-1, build a local sequence prediction model; the local sequence prediction model includes an encoder and a long short-term memory network LSTM, the input is the K×K global grids in the dataset with global grid sequence annotations obtained in step 1, and the output is the predicted sequence of local grids Where 0≤i≤K×K; it is a one-hot vector, representing the prediction result of the i-th global grid at the t-th time; Step 3-2, build a global sequence prediction model; the global sequence prediction model includes an encoder, a long short-term memory network LSTM, a weight adaptive model, a text representation model and a visual representation model. The input is the K×K global grids in the dataset with global grid sequence annotations obtained in step 1, and the predicted sequence of the local grids output in step 3-1. The output K×K global grid prediction sequence S t It is a one-hot vector, which indicates the prediction result at time t.

4. The method for detecting the importance of graphic design elements based on weakly supervised training according to claim 1, characterized in that: The specific method of step 4 includes: Step 4-1, input each global grid of the data set with global grid order annotation obtained in step 1 into the encoder of the local sequence prediction model constructed in step 3-1 to obtain feature representation, input the feature representation and the all-zero vector as START into the long short-term memory network LSTM of the local sequence prediction model constructed in step 3-1, and obtain the probability O of each local grid ranking first i1 , the one with the highest probability is ranked as 1; Step 4-2: The probability O output from step 4-1 i1 The feature representation of the local grid with the highest probability of ranking first is input into the long short-term memory network LSTM of the local sequence prediction model constructed in step 3-1 to obtain the probability of each small grid ranking second, and the one with the highest probability is ranked 2; repeat this step to obtain the prediction sequence of all local grids Where 0≤i≤K×K, O it is a one-hot vector, representing the prediction result of the i-th global grid at the t-th time; Step 4-3, input the plane design in the dataset with global grid order annotation obtained in step 1 into the encoder of the global sequence prediction model to obtain the feature representation of the plane design, and then input the feature representation of the plane design and the all-zero vector as START into the long short-term memory network LSTM of the global sequence prediction model to obtain the probability S1 of each global grid being ranked first; the one with the largest probability is ranked 1, and the hidden state of the long short-term memory network LSTM output is h t-1 ; Step 4-4, calculate the feature representation of the global grid text element and visual element ranked first in step 4-3, and calculate the text feature T through the text representation model t : Use word2vec to map each word to a dimensional word embedding vector, sum the vectors of all words, and input the summation result into the multi-layer perceptron to obtain an ω-dimensional element vector; finally, construct an h×w×ω text feature T by assigning all pixels inside each text element using the element-level vector and setting all remaining elements to zero t , h and w are the height and width of the global grid; Calculate the visual feature V through the visual representation model t :First, delete the text elements in the image and fill the text pixels with the background color of the graphic; then, use an image encoder based on the pre-trained classification network VGG16 to extract an image feature from the generated image, and add a global average pooling layer and two fully connected layers on top of the last convolutional layer of the classification network VGG16 to output an ω-dimensional image vector; finally, construct the h×w×ω visual representation V in the same way as the text representation t , and set other pixels to zero; Step 4-5, based on the text feature T obtained in step 4-4 t and visual features V t And the hidden state h output by step 4-3 t-1 , the text feature T of each position is obtained through the weight adaptive model t The relative weight M t , as follows: M t =f(h t-1 ,V t ,T t ) Among them, h t-1 is the hidden layer state at time t-1, M t has the effect of weighting the contribution of visual and textual representations to the content representation, where M t ∈[0, 1] h×w ; f is a fully connected layer network; Step 4-6, the text feature T obtained in step 4-4 t and visual features V t By using the relative weight M obtained in steps 4-5 t Weighted, we get the feature representation C of the global grid t ; As follows: C t =D c (M t )⊙T t +(1-D c (M t ))⊙V t Among them, ⊙ is element-wise multiplication, D c (·) The function is to repeat M along the characteristic channel t c times; Step 4-7: The feature representation C of the global grid obtained in step 4-6 t The prediction sequence of all local grids output from step 4-2 and the probability S1 of each global grid ranking first output from step 4-3 are input into the global sequence prediction model to obtain the second global grid. Repeat steps 4-4 to 4-6 to obtain the prediction sequence of all global grids. S t is a one-hot vector representing the prediction result at time t. The cross quotient loss is calculated using the global grid prediction sequence and the global grid sequence in the dataset with annotations based on the global grid sequence obtained in step 1, and back-propagation is performed to update the multi-scale sequence prediction model parameters in step 3.

5. The method for detecting the importance of graphic design elements based on weakly supervised training according to claim 1, characterized in that: The specific method of step 5 includes: Step 5-1, mapping the text element or visual element obtained in step 2 to the global grid by calculating the overlap ratio, where the overlap ratio is the intersection area between the global grid and the element divided by the minimum area between the global grid and the element; when the overlap ratio is greater than β, one global grid belongs to one element, and for an element that has not been assigned any global grid, when the center position of the element is within the global grid, the element is set to belong to this global grid, and each element is associated with multiple global grids, and a mapping relationship between each text element and visual element and the global grid is obtained; Step 5-2, based on the visual elements or text elements obtained in step 2, according to the mapping relationship obtained in step 5-1, by combining the text features T of each position obtained in step 4-5 t The relative weight M t The saliency values ​​of all elements are normalized to [0, 1] by adding them together to approximate their saliency values, that is, each saliency value is divided by the maximum saliency value, and the result is convolved with a Gaussian filter to obtain the final saliency map.

6. The method for detecting the importance of graphic design elements based on weakly supervised training according to claim 3, characterized in that: The global grid in step 3-1 refers to each image block obtained by performing K×K grid division on the plane design in the data set with the global grid sequence annotation obtained in step 1; The local grid refers to each image block that is obtained by dividing the global grid into M×M grids.

7. A graphic design element importance detection system based on weakly supervised training based on the method according to any one of claims 1 to 6, characterized in that: include: The dataset acquisition module with global grid order annotation is used to divide the acquired multiple graphic designs into K×K grids, and manually annotate the design order of each cell content in the grid at the grid level to obtain a dataset with global grid order annotation, including graphic designs, K×K global grids and global grid order; An element extraction module is used to extract elements of a graphic design from a dataset with global grid order annotations using the connectivity of components to obtain text elements and visual elements; Multi-scale sequence prediction model building module, used to build a multi-scale sequence prediction model, including a local sequence prediction model and a global sequence prediction model. The input of the multi-scale sequence prediction model is a plane design in a dataset with global grid sequence annotations, and the output is a predicted sequence of a K×K global grid. S t is a one-hot vector, representing the prediction result at time t; The multi-scale sequence prediction model training module is used to train the multi-scale sequence prediction model using a dataset with annotations based on the global grid sequence. During the training process, the K×K global grids of the dataset with annotations based on the global grid sequence are divided into M×M local grids of a fixed size, and each local grid is input into the local sequence prediction model to obtain the prediction sequence of each local grid; the prediction sequence of the local grid and the plane design are input into the global sequence prediction model to obtain the prediction sequence of the global grid, and the cross quotient loss is calculated using the prediction sequence of the global grid and the global grid sequence in the dataset with annotations based on the global grid sequence, and back-propagation is performed to update the parameters of the multi-scale sequence prediction model, during which the visual feature V of each position is obtained. t and text features T t The relative weight of The element importance detection module is used to predict the visual features V of each position in the sequence results using text elements, visual elements and the global grid. t and text features T t The relative weights of the elements are used to calculate the element importance index of the graphic design and obtain the final saliency map.

8. A device for detecting the importance of graphic design elements based on weakly supervised training, characterized in that: include: Memory: a computer program storing a method for detecting the importance of graphic design elements based on weakly supervised training as described in any one of claims 1 to 6, which is a computer-readable device; Processor: used to implement the method for detecting the importance of graphic design elements based on weakly supervised training as described in any one of claims 1 to 6 when executing the computer program.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it can implement a method for detecting the importance of graphic design elements based on weakly supervised training as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Natural scene text recognition method based on sequence transformation correction and attention mechanism

    CN111428727A

  • Weak supervision target detection method based on image attribute learning

    CN112861917A

  • Weak supervision-based image salient target recognition model and method

    CN118644743A

  • A method for training a convolutional neural network for image recognition using image-conditioned masked language modeling

    US20210312628A1

  • Image processing apparatus, image recognition system, and image processing method

    US20230080876A1