Visual design image classification method based on resnet-50 network model
Through the visual design image classification method based on the resnet-50 network model, image segmentation and feature description mining are used, combined with vector interleaving processing and residual mapping, the problem of difficulty in visual design image classification is solved, and the accurate classification of images and efficient classification of models are achieved.
Patent Information
- Application Number
- CN202510257614.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-03-05
AI Technical Summary
There are many visual design images, making it difficult for users to quickly find images of the required categories among many images and create them based on this image.
A visual design image classification method based on the resnet-50 network model is adopted, and image segmentation, shape, color and texture feature description mining is carried out, and a cross entropy loss function is constructed to improve the classification ability of the model.
The accurate classification of visual design images is realized, the classification ability of the model is improved, the images between different categories have a large inter-class distance, and the image search and creation process is simplified.
Smart Images

Figure CN120147740A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image classification, and specifically to a visual design image classification method based on the resnet-50 network model. Background Art
[0002] Visual design includes logo design, packaging design, image design, advertising design, interface design, etc., and is a means and result of subjective forms targeting users' visual functions. Among them, image design refers to using visual design means, through the modeling of logos and specific color expressions, etc., to form an overall image of an enterprise's business philosophy, behavioral concept, management characteristics, product packaging style, marketing guidelines and strategies. Advertising design refers to various manual or computer painting means or imaging technologies, as well as creative image design using composite methods.
[0003] In the process of visual design, it may be necessary to draw inspiration from existing images. However, there are a large number of visual design images, and it is difficult for users to quickly find the images of the required category among them and create based on these images. Therefore, classifying visual design images has become an urgent problem to be solved. Summary of the Invention
[0004] The purpose of the present invention is to solve the problem that there are a large number of visual design images, and it is difficult for users to quickly find the images of the required category among them and create based on these images, and to propose a visual design image classification method based on the resnet-50 network model.
[0005] The technical solution of the present invention to solve the above technical problems is as follows: A visual design image classification method based on the resnet-50 network model, including the following steps: S1. Obtain the visual design image set P and preprocess P to obtain , and divide into a construction set , a training set and a test set ; S2. Perform image segmentation on all images in the construction set to obtain w sub-images ; S3. Perform image shape description mining, image color description mining, and image texture description mining on the w sub-images to respectively obtain m visual design image shape feature vectors , , n visual design image color feature vectors , and p visual design image texture feature vectors , ; S4. Interleave the shape feature vectors of m visual design images , the color feature vectors of n visual design images and the texture feature vectors of p visual design images , and map them to the resnet-50 network model to form a visual design image network model; S5. Construct a cross-entropy loss function to make the visual design images of different categories farther apart; S6. Train the model, input the image set training set in S1 into the visual design image network model in S4, use the SGD optimizer to train the model, and learn and optimize the model parameters by calculating the loss through the loss function in S5; S7: Obtain the visual design image category. After segmenting the images in the test set in S1, input them into the model trained in S6 to obtain the corresponding visual design image category.
[0006] On the basis of the above technical solutions, the present invention can also be improved as follows.
[0007] Preferably, the specific steps of image shape description mining in S3 are as follows: First, perform image enhancement processing on the sub-image , perform grayscale conversion, then introduce histogram equalization to redistribute the grayscale values of the image pixels to obtain an image contour map. Use shape descriptors to quantify the shape features of the objects in the image, use the boundary feature method to obtain the image shape features in the image contour map, use principal component analysis to convert the high-dimensional feature vectors in the image shape features into low-dimensional feature vectors, encode the image shape features, and convert them into shape numerical vectors. Remove redundant shape numerical vectors to finally obtain the shape feature vectors of the visual design images .
[0008] Preferably, the specific steps of image color description mining in S3 are as follows: Obtain the color space of the sub-image , quantize the color space of the sub-image into several bins, then use color automatic segmentation technology to divide the image into several regions. Use each region to index each color component of the quantized color space to form a binary color index set. Then, for the sub-image Convert it into a color histogram, map each color in the image to an index value in the color index set, and count the occurrence frequency of each color in the image to obtain the color histogram. Divide the pixels of each bin of the color histogram into aggregated pixels and non-aggregated pixels. By calculating the proportion and distribution of the aggregated pixels, obtain the color aggregation vector. Select the most representative features from the color aggregation vectors to form a set of color aggregation vectors, and use principal component analysis to perform dimensionality reduction on the high-dimensional color aggregation vectors in the set of color aggregation vectors. Finally, encode and transform the color aggregation vector into a color numerical vector, and the color numerical vector is the visual design image color feature vector 。
[0009] Preferably, the specific steps of image texture description mining in S3 are as follows: For the sub-image Perform grayscale processing to obtain a grayscale image. Use the gray-level co-occurrence matrix to obtain the energy, contrast, entropy, and inverse difference of the grayscale image, and use the energy, contrast, correlation, entropy, and inverse difference of the grayscale image as the texture feature vector. Use principal component analysis to perform dimensionality reduction on the high-dimensional feature vectors among them, and finally obtain the visual design image texture feature vector 。
[0010] Preferably, the image enhancement processing includes the following steps: Process the sub-image image pixels through spatial domain enhancement to enhance the image, and its expression is as follows: In the formula, f(x, y) is the original image, g(x, y) is the processed image, T is the spatial domain operation on the pixels in the neighborhood of the original image at the position (x, y), and the neighborhood of the point (x, y) is defined as a square or rectangular area centered on (x, y).
[0011] Preferably, the histogram equalization includes the following two formulas: Transformation function: Among them, r is the original gray value, s is the transformed gray value, pr(r) is the probability density function of the original gray value, L is the number of gray levels, and the value is 256; Probability density function transformation: Among them, ps(s) is the probability density function of the changed gray value, pr(r) is the probability density function of the original gray value, L is the number of gray levels, and the value is 256.
[0012] Preferably, the formula for constructing the color aggregation vector is as follows: Among them, CCV is the color aggregation vector, is the aggregated pixel number of the Nth color, is the non-aggregated pixel number of N colors.
[0013] Preferably, the calculation formulas for the energy, contrast, entropy, and inverse difference are as follows: Energy: Among them, e and f represent the indices of gray levels, and P(e, f) is the element value in the gray-level co-occurrence matrix; Contrast: Among them, e and f represent the indices of gray levels, and GLGM(e, f) represents the value at the corresponding position in the gray-level co-occurrence matrix; Entropy: Among them, e and f represent the indices of gray levels, and P(e, f) is the element value in the gray-level co-occurrence matrix; Inverse difference: Among them, e and f represent the indices of gray levels, and P(e, f) is the element value in the gray-level co-occurrence matrix.
[0014] Preferably, the specific steps of S4 are as follows: First, use the vector interleaving technique to generate a drawing command set, which first draws the color feature vector , and then draws the texture feature vector at the unoccupied positions of the color feature vector , and finally draws the shape feature vector and the color feature vector at the unoccupied positions of the color feature vector , and finally maps it to the resnet-50 network model through the residual mapping formula; the residual mapping formula is as follows: Among them, F(x) is the initial input feature, W λ is the linear transformation matrix of the visual design image color feature vector, λ is the mapping weight of the visual design image color feature vector, W ω is the linear transformation matrix of the visual design image texture feature vector, ω is the mapping weight of the visual design image texture feature vector, W θ is the linear transformation matrix of the visual design image shape feature vector, and θ is the mapping weight of the visual design image shape feature vector.
[0015] Preferably, the cross-entropy loss function in S5 is: where l(y b = d) takes a value of 0 or 1. When the value is 1, it means that when y b is correctly classified into class d. When the value is 0, it means that y b is not correctly classified into class d. N represents the number of pictures, C represents the picture categories involved, a is a random number in the interval (2, 3), represents the probability that the b-th image is determined to be d, and L is the cross-entropy loss.
[0016] Compared with the prior art, the technical solution of the present application has the following beneficial technical effects: 1. The present invention uses cross-entropy loss to ensure that during the training process of the model, there is a large inter-class distance between images of different categories, improving the classification ability of the model.
[0017] 2. The present invention obtains the shape, color, and texture feature vectors of the visual design images, and performs vector interleaving processing. The fusion of the three ensures the accurate classification of the visual design images, with the shape feature vector having the highest priority, the texture feature vector second, and the color feature vector having the lowest priority, thus ensuring the feature priority in the visual design image network model.
[0018] 3. The present invention maps the shape, color, and texture feature vectors of the visual design images, which are three different vectors, to the resnet-50 network model through residual mapping, solving the problem of inconsistent scales of the three different vectors during the mapping process. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a schematic flowchart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0021] A method for classifying visual design images based on the resnet-50 network model includes the following steps: S1. Obtain the set P of visual design images and preprocess P to obtain , and divide into a construction set , a training set , and a test set ; S2. For the construction set Perform image segmentation on all images within to obtain w sub-images ; S3. For the w sub-images perform mining of image shape description, mining of image color description, and mining of image texture description, and respectively obtain m visual design image shape feature vectors , , n visual design image color feature vectors , and p visual design image texture feature vectors , ; S4. Interleave the m visual design image shape feature vectors , the n visual design image color feature vectors and the p visual design image texture feature vectors and map them to the resnet-50 network model to form a visual design image network model; S5. Construct a cross-entropy loss function to make the visual design images between different categories farther apart; S6. Train the model. Input the image set training set in S1 into the visual design image network model in S4, use the SGD optimizer to train the model, and learn and optimize the model parameters by calculating the loss through the loss function in S5; S7: Obtain the visual design image categories. After segmenting the images in the test set in S1 , input them into the model trained in S6 to obtain the corresponding visual design image categories.
[0022] Among them, the preprocessing steps include removing duplicate, invalid, or irrelevant images; performing denoising processing, using filters (such as mean filtering, Gaussian filtering, etc.) to remove noise in the images; performing image enhancement, highlighting important information in the images by adjusting parameters such as the contrast and brightness of the images; performing normalization processing, scaling the image pixel values to a fixed range for subsequent processing; performing size adjustment, adjusting the images to a unified size for convenient batch processing and model input, and setting the parameters after size adjustment to 448×448, constructing a set , training set and test set in a ratio of 7:2:1.
[0023] Generally, visual design images are classified by shape into geometric shape images and free-form images, by color into monochromatic images and color images, and by texture into natural texture images, artificial texture images, and mixed texture images. Among them, geometric shape images are further subdivided into circles, rectangles, triangles, stars, etc. Monochromatic images are further subdivided into red, orange, yellow, green, cyan, blue, purple, white, black, etc. Color images include random combinations of multiple monochromatic colors. Natural texture images are further subdivided into clouds, smoke, fog, wood grain, rock texture, etc. Therefore, the specific classification of visual design image categories is a combination form of the above classifications, such as free-form white cloud texture images (since there are numerous image classifications and it is difficult to list them all, they are illustrated by examples).
[0024] The specific steps for mining image shape descriptions in S3 are as follows: First, perform image enhancement processing on the sub-image and convert it to grayscale. Then introduce histogram equalization to redistribute the grayscale values of the image pixels to obtain an image contour map. Use shape descriptors to quantify the shape features of the objects in the image, use the boundary feature method to obtain the image shape features in the image contour map, use principal component analysis to convert the high-dimensional feature vectors in the image shape features into low-dimensional feature vectors, encode the image shape features, and convert them into a shape numerical vector. Remove the redundant shape numerical vectors, and finally obtain the visual design image shape feature vector .
[0025] The specific steps for mining image color descriptions in S3 are as follows: Obtain the color space of the sub-image Quantify the color space of the sub-image into several bins, and then use color automatic segmentation technology to divide the image into several regions. Use each region to index each color component of the quantified color space to form a binary color index set. Then convert the sub-image into a color histogram, map each color in the image to an index value in the color index set, and count the occurrence frequency of each color in the image to obtain the color histogram. Divide the pixels of each bin in the color histogram into aggregated pixels and non-aggregated pixels. By calculating the proportion and distribution of the aggregated pixels, obtain the color aggregation vector. Select the most representative features in the color aggregation vector to form a color aggregation vector set, and use principal component analysis to perform dimensionality reduction on the high-dimensional color aggregation vectors in the color aggregation vector set. Finally, encode and convert the color aggregation vector into a color numerical vector, and the color numerical vector is the visual design image color feature vector .
[0026] The specific steps for mining image texture descriptions in S3 are as follows: Sub - image Perform grayscale processing on it to obtain a grayscale image. Use the gray - level co - occurrence matrix to obtain the energy, contrast, entropy, and inverse difference of the grayscale image, and use the energy, contrast, correlation, entropy, and inverse difference of the grayscale image as the texture feature vector. Use principal component analysis to perform dimensionality reduction on the high - dimensional feature vector among them, and finally obtain the texture feature vector of the visual design image 。
[0027] The image enhancement process includes the following steps: Process the sub - image through spatial domain enhancement Image pixels, enhance the image, and its expression is as follows: In the formula, f(x, y) is the original image, g(x, y) is the processed image, T is the spatial domain operation on the pixels in the neighborhood of the original image at the position (x, y), and the neighborhood of the point (x, y) is defined as a square or rectangular area centered on (x, y).
[0028] Histogram equalization includes the following two formulas: Transformation function: Among them, r is the original gray - level value, s is the transformed gray - level value, pr(r) is the probability density function of the original gray - level value, L is the number of gray - levels, and the value is 256; Probability density function transformation: Among them, ps(s) is the probability density function of the transformed gray - level value, pr(r) is the probability density function of the original gray - level value, L is the number of gray - levels, and the value is 256.
[0029] Preferably, the color aggregation vector formula is constructed as follows: Among them, CCV is the color aggregation vector, is the number of aggregated pixels of the Nth color, is the number of non - aggregated pixels of the N colors.
[0030] Preferably, the calculation formulas for energy, contrast, entropy, and inverse difference are as follows: Energy: Among them, e and f represent the indices of gray - levels, and P(e, f) is the element value in the gray - level co - occurrence matrix; Contrast: Among them, e and f represent the indices of gray - levels, and GLGM(e, f) represents the value at the corresponding position in the gray - level co - occurrence matrix; Entropy: where e and f represent the indices of gray levels, and P(e, f) is the element value in the gray-level co-occurrence matrix; Inverse difference: where e and f represent the indices of gray levels, and P(e, f) is the element value in the gray-level co-occurrence matrix.
[0031] The specific steps of S4 are as follows: First, use the vector interleaving technique to generate a set of drawing commands. This set of commands first draws the color feature vector , and then draws the texture feature vector at the positions not occupied by the color feature vector , and finally draws the shape feature vector at the positions not occupied by the drawn color feature vector and the color feature vector ; finally, map it to the resnet-50 network model through the residual mapping formula. The residual mapping formula is as follows: where F(x) is the initial input feature, W λ is the linear transformation matrix of the color feature vector of the visual design image, λ is the mapping weight of the color feature vector of the visual design image, W λ is the linear transformation matrix of the color feature vector of the visual design image, λ is the mapping weight of the color feature vector of the visual design image, W ω is the linear transformation matrix of the texture feature vector of the visual design image, ω is the mapping weight of the texture feature vector of the visual design image, W θ is the linear transformation matrix of the shape feature vector of the visual design image, and θ is the mapping weight of the shape feature vector of the visual design image.
[0032] It should be noted that since the present invention emphasizes the shape feature vector of the visual design image , followed by the texture feature vector of the visual design image , and finally the color feature vector of the visual design image , so generally θ > ω > λ, θ is generally taken as , ω is generally taken as , and λ is generally taken as The cross-entropy loss function in S5 is: where the value of l(y b = d) is 0 or 1. When the value is 1, it means that y b is correctly classified into class d, and when the value is 0, it means that y b is not correctly classified into class d. N represents the number of pictures, C represents the picture categories involved, and a is a random number in the interval (2, 3). It represents the probability that the b-th image is determined as d, and L is the cross-entropy loss.
[0033] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the said element.
[0034] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A visual design image classification method based on the resnet-50 network model, characterized in that: The following steps are involved: S1. Obtain the visual design image set P and preprocess P to obtain ,Will Build Set , training set and test set ; S2. Build a set Segment all images in the image to obtain w sub-images ; S3, for w sub-images Perform image shape description mining, image color description mining, and image texture description mining to obtain m visual design image shape feature vectors respectively , , n visual design image color feature vectors , and p visual design image texture feature vectors , ; S4, transform m visual design image shape feature vectors , n visual design image color feature vectors and p visual design image texture feature vectors Perform vector interleaving and map it to the resnet-50 network model to form a visual design image network model; S5. Construct a cross entropy loss function to make the distance between visual design images of different categories farther; S6, training model, training set of images in S1 , input into the visual design image network model of S4, use the SGD optimizer to train the model, and calculate the loss through the loss function of S5 to learn and optimize the model parameters; S7: Get the visual design image category and test the set in S1 After image segmentation, the image is input into the trained model of S6 to obtain the corresponding visual design image category.
2. According to claim 1, a visual design image classification method based on the resnet-50 network model is characterized in that: The specific steps of image shape description mining in S3 are as follows: First, the sub-image Perform image enhancement and grayscale conversion, then introduce histogram equalization, redistribute the grayscale values of image pixels, obtain image contour map, use shape descriptors to quantify the shape features of objects in the image, use boundary feature method to obtain image shape features in image contour map, use principal component analysis to convert high-dimensional feature vectors in image shape features into low-dimensional feature vectors, encode image shape features, and convert them into shape numerical vectors, remove redundant shape numerical vectors, and finally obtain visual design image shape feature vectors .
3. The visual design image classification method based on the resnet-50 network model according to claim 1 is characterized in that: The specific steps of image color description mining in S3 are as follows: Get sub-image The color space of the sub-image The color space is quantized into several handles, and then the image is divided into several regions using color automatic segmentation technology. Each region is indexed by each color component of the quantized color space to form a binary color index set, and then the sub-image is Convert to color histogram, map each color in the image to an index value in the color index set, and count the frequency of occurrence of each color in the image to obtain a color histogram. Divide the pixels of each handle of the color histogram into aggregated pixels and non-aggregated pixels. By calculating the proportion and distribution of aggregated pixels, a color aggregation vector is obtained. The most representative features in the color aggregation vector are selected to form a color aggregation vector set. The high-dimensional color aggregation vector in the color aggregation vector set is reduced in dimension using principal component analysis. Finally, the color aggregation vector is encoded and converted into a color numerical vector. The color numerical vector is the color feature vector of the visual design image. .
4. The visual design image classification method based on the resnet-50 network model according to claim 1 is characterized in that: The specific steps of image texture description mining in S3 are as follows: Pair of images Grayscale processing is performed to obtain a grayscale image. The energy, contrast, entropy and inverse difference of the grayscale image are obtained using the grayscale co-occurrence matrix. The energy, contrast, correlation, entropy and inverse difference of the grayscale image are used as texture feature vectors. The high-dimensional feature vectors are reduced in dimension using principal component analysis to finally obtain the texture feature vector of the visual design image. .
5. The visual design image classification method based on the resnet-50 network model according to claim 2 is characterized in that: The image enhancement process comprises the following steps: Processing sub-images by spatial domain enhancement Image pixels, enhanced image, its expression is as follows: Where f(x,y) is the original image, g(x,y) is the processed image, T is the spatial domain operation on the pixels in the neighborhood of the original image position (x,y), and the neighborhood of point (x,y) is defined as a square or rectangular area with (x,y) as the center point.
6. The visual design image classification method based on the resnet-50 network model according to claim 2 is characterized in that: The histogram equalization includes the following two formulas: Transformation function: Among them, r is the original gray value, s is the transformed gray value, pr(r) is the probability density function of the original gray value, and L is the gray level, which is 256; Probability density function transformation: Among them, ps(s) is the probability density function of the changed grayscale value, pr(r) is the probability density function of the original grayscale value, and L is the grayscale level, which is 256.
7. The visual design image classification method based on the resnet-50 network model according to claim 3 is characterized by: The formula for constructing the color aggregation vector is as follows: Among them, CCV is the color aggregation vector, is the number of aggregated pixels of the Nth color, is the number of non-aggregated pixels of N colors.
8. The visual design image classification method based on the resnet-50 network model according to claim 1, characterized in that: The calculation formulas for energy, contrast, entropy and inverse difference are as follows: energy: Among them, e and f represent the index of gray level, and P (e, f) is the element value in the gray level co-occurrence matrix; Contrast ratio: Among them, e and f represent the index of gray level, and GLGM (e, f) represents the value of the corresponding position in the gray level co-occurrence matrix; entropy: Among them, e and f represent the index of gray level, and P (e, f) is the element value in the gray level co-occurrence matrix; Inverse Gap: Among them, e and f represent the index of gray level, and P(e,f) is the element value in the gray level co-occurrence matrix.
9. The visual design image classification method based on the resnet-50 network model according to claim 1, characterized in that: The specific steps of S4 are as follows: First, use the vector interleaving technique to generate a drawing command set, which first draws the color feature vector , and secondly in the color feature vector Draw texture feature vectors at unoccupied locations , and finally draw the color feature vector and the color feature vector Draw shape feature vectors at unoccupied locations , and finally mapped to the resnet-50 network model through the residual mapping formula; the residual mapping formula is as follows: Among them, F(x) is the initial input feature, W λ is the linear transformation matrix of the visual design image color feature vector, λ is the mapping weight of the visual design image color feature vector, W ω is the linear transformation matrix of the visual design image texture feature vector, ω is the mapping weight of the visual design image texture feature vector, W θ is the linear transformation matrix of the visual design image shape feature vector, and θ is the mapping weight of the visual design image shape feature vector.
10. The visual design image classification method based on the resnet-50 network model according to claim 1, characterized in that: The cross entropy loss function in S5 is: Among them, l (y b =d) takes a value of 0 or 1. When it takes a value of 1, it means that when y b Correctly classified into d categories, the value is 0, which means y b Not correctly classified into d categories, N is the number of pictures, C is the category of pictures involved, a is a random number in the interval (2,3), It represents the probability that the bth image is judged as d, and L is the cross entropy loss.
Citation Information
Patent Citations
Image conversion training method and device for visually impaired people
CN116797448A
Visual feature constraint-based photovoltaic identification method and related device
CN119540751A
Techniques for fine-tuning a machine learning model to reconstruct a three-dimensional scene
US20240169652A1