Industrial product appearance design image evaluation method based on multi-modal emotion cognition
Through the industrial product appearance design image evaluation method with multi-modal emotional cognition, semi-supervised learning and multi-teacher integrated learning, the problem that users' personalized needs in industrial product appearance design is difficult to meet, and the personalized and differentiated competitive advantages of product appearance design are achieved.
Patent Information
- Application Number
- CN202510258027.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-11
AI Technical Summary
In the appearance design of modern industrial products, it is difficult to accurately meet users' personalized needs, resulting in homogeneous product appearance design, making it difficult for consumers to choose products that meet their needs and preferences.
The image evaluation method of industrial product appearance design based on multimodal emotional cognition is adopted, including sample annotation of semi-supervised learning, sparse deconvolution image processing, multi-teacher integrated learning and multimodal data fusion. By extracting subject colors, emotional image classification and stimulus source data analysis, users' emotional needs are accurately identified.
It improves the accuracy of sample annotation and comprehensiveness of emotional recognition, and can better capture key characteristics and elements in product appearance design, meet users' emotional needs, and improve the personalization level and market competitiveness of the product.
Smart Images

Figure CN120296461A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of modern industry, and in particular to an industrial product appearance design image evaluation method based on multimodal emotion cognition. Background Art
[0002] In the modern industrial field, intelligent manufacturing technology is undergoing rapid changes and development. Advanced technologies with information technology, automation technology, artificial intelligence technology, etc. as the core are constantly integrating, driving the industrial production model from traditional labor-intensive to intelligent, automated, and digital. Under the background of intelligent manufacturing, the iteration speed of industrial product appearance design has significantly accelerated. In order to gain an advantage in the fiercely competitive market, companies have increased their investment in product appearance design, continuously innovating, and striving to attract consumers' attention with product appearance.
[0003] However, with the continuous increase in the number of products on the market and the increasingly rapid dissemination of technology, product appearance design has gradually shown a trend of homogenization in terms of key attributes such as quality, performance and function. The differences in appearance design between products of different brands and manufacturers are getting smaller and smaller, and it is difficult to form a unique competitive advantage in the market by relying on these basic attributes. This homogenization phenomenon means that when consumers choose products, they often face many products with similar appearance and similar functions, and it is difficult to find products that truly meet their needs and preferences.
[0004] At the same time, social and economic development and the improvement of people's living standards have led to a profound change in consumer demand. In the user-centered emotional era, consumers' demand for products is no longer limited to the practical functions of products, but they pay more attention to the emotional value and personalized characteristics contained in the products. They hope that the products they buy can reflect their unique personality, taste and attitude towards life, and become a way of expressing themselves. For product appearance design, how to accurately grasp and meet the personalized needs of users has become an important issue that needs to be solved urgently.
[0005] Therefore, it is necessary to provide an industrial product appearance design image evaluation method based on multimodal emotion cognition to solve the above technical problems. Summary of the invention
[0006] The present invention provides an industrial product appearance design image evaluation method based on multimodal emotion cognition, which solves the problem of how to accurately grasp and meet the personalized needs of users in product appearance design.
[0007] In order to solve the above technical problems, the present invention provides an industrial product appearance design image evaluation method based on multimodal emotion cognition, which comprises the following steps:
[0008] S1. Sample annotation method based on semi-supervised learning;
[0009] S11. To reduce the interference of noise and scattered signals on the original samples and the influence of background factors on the extraction of the main color, first, the sparse deconvolution method is used to improve the image quality;
[0010] S12. Secondly, the samples are made transparent to reduce the influence of background factors on the extraction of the main color. Subsequently, the main color of the transparentized image is extracted;
[0011] S13. Subsequently, the extracted main color is combined with the original image to generate an image based on the original image and color features;
[0012] S14. The color-enhanced image is input into the original semi-supervised learning model framework based on Stone-SAM multi-scale label optimization for training and prediction to obtain the classification information of unlabeled samples;
[0013] S2. Key feature element extraction method based on emotional cognitive classification calculation: Classify industrial product appearance design pictures based on emotional imagery, input the picture data and emotional data into a classification model of multi-teacher ensemble learning for learning and training to obtain an effective classification model. On the basis of the multi-teacher knowledge distillation method, various types of teacher information are introduced for multi-structure teacher distillation to improve the product appearance design picture classification model;
[0014] S3. Emotional recognition method technology for multi-modal cognitive data: Multimodal fusion of stimulus source data, user behavior data, and cognitive physiological data is performed to achieve emotional recognition.
[0015] Preferably, the experimental process of the sparse deconvolution in S11 includes the following steps:
[0016] S111. First, determine the fixed physical parameter values of the system, such as pixel size, wavelength, and numerical aperture parameters;
[0017] S112. Select the initial value of the fidelity, which can be 1500 or 500. The initial sparsity can be selected as one-thirtieth of the fidelity to avoid the loss of some signals in the image due to too high sparsity;
[0018] S113. Keep the ratio of the fidelity to the sparsity and decrease the values of both. During the process of decreasing the fidelity, since its relationship with continuity is reciprocal, the continuity gradually increases and the image will gradually become blurred;
[0019] S114. After determining the value of fidelity, the value of sparsity can be gradually increased. During the process of increasing sparsity, the image will become sparse, showing high-frequency information, and the lower signals in the image may be filtered out accordingly. During the process of increasing sparsity, it is necessary to pay attention to the weaker signals in the image in a timely manner, stop increasing sparsity before the weaker signals in the image are filtered out, and make a balance between the clarity of the image and the filtering of lower signals. To avoid the image becoming too blurred and smooth, the selected fidelity is 200 and the sparsity is 20.
[0020] Preferably, the main color in S12 can reflect the main color characteristics and color composition information of the picture. Since many parts of the appearance design of industrial products have common colors and account for a relatively large proportion of the colors on the surface of the entire product appearance design, in order to prevent the influence of common colors and increase the characteristics of decorative colors, multiple colors need to be extracted when extracting the main color. The five main colors with the highest proportion in the original image are obtained through the adaptive segmentation method, and are divided into color blocks of different sizes according to the proportion of the main colors for combination.
[0021] Preferably, the basic idea of the adaptive segmentation method in S12 is: for the given image, first establish coordinate axes in the color cube according to the direct segmentation method;
[0022] Each component axis is directly segmented into 4 parts, and then the representative color of each cube is obtained by the weighted average method;
[0023] The color component error of each cube is calculated respectively. To make the calculation more accurate, this method introduces a relative coefficient of brightness influence to obtain the adaptive segmentation factor, and arranges it in descending order: the axis with the largest adaptive segmentation factor is segmented again, and so on until the total number of cubes is less than or equal to 256; finally, a color palette is established with the centroid of each cube as the representative color, and the color reconstruction of the image is carried out. The calculation formula of the adaptive segmentation factor based on the relative coefficient of brightness influence:
[0024]
[0025] Preferably, the model structure in S14 is mainly divided into three stages: a semi-supervised learning module, a Stone-SAM optimization module, and a fully supervised learning module:
[0026] S141. In the first stage, the improved SPES-ORCNN algorithm is selected as the detector. This algorithm first introduces a cross-stage local spatial pyramid module in the backbone network to expand the model's receptive field and enhance the network's perception ability; secondly, combines the path aggregation network and the hybrid attention mechanism in the neck network to enrich the feature information contained in the feature map; finally, uses the flexible non-maximum suppression algorithm to optimize the prediction box screening process;
[0027] S142. In the second stage, the SoftTeacher model is used for semi-supervised training of the original dataset. This method combines consistency regularization and pseudo-label methods in semi-supervised learning, and makes full use of the unlabeled data in the dataset to jointly train the object detection model;
[0028] S143. In the third stage, the labels generated in the second stage are input after multi-scale adjustment and used as sparse box prompts to input the Stone-SAM optimization module.
[0029] Preferably, the key feature element extraction method in S2 includes the following steps:
[0030] S21. First, use the preprocessed FashionMNIST dataset to train an initial model with a complex structure, deeper layers, and excellent performance;
[0031] S22. Subsequently, use various types of knowledge such as the softened output, feature map information, and structured output of the initial model as supervision to train multiple models with the same structure as multi-teacher networks. Corresponding weights are given to multiple teachers according to the model performance to fuse the outputs of multiple models. The fused output and the sample label are used together as supervision to train the final target student model, and a lightweight model for classifying product appearance design pictures is obtained.
[0032] Preferably, in S3, the industrial product emotion dataset is used as a case sample for verification. With two emotion words as the judgment criteria, the emotional responses of users to each type of industrial product picture are obtained, and eye movement data and electroencephalogram data are obtained. The above multi-modal data is classified to identify the emotions of users.
[0033] Preferably, the realization of emotion recognition and prediction of multi-modal data in S3 includes the following steps:
[0034] 31. First, based on the feature-level fusion method, by adopting the method of multi-scale temporal convolution, features in multiple frequency bands are fully extracted. According to the characteristics of different feature splicing methods, the features are deeply fused through depthwise separable convolution and temporal convolutional network respectively to fully mine their internal information, so as to extract more accurate electroencephalogram features. The electroencephalogram features are spliced and combined under the planar structure based on the electroencephalogram spatial distribution to construct an electroencephalogram feature set. By analyzing the eye movement cognitive data, the eye movement data features that can represent cognitive psychology are extracted, and the feature data are combined and spliced to construct an eye movement feature set;
[0035] 32. Secondly, based on decision-level fusion, the EEG feature set is processed through the combination of CNN-LSTM to obtain the output of EEG feature data; the eye movement feature set is processed through a deep neural network to obtain the output of eye movement feature data; through the innovation of the VGGNet algorithm, the output of the stimulus source picture features is realized.
[0036] 33. Finally, again based on decision fusion, the three types of multimodal feature data obtained are fused, combined with the user's behavior data, and processed based on a deep neural network to finally realize emotion recognition.
[0037] Preferably, the emotion recognition method based on multimodal cognitive data proposed in S33 mainly includes the following steps:
[0038] S331. Based on the emotion dataset obtained in Technical Point 2, using industrial product pictures as stimulus samples and emotion images as target stimuli, an eye-brain cognitive experiment is constructed.
[0039] S332. For the brain-visual cognitive data obtained in the eye-brain cognitive experiment, study the correlation between EEG signals and eye movement signals based on time series, and study the correlation between each electrode channel of EEG signals and EEG features based on space series, and identify and extract the synchronous information of brain-visual cognition.
[0040] S333. For the stimulus samples in the eye-brain cognitive experiment, extract the feature data based on the product appearance design pictures, study the correlation between the product appearance design sample data based on time series and EEG signals and eye movement signals, fuse the synchronous information of brain-visual cognition of EEG data and eye movement data based on user cognition and the product appearance design structure data information, and use a convolutional neural network to process the data to construct multimodal data.
[0041] S334. According to the multimodal data fusion method, process different modal data respectively, fuse the processed multimodal data, and train again, with emotion information as the output, to construct an emotion recognition model based on multimodal cognitive data.
[0042] Compared with the related technologies, the industrial product appearance design image evaluation method based on multimodal emotion cognition provided by the present invention has the following beneficial effects:
[0043] The present invention provides an industrial product appearance design image evaluation method based on multimodal emotional cognition to improve the accuracy and reliability of sample annotation: in the sample annotation link, the present invention adopts a sample annotation method based on semi-supervised learning, improves image quality through a sparse deconvolution method, effectively reduces the interference of noise and scattered signals on the original samples, and avoids sample annotation errors that may be caused by these interference factors. At the same time, the samples are transparently processed to further reduce the influence of background factors on the extracted main color, so that the extracted main color more accurately reflects the true color characteristics of the product appearance. After the extracted main color is combined with the original image to generate a new image, it is input into a semi-supervised learning model framework based on Stone-SAM multi-scale label optimization for training and prediction. This method makes full use of the information of a small number of labeled samples and a large number of unlabeled samples, can obtain more accurate unlabeled sample classification information, improves the efficiency and accuracy of sample annotation, and provides a reliable data basis for subsequent analysis and evaluation.
[0044] Accurately extract key feature elements: The key feature element extraction method based on emotional cognitive classification calculation can accurately classify industrial product appearance design images based on emotional imagery. By inputting image data and emotional data into the classification model of multi-teacher ensemble learning for learning and training, an effective classification model is obtained. Moreover, on the basis of the multi-teacher knowledge distillation method, various types of teacher information are introduced for multi-structure teacher distillation to improve the product appearance design image classification model, so that the model can better capture the key feature elements in product appearance design. This method can deeply explore the intrinsic connection between product appearance and emotional imagery, extract features that are closely related to user emotional needs, provide more targeted guidance for industrial product appearance design, and help design products that are more in line with user emotional expectations.
[0045] Enhance the comprehensiveness and effectiveness of emotion recognition: The emotion recognition method of multimodal cognitive data integrates stimulus source data, user behavior data and cognitive physiological data in a multimodal manner. This fusion method can obtain user emotional information from multiple dimensions. Compared with single-modal data, it can more comprehensively and accurately reflect the user's true emotional state. By realizing emotion recognition through multimodal fusion, the appearance design of industrial products can better meet the emotional needs of users and enhance user satisfaction and recognition of products. At the same time, this emotion recognition method provides richer and deeper emotional information for product design and optimization, which helps companies make more scientific and reasonable decisions in the product design process and improve the market competitiveness of products.
[0046] Improve the comprehensive level of industrial product appearance design: By integrating the above steps and methods, the present invention can provide a systematic and comprehensive image evaluation system for industrial product appearance design. From ensuring the accuracy of sample annotation, to accurately extracting key feature elements, and then to comprehensively and effectively recognizing emotions, it enables industrial product appearance design to better combine the emotional needs of users and the characteristics of the product itself. This not only helps improve the appearance quality and personalization level of the product, but also enhances the differential competitive advantage of the product in the market, promoting the development of industrial product appearance design towards a more intelligent and user-friendly direction, and meeting the changing market demands and user expectations. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 FIG. is a schematic structural diagram of a preferred embodiment of the industrial product appearance design image evaluation method based on multi-modal emotion recognition provided by the present invention;
[0048] Figure 2 FIG. is a flowchart for extracting key feature elements of product appearance design;
[0049] Figure 3 FIG. is a schematic diagram of the emotion recognition method for multi-modal cognitive data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0050] The present invention will be further described below with reference to the drawings and embodiments.
[0051] First Embodiment
[0052] Please refer to Figure 1 、 Figure 2 and Figure 3 , wherein, Figure 1 FIG. is a schematic structural diagram of a preferred embodiment of the industrial product appearance design image evaluation method based on multi-modal emotion recognition provided by the present invention; Figure 2 FIG. is a flowchart for extracting key feature elements of product appearance design; Figure 3 FIG. is a schematic diagram of the emotion recognition method for multi-modal cognitive data. The industrial product appearance design image evaluation method based on multi-modal emotion recognition includes the following steps:
[0053] S1. A sample annotation method based on semi-supervised learning;
[0054] S11. To reduce the interference of noise and scattered signals on the original sample and reduce the influence of background factors on the extraction of the main color, first use the sparse deconvolution method to improve the image quality;
[0055] S12. Secondly, perform transparency processing on the sample to reduce the influence of background factors on the extraction of the main color, and then extract the main color from the transparently processed image;
[0056] S13. Subsequently, combine the extracted main color with the original image to generate an image based on the original image and color features;
[0057] S14. Input the image with enhanced color into the model framework of the original semi - supervised learning model framework optimized based on Stone - SAM multi - scale labels for training and prediction to obtain the classification information of unlabeled samples;
[0058] S2. Key feature element extraction method based on emotional cognitive classification calculation: Classify industrial product appearance design pictures based on emotional imagery, input the picture data and emotional data into the classification model of multi - teacher ensemble learning for learning and training to obtain an effective classification model. On the basis of the multi - teacher knowledge distillation method, introduce various types of teacher information for multi - structure teacher distillation to improve the product appearance design picture classification model;
[0059] S3. Emotional recognition method technology for multi - modal cognitive data: Perform multi - modal fusion on the stimulus source data, user behavior data, and cognitive physiological data to achieve emotional recognition.
[0060] The experimental process of the sparse deconvolution in S11 includes the following steps:
[0061] S111. First, determine the fixed physical parameter values of the system, such as pixel size, wavelength, and numerical aperture parameters;
[0062] S112. Select the initial value of the fidelity. You can choose 1500 or 500. The initial sparsity can be chosen as one - thirtieth of the fidelity to avoid the loss of some signals in the image due to too high sparsity;
[0063] S113. Keep the ratio of fidelity to sparsity and decrease both values. During the process of decreasing the fidelity, since its relationship with continuity is reciprocal, the continuity gradually increases and the image will gradually become blurred;
[0064] S114. After determining the value of the fidelity, the value of the sparsity can be gradually increased. During the process of increasing the sparsity, the image will become sparse and show high - frequency information. The lower signals in the image may be filtered out accordingly. During the process of increasing the sparsity, pay attention to the weak signals in the image in a timely manner and stop increasing the sparsity before the weak signals in the image are filtered out. Make a balance between the clarity of the image and the filtering of lower signals. To avoid the image becoming too blurred and smooth, the selected fidelity is 200 and the sparsity is 20.
[0065] The main color in S12 can reflect the main color features and color composition information of the picture. Since many different parts of the industrial product's appearance design have common colors, and the proportion of the colors on the surface of the entire product appearance design is relatively large, in order to prevent the influence of the common colors and increase the characteristics of the decorative colors, multiple colors need to be extracted when extracting the main color. The five main colors with the highest proportion in the original image are obtained through the adaptive segmentation method, and they are divided into color blocks of different sizes according to the proportion of the main colors for combination.
[0066] The basic idea of the adaptive segmentation method in S12 is as follows: for the given image, first establish coordinate axes in the color cube according to the direct segmentation method;
[0067] Each component axis is directly segmented into 4 parts, and then the representative color of each cube is obtained by the weighted average method;
[0068] The color component error of each cube is calculated respectively. To make the calculation more accurate, this method introduces the relative coefficient of brightness influence to obtain the adaptive segmentation factor, and arranges them in descending order: the axis with the largest adaptive segmentation factor is segmented again, and so on until the total number of cubes is less than or equal to 256; finally, a color palette is established with the centroid of each cube as the representative color, and the color reconstruction of the image is carried out. The calculation formula of the adaptive segmentation factor based on the relative coefficient of brightness influence:
[0069]
[0070] Among them, ωR = 2.1, ω G = 1.5, ω B = 4.3. After the color image is input, it is decomposed into three primary colors R, G, and B. The color coordinate transformation makes R, G, and B become x ( 1 ) 、 x ( 2 ) 、 x ( 3 ) and other 3 components. s represents selecting N representative colors from the X set after pixel color quantization xn ′ = (x nR ′, x nG ′, x nB ′)(1 ≤ n ≤ N); then the pre-displayed colors are merged into N groups, that is, we divide the entire image into N regions.
[0071] The model structure in S14 is mainly divided into three stages: the semi-supervised learning module, the Stone-SAM optimization module, and the fully supervised learning module:
[0072] S141. In the first stage, the improved SPES-ORCNN algorithm is selected as the detector. First, a cross-stage local spatial pyramid module is introduced into the backbone network to expand the model's receptive field and enhance the network's perception ability. Second, the path aggregation network and the hybrid attention mechanism are combined in the neck network to enrich the feature information contained in the feature map. Finally, the flexible non-maximum suppression algorithm is used to optimize the prediction box screening process.
[0073] S142. In the second stage, the SoftTeacher model is used for semi-supervised training of the original dataset. This method combines the consistency regularization and pseudo-label methods in semi-supervised learning and fully utilizes the unlabeled data in the dataset to jointly train the object detection model.
[0074] S143. In the third stage, the labels generated in the second stage are input after multi-scale adjustment and used as sparse box prompts to input the Stone-SAM optimization module.
[0075] The semi-supervised learning model framework based on Stone-SAM multi-scale label optimization is as Figure 1 shown, where xu is the unlabeled training sample, xg is the labeled dataset, xu is the unlabeled training set. Box is the unannotated dataset xu The label data generated by the semi-supervised object detection model Soft Teacher is used as the position encoding prompt of the optimization module. xo represents the optimized dataset.
[0076] This module adopts the strategy of knowledge distillation to lightweight the network structure of the large vision model SAM and embeds the neural network classifier PP-HGNetV2 to provide the model with the ability of semantic judgment.
[0077] The key feature element extraction method in S2 includes the following steps:
[0078] S21. First, use the preprocessed FashionMNIST dataset to train an initial model with a complex structure, deep layers, and excellent performance.
[0079] S22. Subsequently, use various types of knowledge such as the softened output, feature map information, and structured output of the initial model as supervision to train multiple models with the same structure as multi-teacher networks. According to the model performance, corresponding weights are given to multiple teachers to fuse the outputs of multiple models. The fused output and the sample label are used as supervision to train the final target student model, and a lightweight model for product appearance design picture classification is obtained.
[0080] In S3, an industrial product emotion dataset is used as a case sample for verification. With two emotion vocabularies as the criteria, the emotional responses of users to each type of industrial product picture are obtained, and eye movement data and electroencephalogram (EEG) data are acquired. The above-mentioned multimodal data is classified to identify the emotions of users.
[0081] The realization of emotion recognition and prediction for multimodal data in S3 includes the following steps:
[0082] 31. First, based on the feature-level fusion method, by adopting the multi-scale temporal convolution method, features in multiple frequency bands are fully extracted. According to the characteristics of different feature splicing methods, depthwise separable convolution and temporal convolutional network are respectively used to deeply fuse the features, fully mining their internal information to extract more accurate EEG features. The EEG features are spliced and combined under the planar structure based on the EEG spatial distribution to construct an EEG feature set. By analyzing the eye movement cognitive data, the eye movement data features that can represent cognitive psychology are extracted, and the feature data are combined and spliced to construct an eye movement feature set.
[0083] 32. Second, based on the decision-level fusion, the EEG feature set is processed by combining CNN-LSTM to obtain the output of the EEG feature data; the eye movement feature set is processed by a deep neural network to obtain the output of the eye movement feature data; through the innovation of the VGGNet algorithm, the output of the stimulus source picture features is realized.
[0084] 33. Finally, again based on the decision fusion, the three types of multimodal feature data obtained are fused, combined with the user's behavior data, and processed based on a deep neural network to finally realize emotion recognition.
[0085] The proposed emotion recognition method based on multimodal cognitive data in S33 mainly includes the following steps:
[0086] S331. Based on the emotion dataset obtained in Technical Point 2, using industrial product pictures as stimulus samples and emotion images as target stimuli, an eye-brain cognitive experiment is built.
[0087] S332. For the brain-visual cognitive data obtained in the eye-brain cognitive experiment, study the correlation between the EEG signals and eye movement signals based on time series, study the correlation between each electrode channel of the EEG signals and the EEG features based on spatial series, and identify and extract the synchronous information of brain-visual cognition.
[0088] S333. For the stimulus samples in the eye-brain cognitive experiment, extract the feature data based on the product appearance design pictures, study the correlation between the product appearance design sample data based on time series and the electroencephalogram (EEG) signals and eye movement signals, fuse the brain-visual cognitive synchronous information of EEG data and eye movement data based on user cognition with the product appearance design structure data information, and use a convolutional neural network to process the data to construct multimodal data.
[0089] S334. According to the multimodal data fusion method, process the data of different modalities separately, fuse the processed multimodal data, and train again to construct an emotion recognition model based on multimodal cognitive data with emotion information as the output.
[0090] The process of extracting the key feature elements of the product appearance design is as Figure 2 shown. Extract the salient regions of the product appearance design pictures under specific emotional images, and explore the correlation between the feature elements of this region and the key feature elements extracted by traditional methods. In view of the fact that traditional methods can only extract the common key feature elements of this type of product appearance design, while the changes in brain waves can obtain the key features of the regions of interest that affect the user's emotional image for each product appearance design, this study will verify the key features obtained by deep learning technology with the help of brain wave data testing.
[0091] Decision-level fusion and feature-level fusion are the basis of deep learning hybrid fusion. Through deep learning hybrid fusion, different modalities of data can be processed separately according to their characteristics to achieve the optimal multimodal data fusion effect.
[0092] Compared with related technologies, the industrial product appearance design image evaluation method based on multimodal emotion cognition provided by the present invention has the following beneficial effects:
[0093] The present invention provides an industrial product appearance design image evaluation method based on multimodal emotion cognition, which improves the accuracy and reliability of sample annotation: In the sample annotation link, the present invention adopts a sample annotation method based on semi-supervised learning. By using the sparse deconvolution method to improve the image quality, the interference of noise and scattered signals on the original samples is effectively reduced, and the sample annotation errors caused by these interference factors are avoided. At the same time, the samples are made transparent to further reduce the influence of background factors on the extraction of the main color, so that the extracted main color can more accurately reflect the true color characteristics of the product appearance. After combining the extracted main color with the original image to generate a new image, it is input into the semi-supervised learning model framework based on Stone-SAM multi-scale label optimization for training and prediction. This method makes full use of the information of a small number of labeled samples and a large number of unlabeled samples, can obtain more accurate classification information of unlabeled samples, improves the efficiency and accuracy of sample annotation, and provides a reliable data basis for subsequent analysis and evaluation.
[0094] Accurately extract key feature elements: The key feature element extraction method based on emotional cognitive classification calculation can accurately classify industrial product appearance design images based on emotional imagery. By inputting image data and emotional data into the classification model of multi-teacher ensemble learning for learning and training, an effective classification model is obtained. Moreover, on the basis of the multi-teacher knowledge distillation method, various types of teacher information are introduced for multi-structure teacher distillation to improve the product appearance design image classification model, so that the model can better capture the key feature elements in product appearance design. This method can deeply explore the intrinsic connection between product appearance and emotional imagery, extract features that are closely related to user emotional needs, provide more targeted guidance for industrial product appearance design, and help design products that are more in line with user emotional expectations.
[0095] Enhance the comprehensiveness and effectiveness of emotion recognition: The emotion recognition method of multimodal cognitive data integrates stimulus source data, user behavior data and cognitive physiological data in a multimodal manner. This fusion method can obtain user emotional information from multiple dimensions. Compared with single-modal data, it can more comprehensively and accurately reflect the user's true emotional state. By realizing emotion recognition through multimodal fusion, the appearance design of industrial products can better meet the emotional needs of users and enhance user satisfaction and recognition of products. At the same time, this emotion recognition method provides richer and deeper emotional information for product design and optimization, which helps companies make more scientific and reasonable decisions in the product design process and improve the market competitiveness of products.
[0096] Improve the comprehensive level of industrial product appearance design: Combining the above steps and methods, the present invention can provide a systematic and comprehensive image evaluation system for industrial product appearance design, from ensuring the accuracy of sample annotation, to the precise extraction of key feature elements, to the comprehensive and effective emotion recognition, so that the industrial product appearance design can better combine the user's emotional needs and the characteristics of the product itself. This not only helps to improve the product's appearance quality and personalization level, but also can enhance the product's differentiated competitive advantage in the market, and promote the development of industrial product appearance design in a more intelligent and humanized direction to meet the ever-changing market needs and user expectations.
[0097] The above descriptions are merely embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.
Claims
1. An industrial product appearance design image evaluation method based on multimodal emotion recognition, characterized in that Including the following steps: S1. Sample annotation method based on semi-supervised learning; S11. To reduce the interference of noise and scattered signals on the original samples and the influence of background factors on the extraction of the main color, first use the sparse deconvolution method to improve the image quality; S12. Secondly, perform transparency processing on the samples to reduce the influence of background factors on the extraction of the main color, and then extract the main color from the transparently processed image; S13. Subsequently, combine the extracted main color with the original image to generate an image based on the original image and color features; S14. Input the color-enhanced image into the original semi-supervised learning model framework based on Stone-SAM multi-scale label optimization for training and prediction to obtain the classification information of unlabeled samples; S2. Key feature element extraction method based on emotional cognitive classification calculation: Classify industrial product appearance design pictures based on emotional imagery, input the picture data and emotional data into a classification model of multi-teacher ensemble learning for learning and training to obtain an effective classification model, and introduce various types of teacher information for multi-structural teacher distillation on the basis of the multi-teacher knowledge distillation method to improve the product appearance design picture classification model; S3. Emotional recognition method technology for multi-modal cognitive data: Perform multi-modal fusion of stimulus source data, user behavior data, and cognitive physiological data to achieve emotional recognition.
2. The industrial product appearance design image evaluation method based on multi-modal emotion recognition according to claim 1, characterized in that The experimental process of the sparse deconvolution in S11 includes the following steps: S111. First, determine the fixed physical parameter values of the system, such as pixel size, wavelength, and numerical aperture parameters; S112. Select the initial value of the fidelity, which can be 1500 or 500, and the initial sparsity can be selected as one-thirtieth of the fidelity to avoid the loss of some signals in the image due to too high sparsity; S113. Keep the ratio of the fidelity to the sparsity and reduce the values of both. During the process of reducing the fidelity, since it is the reciprocal relationship with continuity, the continuity gradually increases and the image will gradually become blurred; S114. After determining the value of the fidelity, the value of the sparsity can be gradually increased. During the process of increasing the sparsity, the image will become sparse and show high-frequency information, and the lower signals in the image may be filtered out accordingly. During the process of increasing the sparsity, pay attention to the weaker signals in the image in a timely manner, stop increasing the sparsity before the weaker signals in the image are filtered out, and make a balance between the clarity of the image and the filtering of the lower signals. In order to avoid the image becoming too blurred and smooth, the selected fidelity is 200 and the sparsity is 20.
3. The industrial product appearance design image evaluation method based on multi-modal emotion recognition according to claim 1, wherein, The main color in S12 can reflect the main color characteristics and color composition information of the picture. Since there are common colors in different parts of the appearance design of many industrial products and the color proportion of the entire product appearance design surface is relatively large, in order to prevent the influence of common colors and increase the characteristics of decorative colors, multiple colors need to be extracted when extracting the main color. The five main colors with the highest proportion in the original image are obtained through the adaptive segmentation method, and are divided into color blocks of different sizes for combination according to the proportion of the main color.
4. The industrial product appearance design image evaluation method based on multi-modal emotion recognition according to claim 3, characterized in that The basic idea of the adaptive segmentation method in S12 is as follows: For a given image, first establish coordinate axes in the color cube according to the direct segmentation method; Directly segment each component axis into 4 parts, and then use the weighted average method to find the representative color of each cube; Calculate the color component error of each cube. To make the calculation more accurate, this method introduces a relative coefficient of brightness influence to find the adaptive segmentation factor, and arranges them in descending order: Segment the axis with the largest adaptive segmentation factor again, and so on until the total number of cubes is less than or equal to 256; Finally, establish a color palette with the centroid of each cube as the representative color and perform color reconstruction of the image. The calculation formula of the adaptive segmentation factor based on the relative coefficient of brightness influence:
5. The industrial product appearance design image evaluation method based on multi-modal emotion recognition according to claim 1, characterized in that The model structure in S14 is mainly divided into three stages: the semi-supervised learning module, the Stone-SAM optimization module, and the fully supervised learning module; S141. In the first stage, select the improved SPES-ORCNN algorithm as the detector. This algorithm first introduces a cross-stage local spatial pyramid module in the backbone network to expand the model's receptive field and enhance the network's perception ability; Secondly, combine the path aggregation network and the hybrid attention mechanism in the neck network to enrich the feature information contained in the feature map; Finally, use the flexible non-maximum suppression algorithm to optimize the prediction box screening process; S142. In the second stage, use the SoftTeacher model to perform semi-supervised training on the original dataset. This method combines the consistency regularization and pseudo-label methods in semi-supervised learning, and makes full use of the unlabeled data in the dataset to jointly train the object detection model; S143. In the third stage, take the labels generated in the second stage as inputs for multi-scale adjustment, and use them as sparse box hints to input into the Stone-SAM optimization module.
6. The industrial product appearance design image evaluation method based on multi-modal emotion recognition according to claim 1, characterized in that The key feature element extraction method in S2 includes the following steps: S21. First, use the preprocessed FashionMNIST dataset to train an initial model with a complex structure, deep layers, and excellent performance; S22. Subsequently, use various types of knowledge such as the softened output, feature map information, and structured output of the initial model as supervision, train multiple models with the same structure as multi-teacher networks, and assign corresponding weights to multiple teachers according to the model performance to fuse the outputs of multiple models. The fused output and the sample label are used as supervision to train the final target student model, and a lightweight model for product appearance design picture classification is obtained.
7. The method for evaluating the industrial product appearance design image based on multi-modal emotion recognition according to claim 1, characterized in that In S3, use the industrial product emotion dataset as a case sample for verification, use two emotion words as the judgment to obtain the emotional response of users to each type of industrial product picture, and obtain eye movement data and electroencephalogram data. Classify the above multi-modal data to identify the emotions of users.
8. The industrial product appearance design image evaluation method based on multi-modal emotion recognition according to claim 1, wherein The realization of emotion recognition and prediction of multi-modal data in S3 includes the following steps:
31. First, based on the feature-level fusion method, by adopting the multi-scale temporal convolution method, the features of multiple frequency bands were fully extracted. According to the characteristics of different feature splicing methods, depthwise separable convolution and temporal convolutional network were respectively used to deeply fuse the features, fully mining their internal information to extract more accurate EEG features. And the EEG features were spliced and combined under the planar structure based on the EEG spatial distribution to construct an EEG feature set. By analyzing the eye movement cognitive data, the eye movement data features that can represent cognitive psychology were extracted, and the feature data were combined and spliced to construct an eye movement feature set.
32. Second, based on the decision-level fusion, the EEG feature set was processed by combining CNN-LSTM to obtain the output of the EEG feature data; the eye movement feature set was processed by a deep neural network to obtain the output of the eye movement feature data; through the innovation of the VGGNet algorithm, the output of the stimulus source picture features was realized.
33. Finally, based on the decision fusion again, the three types of multi-modal feature data obtained were fused, combined with the user's behavior data, and processed based on a deep neural network to finally realize the recognition of emotions.
9. The industrial product appearance design image evaluation method based on multi-modal emotion recognition according to claim 1, characterized in that The emotion recognition method based on multi-modal cognitive data proposed in S33 mainly includes the following steps: S331. Based on the emotion dataset obtained in Technical Point 2, using industrial product pictures as stimulus samples and emotion images as target stimuli, an eye-brain cognitive experiment was built. S332. For the brain-visual cognitive data obtained in the eye-brain cognitive experiment, the correlation between the EEG signal and the eye movement signal based on the time series was studied, and the correlation between each electrode channel of the EEG signal and the EEG features based on the spatial series was studied to identify and extract the synchronous information of brain-visual cognition. S333. For the stimulus samples in the eye-brain cognitive experiment, the feature data based on the product appearance design pictures were extracted, the correlation between the product appearance design sample data based on the time series and the EEG signal and the eye movement signal was studied, and the synchronous information of brain-visual cognition of the EEG data and eye movement data based on user cognition was fused with the product appearance design structure data information, and the convolutional neural network was used to process the data to construct multi-modal data. S334. According to the multi-modal data fusion method, different modal data were processed separately, and the processed multi-modal data were fused and trained again, with emotion information as the output, to construct an emotion recognition model based on multi-modal cognitive data.