Automobile appearance design generation method based on perceptual engineering and generative artificial intelligence

By combining sensory engineering and generative artificial intelligence, semantic networks and design generation models are built, the problem of difficult to meet consumers' emotional needs in automotive appearance design is solved, and diverse design images are generated to meet the needs, and the design scheme is enhanced and iterated.

CN120217564AInactive Publication Date: 2025-06-27ZHEJIANG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510686261.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-06-27
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively meet consumer emotional needs in automotive designs, especially when innovative designs are carried out under relatively fixed physical forms, designers face huge challenges.

Method used

Combining sensory engineering and generative artificial intelligence, by obtaining vehicle model review data, extracting emotional words in the comment text, building a semantic network, and building a vehicle appearance design generation model under the framework of this network and large language model, including keyword extraction module, emotional divergence module, design feature translation module, visual module and auxiliary comparison module.

Benefits of technology

This method can provide designers with richer and more comprehensive design references, clarify design directions, generate diverse design images that meet consumers' emotional needs, and generate results through cross-modal and multi-dimensional comparative understanding and comparison, promoting the enhancement and iteration of design solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217564A_ABST
    Figure CN120217564A_ABST
Patent Text Reader

Abstract

The invention discloses an automobile appearance design generation method based on perceptual engineering and generative artificial intelligence. The method comprises the steps of obtaining automobile type comment data of an automobile; constructing an emotion word set corresponding to the vehicle type; constructing a semantic network based on the emotion word set and the labels corresponding to the vehicle types; constructing an automobile appearance design generation model under the framework of the semantic network and the large language model; and inputting the content related to the vehicle model description into the vehicle appearance design generation model to obtain a conceptual design drawing of the vehicle model and a comparison result between the conceptual design drawing and the existing vehicle model. The method provided by the invention can provide richer and more comprehensive design reference for designers in combination with perceptual engineering and artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of intelligent artificial-aided design, and particularly relates to a method for generating automotive exterior designs based on kansei engineering and generative artificial intelligence. Background Art

[0002] The exterior of a product is a key factor influencing consumers' purchasing decisions. The attractiveness of a product's exterior is closely related to its perceived attributes, and it needs to subjectively meet consumers' emotional needs. However, creating an exterior design that meets these needs poses a great challenge to designers. Designers not only need to have an accurate understanding of abstract emotional needs but also be able to translate them into specific and appropriate design elements. This challenge is particularly prominent in automotive exterior design. As a typical industrial product, the physical form of a car is relatively fixed due to its specific functions and high production costs. Innovating on the relatively fixed physical form to meet consumers' emotional needs places high demands on designers' professional skills. However, the exterior is an important consideration for consumers when purchasing a car. Once their emotional needs are not met, it will affect the overall sales volume of the car, and the trial and test costs of car manufacturers will also increase significantly.

[0003] Kansei Engineering (KE), as a technology that transforms consumers' emotional feelings towards products into design elements, has been widely used to analyze consumers' emotional needs and identify corresponding design features to support product exterior design. Among them, "kansei" refers to consumers' psychological feelings or imaginations about the exterior of new products, and these feelings or imaginations are usually condensed into "emotional words" for expression. Kansei engineering has three types: the first type gradually decomposes the zero-level emotional words of a product into more specific multiple emotional words (n-level emotional words) to obtain design details. The second and third types respectively use computer systems and mathematical models to infer the corresponding design features based on emotional words.

[0004] Patent document CN115186368A discloses a method, device, computer device, and car for automotive exterior design. The method includes: obtaining a car image training set, where the car image training set includes several car wireframe diagrams, several functional areas are divided in the car wireframe diagrams, and at least a part of the car wireframe diagrams are replacement wireframe diagrams, and some functional areas of the replacement wireframe diagrams are filled with replacement design elements; using the DCGAN+Attention algorithm to train the car image training set to construct a first car sketch generator and obtain a first design sketch output by the first car sketch generator.

[0005] Patent document CN118194437A discloses a creative design assistance system for automobile taillight shape, which includes an automobile taillight shape home page unit, a generative design unit, a design evaluation unit, and a personal center unit. The automobile taillight shape home page unit presents an automobile taillight shape design data set, which is classified into block, bar, array, and other shapes. The generative design unit calls the progressive generative adversarial network algorithm model to visualize the automobile taillight shape generation design. The design evaluation unit includes an automatic evaluation unit and a manual evaluation unit. The automatic evaluation unit calls the deep residual network algorithm to automatically calculate the taillight design score based on the taillight image feature information uploaded by the user. The manual evaluation unit provides users with real-time scoring and calculates the total score.

[0006] Academic literature A product form design method integrating Kansei engineering and diffusion model[J]. 2023. discloses a method for generating appearance design, which uses diffusion models to randomly generate a batch of product images as design solutions, and then uses a support vector machine (SVM) trained with questionnaire data to score the generated product images according to the sentiment words selected by the designer, and returns the images with higher scores to the designer. Although this method helps to generate design solutions, it still uses sentiment words, design variables and a small amount of user questionnaire data predefined by experts to train SVM, which makes it unable to evaluate design solutions outside the predefined design space, and the accuracy of its evaluation will also be affected by the sample bias of the tested users. Summary of the invention

[0007] The purpose of the present invention is to provide a method for generating automobile appearance design based on Kansei engineering and generative artificial intelligence, which can combine Kansei engineering and artificial intelligence to provide designers with richer and more comprehensive design references.

[0008] In order to achieve the purpose of the present invention, the following technical solution is provided: a method for generating automobile appearance design based on Kansei engineering and generative artificial intelligence, comprising the following steps: Acquire vehicle model review data of the vehicle, wherein the vehicle model review data includes an appearance image of the vehicle model and a review text corresponding to the appearance of the vehicle model; The pre-trained large language model is used to extract sentiment words used to describe the appearance of the vehicle model in the review text to construct a sentiment word set corresponding to the vehicle model. Construct a semantic network based on the set of sentiment words and the tags corresponding to vehicle models. In the semantic network, sentiment words are used as nodes, and relevant edges connecting two nodes are constructed on the condition that two sentiment words appear in the review text of the same vehicle model's appearance. The relevant edges include the number of occurrences of each vehicle model. Construct a vehicle exterior design generation model under the framework of the semantic network and the large language model. The vehicle exterior design generation model includes a keyword extraction module, a sentiment divergence module, a design feature translation module, a visualization module, and an auxiliary comparison module: The keyword extraction module is used to extract the original sentiment words in the input content. The sentiment divergence module is used to perform similarity matching between the extracted original sentiment words and the set of sentiment words to output the target sentiment word with the highest similarity value, and select the top N most relevant divergent sentiment words starting from the target sentiment word in the semantic network. The design feature translation module constructs a corresponding one-way sentiment word path according to the sentiment words selected by the user from the top N divergent sentiment words, and counts the number of occurrences of the vehicle models corresponding to the relevant edges between adjacent sentiment words on the one-way sentiment word path in the semantic network, and outputs the exterior images of the top M vehicle models with the most occurrences of the vehicle models. Extract visual features from the exterior image selected by the user from the top M exterior images to output the text descriptions of each component of the vehicle model in the exterior image. The visualization module performs text-to-image operations according to the text descriptions output by the design feature translation module and the sentiment words selected by the user from the top N divergent sentiment words to generate a concept design drawing and perform visual output. The auxiliary comparison module compares the similarity between the original sentiment words input by the user and the exterior image of the vehicle model with the generated concept design drawing to output a comparison result. Input the content about the vehicle model description into the vehicle exterior design generation model to obtain the concept design drawing of the vehicle model and the comparison result with the existing vehicle models.

[0009] The present invention constructs a corresponding semantic network through the sentiment words in the vehicle model image and the corresponding description, and combines the semantic network with artificial intelligence to construct a vehicle exterior design generation model, and uses the vehicle exterior design generation model to assist designers in the vehicle exterior design work.

[0010] Specifically, the expression of the semantic network is as follows: ; where K represents the set of sentiment words, V represents the set of vehicle models, and respectively represent the i th sentiment word and the jAn emotional word, indicating the i relevant edge between the j th emotional word and the

[0011] th emotional word. The weight value of the relevant edge is set according to the co-occurrence frequency of the two emotional words in different review texts, that is, the weight is the number of times the two emotional words appear in the same review, and the more times, the greater the relevant weight.

[0012] Specifically, the visual feature extraction includes extracting features of the shape, color, and texture of vehicle components in the appearance image.

[0013] Specifically, the auxiliary comparison module includes a graph-to-graph comparison part and a text-to-graph comparison part; The graph-to-graph comparison part calculates the similarity between the visual feature vectors in the appearance image and the concept design drawing; The text-to-graph comparison part calculates the similarity between the divergent emotional words and the feature vectors of each region in the concept design drawing.

[0014] Specifically, the graph-to-graph comparison part uses the ViT model to perform different types of visual feature extraction on the appearance image and the concept design drawing to obtain the shape visual feature vector, color visual feature vector, and texture visual feature vector of the vehicle in the image, and outputs the corresponding type as the image comprehensive feature vector; Reduce the obtained shape visual feature vector, color visual feature vector, texture visual feature vector, and image comprehensive feature vector to two-dimensional space as the coordinates of the corresponding image in the two-dimensional space; Use the offset between the feature vectors of the appearance image and the concept design drawing in the two-dimensional space as a parameter during similarity calculation to obtain the result of the graph-to-graph comparison.

[0015] Specifically, the text-to-graph comparison part extracts features of the input divergent emotional words and the concept design drawing through the CLIP model to obtain keyword feature vectors and image feature vectors in the same dimension, and uses the region flipping algorithm and calculates the similarity based on the keyword feature vectors and image feature vectors to obtain the comparison result between the divergent emotional words and the concept design drawing. The specific process of the region flipping algorithm is as follows: Divide the concept design drawing into multiple image blocks; Extract the keyword feature vector of the divergent emotional word through the CLIP model, as well as the global image feature vector corresponding to the complete concept design drawing and the local image feature vector of the concept design drawing after masking one image block; Calculate the first cosine similarity between the keyword feature vector and the global image feature vector, and calculate the second cosine similarity between the keyword feature vector and the local image feature vector, and use the difference between the first cosine similarity and the second cosine similarity as the correlation score of the local image feature vector corresponding to the masked image patch; Statistically analyze the correlation scores of all image patches and perform normalization processing to output the result of the comparison between text and image.

[0016] Specifically, the calculation formula of the correlation score is as follows: ; ; where, represents the correlation score between the I th image patch of image i and the divergent emotion word, T represents the divergent emotion word, represents the original image, represents the i th inverted image corresponding to the image patch, and the inverted image is created by masking the i th image patch in the complete conceptual design drawing, represents calculating the cosine similarity between image I and the divergent emotion word T , is the image encoder of the CLIP model, is the text encoder of the CLIP model.

[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: combining kansei engineering with artificial intelligence, starting from a specific emotion word describing consumers' emotional needs, exploring the design space, clarifying the design direction, and generating diverse design images that meet the requirements; at the same time, understanding and comparing multiple generation results through cross-modal and multi-dimensional comparisons, thereby promoting the improvement of text-to-image prompts and the enhancement of design schemes, forming an iterative cycle. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 is a flowchart of the method for generating an automotive exterior design based on kansei engineering and generative artificial intelligence provided in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated in the accompanying drawings here can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0020] As Figure 1 shown, a method for generating a car exterior design based on kansei engineering and generative artificial intelligence provided in this embodiment includes the following steps: Obtain the vehicle model review data, where the vehicle model evaluation data includes the exterior images of the vehicle model and the review text corresponding to the exterior of the vehicle model; Extract the emotional words used to describe the sensory perception of the vehicle model's exterior shape from the review text through a pre-trained large language model to construct an emotional word set corresponding to the vehicle model; Construct a semantic network based on the emotional word set and the labels corresponding to the vehicle models. In the semantic network, the emotional words are used as nodes, and the relevant edges connecting the two nodes are constructed on the condition that the two emotional words appear in the review text of the same vehicle model exterior. The relevant edges include the number of times each vehicle model appears; Construct a car exterior design generation model within the framework of this semantic network and the large language model. This car exterior design generation model includes a keyword extraction module, an emotional divergence module, a design feature translation module, a visualization module, and an auxiliary comparison module.

[0021] Among them, the keyword extraction module is used to extract the original emotional words in the input content.

[0022] The emotional divergence module matches the similarity between the extracted original emotional words and the emotional words corresponding to each node in the semantic network to output the top N divergent emotional words with the highest similarity values.

[0023] The design feature translation module constructs a corresponding one-way emotional word path according to the emotional words selected by the user from the top N divergent emotional words, and counts the number of times the vehicle models corresponding to the relevant edges of two adjacent emotional words on the one-way emotional word path appear in the semantic network, and outputs the exterior images of the top M vehicle models with the most appearances of the vehicle models; Extract the visual features of the exterior image selected by the user from the top M exterior images to output the text descriptions of the various components of the vehicle model in the exterior image.

[0024] The visualization module performs text-to-image operations based on the text description output by the design feature translation module and the emotional words selected by the user from the top N divergent emotional words to generate a conceptual design drawing and perform visual output.

[0025] The auxiliary comparison module compares the similarity between the generated conceptual design drawing and the original emotional words input by the user and the exterior images of the vehicle models to output the comparison result.

[0026] Input the content about the vehicle model description into the vehicle exterior design generation model to obtain the conceptual design drawing of the vehicle model and the comparison result with the existing vehicle models.

[0027] Furthermore, in order to construct the vehicle exterior design generation model provided in the above embodiments, it is necessary to be based on kansei engineering and semantic network. From the perspective of kansei engineering, in order to capture the true emotional feedback of users towards various vehicle exterior designs, the review texts of users for multiple active vehicle models are crawled through mainstream automotive review websites or apps, and at the same time, the exterior design images of the corresponding vehicle models are attached when organizing the data.

[0028] Use a large language model to extract the emotional words that reflect users' perception of the exterior design of each vehicle model from the review texts of each vehicle model. In this embodiment, a total of 9,137 emotional words are obtained from 14,709 reviews of 2,922 vehicle models. For example, the three emotional words "good-looking", "young", and "dynamic" are extracted from the review "The shape of this car is very good-looking, young and full of vitality".

[0029] In addition, in order to avoid data duplication or ambiguity, frequency analysis is performed on the collected emotional words, and words that appear less than three times or have a length exceeding five characters are excluded because most of these are mis-extracted. Then, the large language model is used again to merge the remaining emotional words. For example, "luxury" and "sumptuous" are merged into "luxury". After the merged emotional words are manually reviewed, a set of 1,255 emotional words used in the kansei engineering engine is finally obtained.

[0030] The semantic network is constructed based on the extracted emotional words and related vehicle models. Each node in this network represents an emotional word. When two words appear in the reviews of the same vehicle model at the same time, they are defined as related, and they are connected by an edge with the corresponding vehicle model label. A total of 88,568 edges are finally formed in the semantic network. Therefore, a relationship in the semantic network can be represented as , where K represents the set of emotional words, V represents the set of vehicle models. The weight of the edge between two nodes is determined by the number of associated vehicle models, that is . Therefore, two emotional words and The higher the co-occurrence frequency in different vehicle model reviews, the heavier the weight of the edge, indicating a closer relationship between them.

[0031] The emotional divergence module provided in this embodiment aims to assist designers in exploring the design space and clarifying the design direction starting from an emotional word that describes the emotional needs of users. The specific implementation process is as follows: perform a divergence operation based on the emotional word in the user input content, and use the text-embedding-3-small model proposed by OpenAI to calculate the embedding vector of the emotional word. Then, the engine will retrieve the closest emotional word in the semantic network constructed previously, which allows designers to input any personalized vocabulary outside the semantic network. After that, the engine retrieves the top 10 most strongly related emotional words from the semantic network and provides them to the large language model as the basic knowledge for reasoning. The large language model then infers the five emotional words highly related to the emotional word specified by the designer . This method combining the semantic network and the large language model ensures the reliability of the results through the factual basis stored in the semantic network, while exploring potential and underutilized design spaces using the generation randomness of the large language model. The present invention presents the emotional word divergence results to the designer in the form of a mind map, and the designer can continue to select the inferred emotional words for further divergence.

[0032] Taking "elegant" as an example, the divergence result of the emotional word "elegant" by the emotional divergence module in this embodiment, that is, the divergence results are various emotional words such as "luxury", "high-end", "exquisite", "streamlined", etc. In this embodiment, the divergence distance is set to 2 steps.

[0033] The specific execution process of the design feature translation module provided in this embodiment is as follows. After completing the emotional word divergence, the emotional words that the user is satisfied with can be selected. In this embodiment, it is , and a series of relevant relationships in the semantic network are used to represent this unidirectional emotional word path, and the three vehicle models that most frequently appear on the relevant edges in the semantic network are counted. The appearance images of these three vehicle models are output, and visual feature extraction is performed according to the selected appearance images, that is, for the grille, air intake, front headlights, body, rearview mirrors, wheels, door handles, and rear lights in the vehicle in the appearance images, visual feature extraction is performed from the three angles of shape, color, and texture to obtain the corresponding text descriptions.

[0034] After obtaining the text description, the selected divergent sentiment words are input into the large language model to construct the prompt words of the text graph. The text graph model used in this embodiment is the text graph model Dall-E 3. The constructed prompt words are input into the text graph model Dall-E 3 to generate a conceptual design diagram, thereby providing a reference for designers.

[0035] The quality of the prompt words will greatly affect the generated results. However, it is difficult for designers to express their abstract design intentions by writing accurate high-quality text prompt words in a short period of time. In order to improve the quality of the generated prompt words, this embodiment also uses a small amount of high-quality prompt words collected from the Internet to fine-tune the large language model.

[0036] The auxiliary comparison module in this embodiment includes a picture-to-picture comparison part and a text-to-picture comparison part, wherein the picture-to-picture comparison allows designers to compare all generated and recommended design solutions from the perspective of image similarity. The text-to-picture comparison visualizes the relationship between the text prompt words and the generated design solutions. With these functions, designers can quickly compare and filter all generated design solutions, extract useful design features, and iterate the text prompt words again based on the text-to-picture comparison results to continuously optimize the generated design solutions.

[0037] Most existing text-to-image models use CLIP text to encode text prompt words. The CLIP model models a joint embedding space of visual images and text language, thereby achieving alignment of text and images.

[0038] Therefore, this embodiment uses ViT as the visual backbone model. Different attention heads in ViT extract different semantic information in the image, among which 22 layers 1 head extracts shape features, 22 layers 11 heads extract color information, and 23 layers 12 heads extract texture information. By inputting an image of a car model into ViT and calculating the output of the image classifier on the above-mentioned different attention heads, a vector representing the shape, color and texture features of the car image can be obtained. In addition, the output of the classifier in the last layer is used as a vector representing the comprehensive features of the image. Then, t-Distributed Stochastic NeighborEmbedding (TSNE) is used to reduce these high-dimensional feature vectors of all generated design solutions to two-dimensional space as the coordinates corresponding to the design solution in two-dimensional space, and visualize them on the canvas, and compare the generated design solutions from four dimensions: shape, color, texture and comprehensive. At this time, the design solution images with similar features in a certain dimension will be clustered together, while the distance between the solution images with large feature differences will be larger. This can help designers explore the design space extensively and compare the differences between multiple solutions.

[0039] Regarding the part of text-image comparison, this embodiment adopts a regional inversion algorithm. That is, for any divergent emotional word and the generated conceptual design drawing, this algorithm first divides the image into a total of 64 local regions. For each region, the cosine similarity between the feature vector of the original complete image and the feature vector of the keyword, as well as the cosine similarity between the feature vector of the remaining image excluding this local region and the keyword vector, are calculated respectively. The difference between the two represents the influence of the design features in this region on the performance of the entire product design in terms of the target keyword, that is, the correlation between this region and the keyword. The calculation of the correlation can be expressed by the formula: ; ; In the formula, represents the I th i image patch of the image T and the correlation score between the divergent emotional word, represents the original image, represents the i th inverted image corresponding to the image patch. The inverted image is created by masking the i th image patch in the complete conceptual design drawing, represents calculating the cosine similarity between the image I and the divergent emotional word T , is the image encoder of the CLIP model, is the text encoder of the CLIP model.

[0040] Table 1 shows the eight main emotional words in automotive exterior design summarized in the academic paper "Kansei engineering for new energy vehicle exterior design: An internet big data mining approach". We input these eight emotional words into the KE engine and the baseline method proposed in the present invention for testing, and require them to recommend five existing vehicle models for each emotional word. The baseline method requires GPT-4 to imitate professional automotive designers and recommend existing vehicle models according to the input emotional words. Taking the emotional word "elegant" as an example, the prompt input to GPT-4 by the baseline method is as follows: You are an automotive exterior designer proficient in kansei engineering. Please recommend five existing vehicle models according to the emotional word "elegant".

[0041] For the KE engine, we require it to first perform two rounds of divergence on the input sentiment words, generating a total of 25 sentiment word paths. Then, we extract the vehicle models on the edges of these paths in the semantic network, that is, the vehicle models related to the diverged sentiment words, and sort them according to the frequency of the vehicle models. The top five vehicle models with the highest frequency are used as the output of the KE engine. Each method generated 40 vehicle models in total. We manually retrieved the front 45-degree perspective appearance images of these vehicle models from the Internet and used CLIP to calculate the sentiment similarity between the output of each image and the corresponding sentiment word input, that is, calculate the cosine similarity between the feature vectors of the two. To reduce interference, we used object recognition technology to reset the background of all car images to white. Table 1 shows the evaluation results of the KE engine and the baseline method. As can be seen from Table 1, the KE engine performs better on most target sentiments, and the average sentiment similarity between the recommended vehicle model images and the target sentiment words is higher. This shows that compared with the baseline method that directly relies on the large language model to recommend design solutions, the results obtained by the KE engine by combining the semantic network with the large language model are more accurate.

[0042] 。

[0043] Twenty vehicle models were randomly selected from the existing collected design solutions, and each vehicle model included images of its front, rear, side, and front 45-degree perspectives. Two designers were invited to write design feature descriptions for the randomly selected 10 vehicle models respectively based on their professional knowledge. At the same time, these 20 groups of images were also input into the KE engine to extract design feature descriptions. Subsequently, two expert designers were invited to annotate the design feature descriptions extracted by the KE engine and the designers. This process included identifying the total number of features, the number of correct features, and the number of overlapping features between the two groups. The results in Table 2 show that the design features extracted by the KE engine are comparable to those of the designers in terms of accuracy and have an obvious advantage in terms of quantity.

[0044] 。

[0045] However, designers often use abstract and concise language to describe the overall style and contour of cars, such as "elegant and low", "sculptural art form", etc. The descriptions extracted by the KE engine not only include some features considered by the designers, such as the "sculptural line" in the body shape, but also perform well in using more concrete language to describe local features in detail, such as "grille shape: large hourglass shape; taillight shape: L-shaped surround design". This shows that the KE engine can not only help designers expand the scope of design feature descriptions, but also use more specific and accurate language to express the designers' design intentions to the text-to-image model.

[0046] Design comparison and experimental group: 16 subjects (7 females, 9 males) were recruited through social media. All subjects are currently working designers in the automotive industry, with an average work experience of 2.5 years. We chose ChatGPT as the baseline method for comparison with the present invention because its underlying generation model is the same as the model used in the present invention. To evaluate the effectiveness of the proposed method flow of the present invention rather than the performance of the underlying generation model, when using the baseline system, subjects can not only use the text-to-image DALL-E-3 for automotive image generation, but also use the text generation function of GPT-4 and the image recognition function of GPT-4V. This simulates the scenario where designers use generation models in real work. They can also search for automotive or other design images on the Internet to obtain inspiration to complete design tasks.

[0047] Each subject was required to perform a design task using both the method proposed in the present invention and the baseline method. The task required the designer to collaborate with the generation model to create an automotive exterior design for the target consumer. Finally, the designer was asked to select the image they considered the most satisfactory and inspiring from all the generated automotive images as their conceptual design solution. The target consumer for Task 1 was a young female who liked elegant cars, while the target for Task 2 was a middle-aged male who liked powerful cars. The order in which the subjects used the two methods and the task settings were balanced throughout the experiment. At the beginning of the experiment, we introduced the experimental content and the usage skills of the two methods to the subjects, and then started two rounds of design tasks. Subsequently, with a clear understanding of the user profile and emotional needs of the target consumer, the subjects carried out design conceptions with the support of the given method, with a time limit of 40 minutes.

[0048] To collect the effectiveness of the proposed method in supporting creativity and enhancing the human-computer interaction experience, the questionnaire included the following seven indicators: design space exploration (whether the method can support designers to widely explore the relevant design space), design intention expression (whether the method can support designers to accurately express their design intentions), result matching degree (whether the generated content matches the designer's goals), process transparency (whether the method is transparent and interpretable in how the final design solution is obtained), result satisfaction (whether the designer is satisfied with the final design solution), ease of use (whether the method is easy to use), and usefulness (whether the method is useful in actual work). The subjects rated these indicators on a 5-point Likert scale (1: strongly disagree, 5: strongly agree). In addition, three indicators were selected from NASA-TLX to evaluate the task load (mental burden, effort level, and frustration), and they were rated on a 7-point scale, as shown in Table 3.

[0049] 。

[0050] After collecting the 32 design solutions created by all the participants, 25 female consumers aged between 25 - 35 (average age 28.84, variance 2.62) were recruited via social media to evaluate the emotional expressions of the design solutions generated in Task 1, and 25 male consumers aged between 35 - 45 (average age 40.84, variance 2.49) were recruited to evaluate the emotional expressions of the design solutions generated in Task 2. The evaluation of all indicators was conducted using a 5 - point Likert scale. For the results in the questionnaires filled out by the participants, a two - sample paired t - test was performed. For the evaluation results of consumers and experts, we conducted a two - sample t - test. Male consumers aged between 35 - 45 (average age 40.84, variance 2.49) evaluated the emotional expressions of the design solutions generated in Task 2. We also invited four experts with more than five years of design work experience to evaluate the aesthetic quality and novelty of the design solutions. The specific indicators included shape proportion, color coordination, unity, and innovativeness. The evaluation of all indicators was conducted using a 5 - point Likert scale. A two - sample paired t - test was performed on the results in the questionnaires filled out by the participants. A two - sample t - test was performed on the evaluation results of consumers and experts.

[0051] As shown in Table 4, the evaluation results of consumers' emotional expressions of the design solutions created using the present invention and the baseline method. It can be seen that the designs produced using the present invention more effectively conveyed the target emotions. In Task 1 (i.e., designing an "elegant" car), the results of the present invention were significantly better than the baseline method, while in Task 2 (i.e., designing a "powerful" car), the results of the present invention were slightly better than the baseline method, but no statistical significance was observed. Based on our observations of the design solutions generated for this task, the possible reason is that designers using both methods tended to create similar results. Many people chose to design an SUV, and rough lines, angular contours, and darker colors were frequently used in multiple design solutions. This may be attributed to the relatively limited design space associated with the emotional term "powerful".

[0052] 。

[0053] As the results in Table 5 show, the design solutions created using the method provided in this embodiment had better performance in terms of shape proportion, color coordination, design unity, and innovativeness, and were statistically significant in all these indicators. This indicates that the present invention can effectively help designers improve the aesthetic quality and innovativeness of the generated design solutions.

[0054] 。

[0055] As shown in Table 6, the feedback from the participating designers on the effectiveness of the method provided in this embodiment in supporting design concepts and enhancing the user experience is presented. The results indicate that compared with the baseline method, designers believe that the present invention can effectively promote a more extensive exploration of the design space and allow them to express their design intentions more easily and accurately. In addition, they point out that the design generation process using the present invention is significantly more transparent and easier to understand, and the generated results are more in line with their inputs. This improvement may be attributed to the fact that the present invention enables designers to provide more detailed and specific prompts to the generation model, thereby enhancing the consistency between the text prompts and the generated images. Designers show a higher level of satisfaction with the design solutions generated using the method provided in this embodiment, which is consistent with the evaluations of consumers and expert reviewers. In terms of ease of use, the method provided in this embodiment and the baseline method received the same score, while in terms of practicality in actual work, the present invention is considered more useful.

[0056] 。

[0057] Table 7 shows the statistical results of the main inspiration sources of the participants when generating design solutions using the two methods. The data shows that when using the method provided in this embodiment, all the inspiration of designers comes from the output of the method, especially the recommended design features and the generated images. When using the baseline method, approximately 44% of the participants obtained inspiration from external sources, including their own creativity and images searched on the Internet. This indicates that the method provided in this embodiment can provide more comprehensive creative support for designers. In contrast, the method provided in this embodiment helps designers save time on these activities and promotes a more extensive creative exploration and more effective inspiration generation.

[0058] 。

[0059] In summary, through two quantitative experiments, the performance of the kansei engineering engine proposed in the present invention in the translation of emotional words - design solutions and the extraction of design features is evaluated respectively, and through a user experiment, it is evaluated whether the present invention can help designers create appearance designs that better meet the emotional needs of consumers and have higher quality, and whether the present invention can enhance the experience of designers in the co - creation process with AI.

[0060] In addition, the terms "upper", "lower", "inner", "outer", "front", and "rear" are for descriptive purposes only and should not be construed as indicating or implying relative importance. Unless otherwise specifically stated, the relative steps, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of the present invention.

[0061] Certainly, the above are only specific embodiments of the present invention and are not intended to limit the scope of implementation of the present invention. Any equivalent changes or modifications made according to the structure, features, and principles described in the scope of the patent application of the present invention should be included in the scope of the patent application of the present invention.

[0062] Finally, it should be noted that the above embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that any person skilled in the art within the technical scope disclosed by the present invention can still modify the technical solutions described in the foregoing embodiments or easily conceive of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes, or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.

Claims

1. A method for generating automotive exterior design based on kansei engineering and generative artificial intelligence, characterized in that, It includes the following steps: Obtain the vehicle model review data and the set of emotional words describing the sensory perception of the vehicle model's appearance; Construct a semantic network based on the set of emotional words and the labels corresponding to the vehicle models; Construct a vehicle exterior design generation model under the framework of the semantic network and the large language model, including a keyword extraction module, an emotional divergence module, a design feature translation module, a visualization module, and an auxiliary comparison module: The keyword extraction module is used to extract the original emotional words in the input content; The emotional divergence module is used to perform similarity matching between the extracted original emotional words and the set of emotional words to output the top N divergent emotional words with the highest correlation; The design feature translation module constructs a corresponding unidirectional emotional word path according to the emotional words selected by the user from the top N divergent emotional words, and counts the occurrence times of the vehicle models corresponding to the relevant edges between two adjacent emotional words on the unidirectional emotional word path in the semantic network, and outputs the exterior images of the top M vehicle models with the most occurrences; Extract visual features from the exterior image selected by the user from the top M exterior images to output the text descriptions of the various components of the vehicle model in the exterior image; The visualization module performs text-to-image operations according to the text description output by the design feature translation module and the emotional words selected by the user from the top N divergent emotional words to generate a concept design drawing and perform visual output; The auxiliary comparison module compares the similarity between the original emotional words input by the user and the exterior image of the vehicle model with the generated concept design drawing to output a comparison result.

2. The method for generating an automotive exterior design based on kansei engineering and generative artificial intelligence according to claim 1, wherein, The expression of the semantic network is as follows: ; where K represents the set of sentiment words, V represents the set of vehicle models, and respectively represent the i th sentiment word and the j th sentiment word, represents the relevant edge between the i th sentiment word and the j th sentiment word. The weight value of the relevant edge is set according to the frequency of co-occurrence of the two sentiment words in different review texts.

3. The vehicle exterior design generation method based on kansei engineering and generative artificial intelligence according to claim 1, wherein the emotional divergence module performs similarity matching between the extracted original emotional words and the set of emotional words to output the target emotional word with the highest similarity value, and selects the top N divergent emotional words with the highest correlation starting from the target emotional word in the semantic network.

4. The method for generating an automotive exterior design based on kansei engineering and generative artificial intelligence according to claim 1, characterized in that, The visual feature extraction includes extracting features of the shape, color, and texture of the vehicle components in the exterior image.

5. The method for generating an automotive exterior design based on kansei engineering and generative artificial intelligence according to claim 4, characterized in that, The vehicle components include the vehicle's grille, air intake, front headlights, body, rearview mirror, wheels, door handles, and rear lights.

6. The method for generating an automotive exterior design based on kansei engineering and generative artificial intelligence according to claim 1, wherein The auxiliary comparison module includes a graph-to-graph comparison part and a text-to-graph comparison part; The graph-to-graph comparison part calculates the similarity between the visual feature vectors of each in the exterior image and the concept design drawing; The text-to-graph comparison part calculates the similarity between the divergent emotional words and the feature vectors of each region in the concept design drawing.

7. The method for generating an automotive exterior design based on kansei engineering and generative artificial intelligence according to claim 6, wherein The graph-to-graph comparison part uses the ViT model to perform different types of visual feature extraction on the exterior image and the concept design drawing to obtain the exterior visual feature vector, color visual feature vector, and texture visual feature vector of the vehicle in the image, and outputs the corresponding type as the image comprehensive feature vector; Reduce the obtained exterior visual feature vector, color visual feature vector, texture visual feature vector, and image comprehensive feature vector to a two-dimensional space to serve as the coordinates of the corresponding image in the two-dimensional space; The offset between the feature vectors of the appearance image and the conceptual design drawing in the two-dimensional space is used as a parameter in calculating the similarity to obtain the result of comparing the drawings.

8. The method for generating an automotive exterior design based on kansei engineering and generative artificial intelligence according to claim 7, characterized in that, In the part of comparing text with drawing, the CLIP model is used to extract features of the input divergent emotion words and the conceptual design drawing to obtain the keyword feature vectors and image feature vectors in the same dimension. The regional flipping algorithm is adopted and the similarity is calculated based on the keyword feature vectors and the image feature vectors to obtain the comparison result between the divergent emotion words and the conceptual design drawing. The specific process of the regional flipping algorithm is as follows: The conceptual design drawing is divided into multiple image blocks; The keyword feature vectors of the divergent emotion words are extracted through the CLIP model, as well as the global image feature vector corresponding to the complete conceptual design drawing and the local image feature vector of the conceptual design drawing with one image block masked; The first cosine similarity between the keyword feature vector and the global image feature vector is calculated, and the second cosine similarity between the keyword feature vector and the local image feature vector is calculated. The difference between the first cosine similarity and the second cosine similarity is used as the correlation score of the image block corresponding to the local image feature vector; The correlation scores of all image blocks are statistically analyzed and normalized to output the result of comparing text with drawing.

9. The method for generating an automotive exterior design based on kansei engineering and generative artificial intelligence according to claim 8, characterized in that, The calculation formula of the correlation score is as follows: ; ; In the formula, represents the correlation score between the I th image patch of the image i and the divergent emotion word, T represents the divergent emotion word, represents the original image, represents the inverted image corresponding to the i th image patch. The inverted image is created by masking the i th image patch in the complete conceptual design drawing, represents calculating the cosine similarity between the image I and the divergent emotion word T , is the image encoder of the CLIP model, is the text encoder of the CLIP model.

10. The method for generating an automotive exterior design based on kansei engineering and generative artificial intelligence according to claim 1, wherein In the semantic network, emotion words are used as nodes, and relevant edges connecting two nodes are constructed on the condition that two emotion words appear in the vehicle model review data of the same vehicle model appearance. The relevant edges include the number of occurrences of each vehicle model.

Citation Information

Patent Citations

  • Automobile shape design method and device, computer equipment and automobile

    CN115186368A

  • Auxiliary system for creative design of appearance of automobile tail lamp

    CN118194437A

  • FBS theory-based man-machine collaboration concept design generation method and system

    CN117874847A

  • Man-machine collaborative creation method and system supporting combined creativity and electronic equipment

    CN117932048A

  • Image material generation method and device, medium and computing equipment

    CN119850789A