A jeans appearance emotional preference analysis method

By using the latent Dirichlet distribution topic model and hierarchical analysis method, combined with the fuzzy comprehensive evaluation method, the efficiency and accuracy issues in the sentiment analysis of jeans appearance were solved, the user's multi-emotional needs were quantified, efficient design scheme evaluation was achieved, and user satisfaction and market fit were improved.

CN120561274BActive Publication Date: 2025-10-10TIANJIN POLYTECHNIC UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511045173.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-10-10
Estimated Expiration
2045-07-29

AI Technical Summary

Technical Problem

Existing technologies are inefficient and inaccurate in analyzing the emotional appearance of jeans. They find it difficult to quantify the priorities of users' multiple emotional needs and lack effective means for evaluating design solutions, resulting in long design cycles, increased costs, and low market fit.

Method used

The latent Dirichlet distribution topic model and hierarchical analysis method are combined with the fuzzy comprehensive evaluation method. The e-commerce platform review data is collected through crawler technology to construct a set of emotional keywords, quantify the keyword weights, and evaluate the emotional preference scores of the design schemes.

Benefits of technology

It improves the efficiency and accuracy of emotional keyword recognition, clarifies the priority of users' multi-emotional needs, realizes quantitative evaluation of design solutions, significantly improves user satisfaction and market fit, shortens the design cycle, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561274B_ABST
    Figure CN120561274B_ABST
Patent Text Reader

Abstract

The application provides a jeans appearance emotional preference analysis method, and belongs to the technical field of emotional analysis. Massive real and dynamic user online comment data is collected to obtain original comment data sets, and structured text corpus is formed after preprocessing. A latent Dirichlet allocation model is used to model the theme of the preprocessed text corpus, and a domain knowledge filtering mechanism is introduced to establish a jeans appearance optimal emotional keyword set. Then, the weight values of the emotional keywords in each theme are calculated. Finally, based on the weight calculation results of the emotional keywords, fuzzy comprehensive evaluation method is used to perform fuzzy operation on the candidate design scheme to obtain the emotional preference scores of different design schemes, so that quantitative evaluation and optimization of the design scheme are realized. The application can improve the efficiency and accuracy of emotional keyword recognition, clearly define the priority of user multi-emotional needs, output the emotional preference scores of the user on the product design scheme, and provide the optimal design scheme that meets the emotional needs of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of sentiment analysis, and in particular relates to a method for analyzing sentiment preferences for jeans appearance. Background Art

[0002] Against the backdrop of the deep integration of the digital economy and the experience economy, the apparel industry is undergoing an upgrade from providing functional products to creating emotional experiences. As a unique consumer product that combines the attributes of "physical medium" and "social symbol," clothing's visual appearance is a core dimension of the user's emotional experience.

[0003] Consumers have diverse emotional preferences for jeans, including preferences for style, color, material, and other aspects of appearance, which directly influence their purchasing decisions and brand loyalty. Statistics show that over 85% of online apparel purchases are driven by product visuals, and 76% of returns are attributed to the emotional difference between the actual product and the image. Therefore, a deep understanding of user emotional preferences is crucial for product design and development. However, current approaches to capturing and analyzing user emotional preferences present numerous challenges:

[0004] On the one hand, traditional methods of acquiring user emotional preferences are inefficient and inaccurate. For example, collecting data through questionnaires and other methods is not only time-consuming and labor-intensive, but also difficult to fully and accurately capture users' true emotions due to the limitations of questionnaire design and differences in user subjective understanding. Therefore, current sentiment mining technology based on online reviews has become an important way to gain user insights. Compared with traditional research methods, the massive user review data accumulated by e-commerce platforms has the advantages of real-time, spontaneity, and authenticity, providing an ideal data source for sentiment mining. However, the application of existing sentiment mining technology in the clothing field suffers from low efficiency, lack of accuracy, and lack of interpretability, which are specifically manifested as follows:

[0005] (1) Keyword extraction method based on term frequency-inverse document frequency (TF-IDF). TF-IDF achieves keyword screening by combining the frequency of each word in the document (TF) and the distribution frequency in the entire corpus (IDF). Its statistical characteristics lead to insufficient adaptability and accuracy for short texts of clothing reviews. The high-frequency word screening mechanism is difficult to capture low-frequency but high-semantic-value sentiment keywords in the clothing field, and ignores the semantic relevance between word vectors.

[0006] (2) Word2vec word vector method: Word2vec constructs word vector space through context prediction. However, in clothing review analysis, the local window mechanism makes it difficult to capture semantic associations across sentences. The sparsity of short texts leads to poor stability of word vector representation. In addition, the limited context information leads to a lack of interpretability of the results. Designers cannot directly understand the semantic clustering results in the vector space.

[0007] On the other hand, the existing research and method lack quantitative analysis of the priority of the multi-emotional needs of users. Users may have multiple emotional needs for jeans, such as pursuing fashion and comfort at the same time, but it is difficult to determine the relative importance of these different emotional needs in the minds of users at present, which makes it difficult for designers and clothing businesses to effectively use emotional analysis results. In addition, in the evaluation link of jeans design scheme, there is a lack of effective quantitative evaluation means, and designers often judge whether the design scheme meets the emotional needs of users according to experience, which may cause a large deviation between the design scheme and the actual needs of users, resulting in low market fit of products, long design cycle and increased cost. SUMMARY

[0008] The problem to be solved by the present application is to provide a jeans appearance emotional preference analysis method capable of improving the efficiency and accuracy of emotional keyword recognition, determining the priority of multi-emotional needs of users, and outputting the emotional preference score of users for product design scheme, so as to provide an optimal design scheme that meets the emotional needs of users, so as to significantly improve user satisfaction and product market fit, effectively shorten the design cycle, improve design quality and reduce cost.

[0009] To solve the above technical problems, the technical scheme adopted by the present application is: a jeans appearance emotional preference analysis method, comprising the following steps:

[0010] S1, constructing a user comment corpus: collecting online comment data of users on purchased jeans products in an e-commerce platform, creating a structured text corpus through preprocessing steps of word segmentation, stop word removal and part-of-speech tagging;

[0011] S2, establishing topic modeling and emotional keyword set: using latent Dirichlet allocation topic model to perform topic modeling on the preprocessed text corpus, for automatically identifying the latent topic structure in the text corpus, obtaining visualized "document-topic" distribution information and "topic-word" distribution information; filtering noise words by combining jeans domain knowledge, and screening to obtain emotional keywords under each topic, establishing an optimal emotional keyword set of jeans appearance;

[0012] S3, priority quantification of emotional keywords: based on the optimal emotional keyword set of jeans appearance, calculating the weight values of several emotional keywords under each topic by analytic hierarchy process, to obtain the priority ranking of different emotional keywords in user perception;

[0013] S4. Evaluation and Optimization of Design Schemes: Based on the weight calculation results of emotional keywords using the hierarchical analysis method, the candidate jeans design schemes composed of different emotional keywords were quantitatively evaluated and optimized using the fuzzy comprehensive evaluation method. The emotional preference scores of different jeans design schemes were calculated to determine the quality of their emotional design.

[0014] Furthermore, in step S1, the following steps are included:

[0015] S11. Use crawler technology to collect massive amounts of real, dynamic online user review data to obtain an original review dataset. The online user review data comes from a wide range of sources, including major e-commerce platforms, fashion forums, and jeans-related review areas.

[0016] S12, cleaning and organizing the original comment data of the original comment dataset, removing irrelevant characters, HTML tags, special symbols and other noise;

[0017] S13. Segment the cleaned and organized text data into words using a professional word segmentation tool to segment the text into independent words;

[0018] S14. Combine the general stop word list and the custom stop word list to remove high-frequency words that have no practical significance for sentiment analysis and are irrelevant to the appearance of jeans, and add a custom proprietary dictionary based on jeans domain knowledge;

[0019] S15. Tag each word with parts of speech, including nouns, verbs, adjectives, etc., for subsequent analysis.

[0020] Furthermore, in step S2, the following steps are included:

[0021] S21, read the text corpus, and use the latent Dirichlet allocation technology to perform topic modeling on the preprocessed text corpus;

[0022] S22. Set the initial number of topics k, the number of iterations, and the random seed, define the Dirichlet prior parameter α of the document-topic distribution and the Dirichlet prior parameter β of the topic-word distribution, and perform model training.

[0023] S23, model evaluation and tuning, traverse the topic number interval K, evaluate the fit of the model to the training data through perplexity, measure the semantic coherence and interpretability of the topics through consistency score, and determine the optimal number of topics by combining perplexity and consistency indicators;

[0024] S24. Based on the optimal number of topics, run and output visualized “document-topic” distribution information and “topic-word” distribution information;

[0025] S25. Select several keywords with the highest probability of appearing in each topic as candidate emotional keywords for the topic, and modify the emotional keywords under each topic based on domain knowledge. Remove words that exist in multiple topics, have similar meanings, are weakly associated, and have limited value for sentiment analysis, to obtain the optimal emotional keyword set for the appearance of jeans.

[0026] Furthermore, in step S3, the following steps are included:

[0027] S31. Based on the optimal emotional keyword set of the topic structure and the appearance of jeans, decompose it into a hierarchical structure of goals, criteria, and sub-criteria according to the decision-making goal. The user's emotional demand for the appearance of jeans is set as the goal layer, several topics are set as the criterion layer, and the emotional keywords contained in each topic are their respective sub-criteria layers, thereby establishing a hierarchical structure model.

[0028] S32. Obtain the distribution probability of each topic output by the latent Dirichlet topic model, and use the distribution probability as the initial weight of each topic in the criterion layer;

[0029] S33, comparing and rating the relative importance of the emotional keywords under each topic, and constructing a sub-criteria layer judgment matrix A using a 1-9 scaling method. The judgment matrix is ​​obtained by conducting a questionnaire survey on the target user group / experts to obtain the importance judgment of the emotional keywords;

[0030] S34. Perform consistency check on each judgment matrix A and eliminate matrices that do not meet the consistency threshold (CR < 0.1) to ensure the rationality of the judgment results;

[0031] S35. Aggregate several individual judgment matrices to obtain an integrated matrix for each topic Then, the geometric mean method is used to calculate the weight value of each emotional keyword under each topic, and a consistency test is performed. If CR < 0.1, it means that the judgment matrix passes the consistency test, thereby quantifying the relative importance of each emotional keyword in user perception.

[0032] Furthermore, in step S4, the following steps are included:

[0033] S41. Based on the calculation results of the emotional keyword weights, the required emotional keywords are selected from each theme according to the decision-making objectives to form candidate design solutions, and the factor set Y and weight vector W are determined. The factor set is the set of emotional keywords obtained by the previous screening, and the weight vector is the AHP weight calculation result of each emotional keyword;

[0034] S42. Establish an evaluation grade set V, clarify the evaluation grades, and assign values ​​to each evaluation grade, thereby establishing an evaluation grade value set C;

[0035] S43, evaluating each evaluation factor of the candidate solution according to the evaluation level, wherein the evaluation factor is an emotional keyword; counting the evaluation frequency of each evaluation factor at each evaluation level, and calculating the membership vector of each evaluation factor to obtain the membership matrix R of the solution;

[0036] S44. Obtain a comprehensive evaluation vector X of the scheme using a weighted average operator based on the scheme membership matrix R and the normalized weight vector W' of each evaluation factor.

[0037] S45. Calculate the emotional preference comprehensive score S of the candidate design schemes based on the comprehensive evaluation vector X and the evaluation grade assignment set C, obtain the corresponding emotional evaluation results of the design schemes, and select the design scheme with the highest score as the optimal design scheme that meets the user's emotional needs.

[0038] Due to the adoption of the above technical solution, the present invention has the following beneficial effects:

[0039] (1) Improve the efficiency, accuracy and interpretability of sentiment keyword recognition.

[0040] The present invention makes up for the shortcomings of traditional small sample data by collecting massive amounts of real review data and scientifically preprocessing it, thereby improving the authenticity and domain adaptability of the data. By combining the latent Dirichlet distribution topic modeling technology and the application of domain knowledge, it automatically identifies potential topic structures from high-dimensional user reviews and captures deep semantic information in the text. Even if the text is short, it can effectively identify its topic and quickly and accurately screen out keywords related to the emotional appearance of jeans, thereby improving recognition efficiency and accuracy compared to traditional methods.

[0041] (2) Clarify the priority of users’ multi-emotional needs.

[0042] This paper uses the AHP method to calculate the weight values ​​of emotional keywords, quantifying the relative importance of different emotional needs in the minds of users, enabling designers to clearly understand the user's preference order among multiple emotional needs and provide more targeted guidance for product design.

[0043] (3) Achieve quantitative evaluation and optimization of design solutions.

[0044] This method uses a fuzzy comprehensive evaluation method to evaluate candidate design solutions and output a specific emotional preference score. This allows for intuitive comparison of the pros and cons of different design solutions, thereby selecting the one that best meets the user's emotional needs, significantly improving user satisfaction and product-market fit. Furthermore, by more accurately understanding user needs, it effectively shortens the design cycle, improves design quality, and reduces costs.

[0045] In summary, the present invention can improve the efficiency and accuracy of emotional keyword recognition, clarify the priority of users' multiple emotional needs, realize high-precision capture of users' implicit emotional needs, and output users' emotional preference scores for product design schemes, provide optimal design schemes that meet users' emotional needs, significantly improve user satisfaction and product-market fit, effectively shorten the design cycle, improve design quality, reduce costs, and provide method support for the realization of accurate and efficient emotional product design and manufacturing, and marketing and personalized recommendations based on user preferences. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] The present invention will be described in detail below with reference to the accompanying drawings and in combination with examples, and the advantages and implementation modes of the present invention will become more apparent. The contents shown in the accompanying drawings are only used to illustrate the present invention and do not constitute any limitation to the present invention. In the accompanying drawings:

[0047] Figure 1 It is a flow chart of the present invention. DETAILED DESCRIPTION

[0048] like Figure 1 As shown, the present invention provides a method for analyzing emotional preferences for jeans appearance, comprising the following steps:

[0049] S1. Build a user review corpus: Collect online review data from users on the e-commerce platform about the jeans they purchased. Create a structured text corpus through pre-processing steps such as word segmentation, stop word removal, and part-of-speech tagging.

[0050] Among them, the online review data of jeans users are obtained through crawler technology.

[0051] S2. Establishing Topic Modeling and Sentiment Keyword Sets: We use the Latent Dirichlet Allocation (LDA) topic model to perform topic modeling on the preprocessed text corpus. This automatically identifies the underlying topic structure within the text corpus and generates visual "document-topic" and "topic-word" distribution information. We also filter out noise words by combining domain knowledge about jeans, identifying sentiment keywords for each topic, and establishing an optimal sentiment keyword set for jeans appearance.

[0052] The latent Dirichlet distribution topic model is obtained by calling the models.LdaModel module of the Python gensim library; the visual "document-topic" distribution information and "topic-word" distribution information are obtained by calling the pyLDAvis interactive tool in Python.

[0053] S3, priority quantification of emotional keywords: based on the optimal emotional keyword set of the appearance of jeans, the weight values of several emotional keywords under each theme are calculated by the analytic hierarchy process (AHP), and the priority ranking of different emotional keywords in user perception is obtained;

[0054] AHP is a decision-making method that decomposes elements related to decision-making into levels such as goals, criteria, and schemes, and conducts qualitative and quantitative analysis based on this.

[0055] S4, evaluation and optimization of design scheme: based on the weight calculation results of emotional keywords by AHP, the fuzzy comprehensive evaluation method is used to quantitatively evaluate and optimize the candidate jeans design scheme composed of different emotional keywords, and the emotional preference score of different jeans design schemes is calculated to judge the emotional design level, the higher the score, the higher the emotional satisfaction level of the scheme;

[0056] Fuzzy comprehensive evaluation method is a comprehensive evaluation method based on fuzzy mathematics, which can make a general evaluation of things or objects subject to multiple factors.

[0057] In step S1, the following steps are included:

[0058] S11, use crawler technology to collect massive real and dynamic user online comment data to obtain the original comment data set; user online comment data is widely sourced from various e-commerce platforms, fashion forums and other comment areas related to jeans;

[0059] This embodiment is based on the jeans comment data of a certain e-commerce platform. Specifically, using Python web crawler technology, search with "jeans" as the keyword, collect the comment data about users' comments on the purchased jeans in the past year, a total of 8910 comment data, covering jeans of different brands, styles and price ranges, to obtain the original comment data set, stored as an Excel file.

[0060] S12, clean and organize the original comment data of the original comment data set, delete irrelevant characters, HTML tags, special symbols and other noise;

[0061] In this embodiment, the original comment data set is cleaned, and duplicate data, invalid content and outliers are deleted, and secondary comments, too short content, HTML tags, special symbols and null values are removed.

[0062] S13, segment the text data after cleaning and organizing, and use professional segmentation tools to segment the text into independent words;

[0063] In this embodiment, the jieba word segmentation tool of Python is used to perform word segmentation processing on the text data, and the text is divided into independent words and stored as an Excel file.

[0064] S14. Combine the general stop word list and the custom stop word list to remove high-frequency words that have no practical significance for sentiment analysis and are irrelevant to the appearance of jeans, and add a custom proprietary dictionary based on jeans domain knowledge;

[0065] In this embodiment, a stop word list is constructed by combining the stop word list of Harbin Institute of Technology and custom stop words (words that are not related to the appearance of jeans, such as "cheap", "buy", "express delivery", "customer service", "free shipping", etc.), stop words are read from the pre-constructed stop word list, and stop words in the text after word segmentation (high-frequency but meaningless words, including numbers, punctuation marks, and function words such as "this", "of", "le", and "is"); domain knowledge of jeans is introduced to add a custom proprietary dictionary, such as "slimming", "lengthening legs", and "lengthening proportions" that appear many times in comments, to improve domain adaptability.

[0066] S15. Tag each word with its part of speech, including noun, verb, adjective, etc., for subsequent analysis; for example, "jeans" is tagged as a noun, and "fashion" is tagged as an adjective. After the above preprocessing steps, a structured text corpus is formed.

[0067] In this embodiment, the posseg function provided by Jieba is used to perform part-of-speech tagging on the text after removing stop words, including verbs, nouns, and adjectives.

[0068] Wherein, in step S2, the following steps are included:

[0069] S21, read the text corpus, and use the latent Dirichlet allocation technology to perform topic modeling on the preprocessed text corpus;

[0070] In this embodiment, topic modeling is performed by calling the models.LdaModel module of the Python gensim library.

[0071] S22. Set the initial number of topics k (default value, which can be optimized in step S23), the number of iterations, the random seed, define hyperparameters such as the Dirichlet prior parameter α of the document-topic distribution and the Dirichlet prior parameter β of the topic-word distribution, and perform model training.

[0072] In this embodiment, the initial number of topics k is set to 2, the number of iterations is set to 50, the random seed is set to 0, and the Dirichlet prior parameter α of the document-topic distribution is defined to be 0.1, and the Dirichlet prior parameter β of the topic-word distribution is defined to be 0.01.

[0073] S23, model evaluation and tuning, traverse the topic number interval K, evaluate the fit of the model to the training data through perplexity, measure the semantic coherence and interpretability of the topics through consistency score, and determine the optimal number of topics by combining perplexity and consistency indicators;

[0074] In this embodiment, an integer in the interval [2-14] is selected as the candidate number of topics for training, and the topic number interval K∈[2-14] is traversed. The optimal number of topics k=4 is determined based on the perplexity and consistency indicators;

[0075] The perplexity calculation formula is:

[0076]

[0077] Where D represents the document set of the corpus; M represents the number of documents; Represents the words in document d; N d represents the total number of words in each document d; P( ) represents the word in the document Probability of occurrence.

[0078] Consistency calculation formula:

[0079]

[0080] In the formula, k represents the number of topics, m is the number of high-frequency keywords selected in each topic, It represents the probability value of the jth most frequent keyword in the i-th topic.

[0081] S24. Based on the optimal number of topics, run and output the visualized "document-topic" distribution information and "topic-word" distribution information; the visualized "document-topic" distribution information and "topic-word" distribution information are obtained by calling the pyLDAvis interactive tool in Python.

[0082] In this embodiment, based on the optimal number of topics obtained above, the probability distribution of the four topics in the document and the top 30 keywords in each topic are output; taking topic I as an example, its distribution probability in the entire document is 0.3827, and the keywords include "retro", "fashion", "breathable", "soft", "slim", "show leg length", "trendy", "washed", "faded", "straight", "exquisite", "printed", etc.

[0083] S25. Select several keywords with the highest probability of appearing in each topic as candidate emotional keywords for the topic, and modify the emotional keywords under each topic based on domain knowledge. Remove words that exist in multiple topics, have similar meanings, are weakly associated, and have limited value for sentiment analysis, to obtain the optimal emotional keyword set for the appearance of jeans.

[0084] In this embodiment, the top 10 keywords with the highest appearance probability in each topic are selected as candidate sentiment words of the topic. The sentiment keywords in each topic are modified in combination with the knowledge in the field of jeans, and the words existing in multiple topics, with similar meanings, weak correlations and limited value for sentiment analysis are removed. Taking topic I as an example, “fashion” and “trend” have similar meanings, and therefore only the keyword “fashion” with a higher probability is retained, and the keywords “water washing”, “fading”, and “straight tube” describing the specific design features of jeans are removed. Finally, the optimal sentiment keyword set of the appearance of jeans includes 20 sentiment keywords, topic I includes “retro”, “fashion”, “breathable”, “soft”, “body shaping”, and “leg length revealing”; topic II includes “simple”, “versatile”, “loose”, “slim”, and “natural”; topic III includes “durable”, “leisure”, “creative”, and “solid”; and topic IV includes “comfortable”, “light”, “skin-friendly”, “elegant”, and “warm”.

[0085] In step S3, the following steps are included:

[0086] S31, based on the topic structure and the optimal sentiment keyword set of the appearance of jeans, the target, criterion, sub-criterion and other hierarchical structures are decomposed according to the decision target. Specifically, the emotional needs of users for the appearance of jeans are set as the target layer, a plurality of topics are set as the criterion layer, and the sentiment keywords contained in each topic are set as the respective sub-criterion layer, and a hierarchical structure model is established.

[0087] In this embodiment, the criterion layer is topic I, topic II, topic III and topic IV, and the sub-criterion layer is the sentiment keywords contained in each topic.

[0088] S32, the distribution probability of topics I-IV output by the latent Dirichlet topic model is obtained, and the distribution probability is taken as the initial weight of each topic in the criterion layer.

[0089] In this embodiment, the distribution probability of topic I is 0.3827, the distribution probability of topic II is 0.3204, the distribution probability of topic III is 0.1718, and the distribution probability of topic IV is 0.1252.

[0090] S33, the relative importance of the sentiment keywords under each topic is compared and rated, and the 1-9 scale method is used to construct the sub-criterion layer judgment matrix A:

[0091]

[0092] In the matrix, n represents the number of elements (sentiment keywords) participating in the pairwise comparison, the order of the judgment matrix is n, and the dimension of the matrix (n x n) is determined.

[0093] For example, for the two emotional keywords "fashion" and "comfort", users / experts are asked to judge which one is more important when buying jeans, and the proportion of their importance. The judgment matrix obtains the importance judgment of emotional keywords by conducting a questionnaire survey on the target user group / experts;

[0094] In this embodiment, one of the judgment matrices A in Topic I can be obtained as follows:

[0095] .

[0096] S34. Perform consistency check on each judgment matrix A and eliminate matrices that do not meet the consistency threshold (CR < 0.1) to ensure the rationality of the judgment results;

[0097] In this embodiment, each judgment matrix passes the consistency test.

[0098] S35. Aggregate several individual judgment matrices to obtain an integrated matrix for each topic ,

[0099]

[0100] Where P represents the number of experts, represents the relative importance ratio given by the pth user / expert to element i and element j.

[0101] In this embodiment, the integration matrix of topic I can be obtained for:

[0102]

[0103] Then, the geometric mean method is used to calculate the weight value of each emotional keyword under each topic, and a consistency test is performed. If CR < 0.1, it means that the judgment matrix passes the consistency test, thereby quantifying the relative importance of each emotional keyword in user perception; for example, the weight value of "fashion" is calculated to be 0.6, and the weight value of "comfort" is 0.4, indicating that fashion is relatively more important in the user's emotional demand for the appearance of jeans.

[0104] The geometric mean weight calculation formula is:

[0105]

[0106] Where W i represents the weight of the i-th element, a ij It is the element in the i-th row and j-th column of the judgment matrix A, which represents the ratio of the importance of the i-th element to the j-th element.

[0107] In this embodiment, the weight of each keyword in the optimal emotional keyword set of jeans is calculated, that is, n=20, and the weight value of "retro" in theme I is calculated to be 0.1568, the weight value of "fashion" is 0.1922, the weight value of "breathable" is 0.1487, the weight value of "soft" is 0.1171, the weight value of "slim" is 0.2047, and the weight value of "showing long legs" is 0.1804. In theme II, the weight value of "simple" is 0.2199, the weight value of "versatile" is 0.2367, and the weight value of "loose" is 0. The weight value of "slim" in theme III is 0.1722, the weight value of "slim" is 0.1879, and the weight value of "natural" is 0.1832; the weight value of "durable" in theme III is 0.3018, the weight value of "casual" is 0.2558, the weight value of "creativity" is 0.2111, and the weight value of "sturdy" is 0.2313; the weight value of "comfort" in theme IV is 0.2474, the weight value of "lightness" is 0.1981, the weight value of "skin-friendly" is 0.2079, the weight value of "elegance" is 0.1826, and the weight value of "warmth" is 0.164.

[0108] Consistency CR calculation formula:

[0109]

[0110] In the formula, CR is defined as the ratio of consistency index (CI) to random index (RI), , RI is the randomness index, the specific values ​​are shown in Table 1, λ max Represents the maximum eigenvalue of the judgment matrix.

[0111]

[0112] Where, To judge the integration matrix, W i represents the weight of the i-th element, is a matrix The i-th component of .

[0113]

[0114] In this embodiment, the four subject judgment integrated matrix They were 6.0274, 5.0103, 4.0008, and 5.004 respectively, with CIs of 0.0055, 0.0026, 0.0003, and 0.001 respectively, and CRs of 0.0043, 0.0023, 0.0003, and 0.0009 respectively.

[0115] Wherein, in step S4, the following steps are included:

[0116] S41. Based on the calculation results of emotional keyword weights, select the required emotional keywords from each topic according to the decision-making objectives to form candidate design solutions, clarify the factor set and weight vector, the weight represents the importance of a certain evaluation factor to the evaluation result, and the size of the weight has a direct impact on the final evaluation result. Construct the weight vector W, W={w1,w2,…,w n′}, factor set Y, Y={y1,y2,…,y n′}, where n′ is the number of evaluation factors (selected emotional keywords), w i represents the weight of the i-th evaluation factor, y i is the i-th evaluation factor.

[0117] The factor set Y is the set of selected emotional keywords, such as "slimming" and "comfort", and the weight vector W is the AHP weight calculation result corresponding to the selected emotional keywords.

[0118] In this embodiment, a scheme Q is set, the number of evaluation factors n′=4, and the scheme factor set Y Q ={slim, versatile, durable, comfortable}, weight vector W Q ={0.2047,0.2367,0.3018,0.2474};

[0119] S42. Establish an evaluation level set V, clarify the evaluation level, and assign a value to each evaluation level. The evaluation level represents the various evaluation results of the evaluated factors. Given a finite set V to represent the evaluation level set, V={v1,v2,…,v f}, v is the evaluation level, f is the number of evaluation levels, and the specific number of levels is determined according to actual needs; then assign a value to each evaluation level and establish an evaluation level assignment set C={c1,c2,…,c f}, c is the evaluation grade assignment.

[0120] In this embodiment, f=5, V Q ={v1,v2,v3,v4,v5}={excellent, good, average, poor, very poor}, the evaluation grade assignment set is C Q ={c1,c2,c3,c4,c5}={90,70,50,30,10}.

[0121] S43. Evaluate each evaluation factor of the candidate solution according to the evaluation level, where the evaluation factors are emotional keywords; count the evaluation frequency of each evaluation factor at each level V, and calculate the membership vector of each evaluation factor to obtain the solution's membership matrix (fuzzy relationship matrix) R:

[0122]

[0123] Where rij is the element of the membership matrix R (dimension n' x f) in the i-th row and j-th column, representing the membership of the i-th evaluation factor to the j-th evaluation level v j , i = 1, 2, …, n', j = 1, 2, …, f.

[0124] In this embodiment, the evaluation factor n' includes 4 emotional keywords, namely self-cultivation, all-match, wear-resistant and comfortable, f = 1, 2, 3, 4, 5, corresponding to excellent, good, general, poor and very poor in the evaluation level, for the design scheme Q, the evaluation group's evaluation result of the "self-cultivation" factor is that 7 people choose "excellent", 5 people choose "good", 5 people choose "general", 0 people choose "poor" and "very poor", and the corresponding membership vector is [0.4118 0.2941 0.2941 0 0], and the membership vectors of all evaluation factors are calculated in the same way. These evaluation results are converted into the membership matrix R, and the membership matrix of the scheme Q is:

[0125] .

[0126] S44, for the weight vector W, the normalization condition must be satisfied, therefore the original weight vector W is modified, the original weight of each evaluation factor is divided by the weight sum, to obtain the normalized weight vector W', according to the scheme membership matrix R and the normalized weight vector W', the comprehensive evaluation vector X of the scheme is obtained by the weighted average operator:

[0127]

[0128] In the formula, W' is the normalized weight row vector (dimension 1 x n'), is the normalized weight of the i-th evaluation factor (i = 1, 2, …, n'), x j is the comprehensive membership of the evaluation object to the j-th evaluation level.

[0129] In this embodiment, the normalized weight vector is obtained according to and the membership matrix R Q , that is, the comprehensive evaluation vector X Q :

[0130] .

[0131] S45, according to the comprehensive evaluation vector X and the evaluation level value set C, the emotional preference comprehensive score S of the candidate design scheme is calculated, and the emotional evaluation result of the scheme is obtained:

[0132]

[0133] Where C T Represents the transpose of the evaluation grade assignment set C.

[0134] For example, for design option A', the emotional preference score obtained after calculation is 65 points, option B' is 72 points, and option C' is 58 points. By comparing the scores, design option B' is selected as the optimal design option to meet the user's emotional needs.

[0135] In this embodiment, the sentiment preference score S of solution Q is Q for:

[0136]

[0137] This will determine the final evaluation results.

[0138] In the present invention, LDA: a probabilistic topic modeling technology, used to mine potential topic structures from text data. AHP: hierarchical analysis method, a multi-criteria decision-making method, which quantifies the relative importance (weight) of indicators (emotional keywords) by constructing a judgment matrix. FCE: fuzzy comprehensive evaluation method, a quantitative evaluation method based on fuzzy mathematics, used to deal with uncertainty problems, to verify the effectiveness of obtaining emotional preferences, and to achieve emotional satisfaction evaluation and optimization of design schemes. Emotional keywords: words related to the appearance of jeans that can reflect the user's emotional preferences. Perplexity: the evaluation level of the topic model, which measures the degree of fit of the model to the training data. The smaller the value, the stronger the model's prediction ability. Consistency: a quantitative indicator of topic interpretability, which reflects the semantic consistency of words within the topic. The higher the value, the greater the clarity and interpretability of the topic. Document-topic distribution information: represents the probability distribution of each topic in each document. Topic-word distribution information: represents the probability distribution of each word in each topic.

[0139] This invention can be applied to: Apparel brand product design: providing women's clothing brands with emotional preference data on jeans appearance to support design optimization; Personalized recommendations on e-commerce platforms: improving recommendation algorithms based on user emotional preferences to increase recommendation conversion rates; Market trend forecasting: analyzing dynamic changes in review data to predict jeans appearance trends.

[0140] The embodiments of the present invention are described in detail above, but the contents are only preferred embodiments of the present invention and should not be considered to limit the scope of the present invention. All equivalent changes and improvements made within the scope of the present invention should still fall within the scope of the present invention.

Claims

1. A method for analyzing emotional preferences for jeans appearance, characterized by: The following steps are involved: S1. Build a user review corpus: Collect online review data from users on the e-commerce platform about the jeans they purchased. Create a structured text corpus through pre-processing steps such as word segmentation, stop word removal, and part-of-speech tagging. S2. Establishing Topic Modeling and Sentiment Keyword Sets: We use the Latent Dirichlet Allocation topic model to perform topic modeling on the preprocessed text corpus. This is used to automatically identify the underlying topic structure within the text corpus and obtain visual document-topic distribution information and topic-word distribution information. We also filter out noise words by combining domain knowledge about jeans, identifying sentiment keywords under each topic, and establishing an optimal sentiment keyword set for jeans appearance. S3. Prioritization of emotional keywords: Based on the optimal emotional keyword set for jeans appearance, the weight values ​​of several emotional keywords under each theme are calculated using the hierarchical analysis method to obtain the priority ranking of different emotional keywords in user perception; S4. Evaluation and optimization of design schemes: Based on the weight calculation results of emotional keywords using the hierarchical analysis method, the candidate jeans design schemes composed of different emotional keywords are quantitatively evaluated and optimized using the fuzzy comprehensive evaluation method. The emotional preference scores of different jeans design schemes are calculated to judge the quality of their emotional design.

2. The method for analyzing emotional preference for jeans appearance according to claim 1, characterized in that: In step S1, the following steps are included: S11. Use crawler technology to collect massive amounts of real and dynamic user online review data to obtain the original review dataset; S12, cleaning and organizing the original comment data of the original comment dataset, deleting irrelevant characters, HTML tags, and special symbols; S13, performing word segmentation on the cleaned and organized text data, using a word segmentation tool to segment the text into independent words; S14. Combine the general stop word list and the custom stop word list to remove high-frequency words that have no practical significance for sentiment analysis and are irrelevant to the appearance of jeans, and add a custom proprietary dictionary based on jeans domain knowledge; S15. Tag each word with part of speech for subsequent analysis.

3. The method for analyzing emotional preference for jeans appearance according to claim 1, characterized in that: In step S2, the following steps are included: S21, read the text corpus, and use the latent Dirichlet allocation technology to perform topic modeling on the preprocessed text corpus; S22. Set the initial number of topics, the number of iterations, and the random seed, define the Dirichlet prior parameters for the document-topic distribution and the Dirichlet prior parameters for the topic-word distribution, and perform model training. S23, model evaluation and tuning, traverse the range of topic numbers, evaluate the degree of fit of the model to the training data through perplexity, measure the semantic coherence and interpretability of the topics through consistency score, and determine the optimal number of topics by combining perplexity and consistency indicators; S24. Based on the optimal number of topics, run and output visualized document-topic distribution information and topic-word distribution information; S25. Select several keywords with the highest probability of appearing in each topic as candidate emotional keywords for the topic, and modify the emotional keywords under each topic based on domain knowledge. Remove words that exist in multiple topics, have similar meanings, are weakly associated, and have limited value for sentiment analysis, to obtain the optimal emotional keyword set for the appearance of jeans.

4. The method for analyzing emotional preference for jeans appearance according to claim 1, characterized in that: In step S3, the following steps are included: S31. Based on the optimal emotional keyword set of the topic structure and the appearance of jeans, the user's emotional demand for the appearance of jeans is set as the target layer, several topics are set as the criterion layer, and the emotional keywords contained in each topic are their respective sub-criterion layers, thereby establishing a hierarchical structure model; S32. Obtain the distribution probability of each topic output by the latent Dirichlet topic model, and use the distribution probability as the initial weight of each topic in the criterion layer; S33, comparing and rating the relative importance of the emotional keywords under each topic, and constructing a sub-criteria layer judgment matrix using a 1-9 scaling method, wherein the judgment matrix is ​​obtained by conducting a questionnaire survey on the target user group or experts to obtain the importance judgment of the emotional keywords; S34. Perform consistency check on each judgment matrix and eliminate matrices that do not meet the consistency threshold to ensure the rationality of the judgment result; S35. Aggregate several individual judgment matrices to obtain an integrated matrix for each topic, then use the geometric mean method to calculate the weight values ​​of several emotional keywords under each topic, and perform a consistency test. If CR < 0.1, it means that the judgment matrix passes the consistency test, thereby quantifying the relative importance of each emotional keyword in user perception.

5. The method for analyzing emotional preference for jeans appearance according to claim 1, characterized in that: In step S4, the following steps are included: S41. Based on the calculation results of the emotional keyword weights, select the required emotional keywords from each theme according to the decision-making goal to form a candidate design scheme, and clarify the factor set and weight vector. The factor set is the emotional keyword set obtained by the previous screening, and the weight vector is the AHP weight calculation result of each emotional keyword; S42, establishing an evaluation grade set, clarifying the evaluation grade, and assigning a value to each evaluation grade to establish an evaluation grade assignment set; S43, evaluating each evaluation factor of the candidate solution according to the evaluation level, wherein the evaluation factor is a sentiment keyword; counting the evaluation frequency of each evaluation factor at each evaluation level, and calculating the membership vector of each evaluation factor to obtain a membership matrix of the solution; S44, obtaining a comprehensive evaluation vector of the scheme by using a weighted average operator based on the scheme membership matrix and the normalized weight vectors of the evaluation factors; S45. Calculate the emotional preference comprehensive score of the candidate design schemes based on the comprehensive evaluation vector and the evaluation grade assignment set, obtain the corresponding emotional evaluation results of the design schemes, and select the design scheme with the highest score as the optimal design scheme that meets the user's emotional needs.

Citation Information

Patent Citations

  • User demand comprehensive analysis method based on online user comment data

    CN120146699A

  • KR20220107399A